All work

ESG & sustainabilityIn production

Institutional ESG Analytics Built an AI Filing Intelligence Platform

How analysts stopped rebuilding the same research out of thousand-page filings, with every extracted number carrying a link back to the page it was read from.

Challenge
ESG data is trapped in long PDFs whose layouts defeat naive extraction, and search returns fragments that cannot tell an analyst whether a number belongs to the right entity, year, unit or framework.
Solution
A pipeline that parses without flattening, extracts with per-field confidence so uncertain values route to review, maps to frameworks with versioned mappings, and links facts across reports.
Impact
Analysis cycles compressed from weeks to hours, with every metric resolving back to the source page it was read from.
Status
In production. The cycle-time change is an internal working result not yet reconfirmed on current volumes.
Weeks → hours
ESG analysis cycle

Internal working result from the build. Not yet reconfirmed with the client on current document volumes.

Sent to review
Low-confidence extractions

Every extracted field carries a confidence score, and an uncertain one is flagged rather than entering a dashboard as accepted data. Demonstrable in the product.

To the page
Lineage on every metric

Each number links back to the source page it was read from. Verify coverage on your own corpus before treating lineage as universal.

Stack

Layout-aware OCR + NLPFramework mapping (SDG, GRI, CSRD)Knowledge graph + vector searchGoverned analytics, full lineage

Hand an institution's densest filings to a general chatbot and it will answer confidently and wrongly. That is the trap DocVerse set out to avoid.

The job was never to summarise sustainability reports. It was to ground real analysis in them, with every number traceable back to the page it came from — because in institutional work, a number you cannot source is a number you cannot use.

About the Client

Client
Institutional ESG Analytics
Industry
ESG & institutional
Stage
Institutional · ESG
Service
Document Intelligence

DocVerse transforms complex, unstructured documents — sustainability reports, ESG disclosures, annual filings — into structured, decision-ready intelligence. It understands domain context, connects information across documents and across years, and serves it to analysts, policymakers and enterprise teams. It is a structured platform rather than a chat box: you build a report, pin it to a dashboard, and spin up custom versions.

The old model

Most ESG and impact data is trapped in long PDFs, scans and slide decks with inconsistent layouts, tables, annexures and footnotes that defeat naive extraction. So analysts do the work by hand: find the metric, copy it into a spreadsheet, note where it came from if there is time, and move on.

That is slow and error-prone, and it is almost impossible to audit afterwards. But the extraction problem is only the first layer, and it is not the expensive one.

The mapping problem sits on top of it. Aligning disclosures to reporting frameworks — the UN SDGs, GRI, CSRD, a client's internal KPIs — is inconsistent across versions, entities and years. So even perfectly clean data is hard to compare, because two companies reporting the same underlying fact against different framework versions have not reported the same number.

And because there is no reliable single view across entities, geographies or time, audits, board reporting and impact tracking are painful. Search does not solve this. Search returns fragments, and a fragment does not tell you whether the number you found belongs to the right company, the right year, the right units, or the right framework. All four have to be right, and being wrong about any one of them produces a figure that looks entirely plausible.

A wrong number travels a long way before anyone catches it. That is the actual risk being managed here.

  • Ingest mixed-format documents at volume into one structured model

  • Extract metrics, units, entities and periods with layout awareness and per-field confidence

  • Map every signal to frameworks with traceable, version-controlled mappings

  • Serve it under governance, so no figure changes silently and every one is verifiable

What changed

We built a complete ESG-intelligence pipeline: ingest unstructured documents at scale, extract structured metrics with ESG context, map them to global frameworks, and serve them through a controlled analytics layer.

The principle underneath is lineage. A number is only as useful as your ability to prove where it came from, so the system is built so that proof is never lost. Every metric carries a link back to the exact page it was read from, and a reviewer can always open the source.

The second decision is the one that changes how analysts behave. Every extracted field carries a confidence score, and a low-confidence field goes to review rather than into a dashboard as accepted data. That sounds like a small quality feature and is actually the thing that makes the output reusable. An analyst who knows uncertain figures are flagged can build on last quarter's work; an analyst who cannot tell which figures were shaky re-does the extraction to be safe, which is exactly the work the system was supposed to remove.

New filings are also fact-checked against what the platform already knows, and a deep-research mode reasons across the whole corpus rather than answering from a single retrieval — because the question an analyst actually has is usually comparative, and a single retrieval cannot answer a comparison.

Design decision: generated analysis must keep its citations and source-page links. An unsupported conclusion is not presented as a fact, even when it is probably right — "probably right" is not a category institutional work has room for.

How the extraction, mapping and lineage layers are built

The new workflow

Documents come in as they are

Layout-aware parsing takes each document through intake, OCR and language processing, producing structured chunks, sections, tables and footnotes — machine-readable without flattening the structure that gave it meaning.

Extraction carries its own uncertainty

Entities, metrics, units, time periods, tables and narrative claims come out with a confidence score attached, so an uncertain field is a flag rather than a silent guess.

Everything maps to a framework

Classifiers and rules align KPIs and statements to SDGs, GRI, CSRD or a custom framework, with traceable mappings — which is what makes one report comparable to another across years and entities.

Facts connect across documents

Entities and KPIs link into a knowledge graph, paired with vector search, so an analyst can ask cross-report and time-series questions instead of reading each filing in isolation.

Analysis arrives with its evidence

Dashboards, heatmaps, benchmarks and anomaly views over a versioned store, with every figure resolving back to a source page.

Control and evidence

What can be used as a fact

Source authority, document metadata, reporting period, framework mapping, role-based access and lineage together determine which claims and metrics can be used at all. Low-confidence extraction is flagged for review rather than accepted. Generated analysis retains citations and source-page links instead of presenting unsupported conclusions as findings.

What runs, in what order

Intake, layout-aware OCR, structured extraction, framework mapping, graph construction, vector retrieval, analysis, fact-checking, presentation. For a research question the system decomposes it, selects sources, combines graph and vector evidence, constructs claims, checks them, and attaches provenance — rather than retrieving once and generating.

What you can see afterwards

The observable unit is the claim-to-evidence path. A reviewer can inspect the document, the page, the parsed structure, the extracted entity or metric, its confidence, the retrieval decision, the framework mapping, and the generated conclusion. Versioned data and audit logs let two revisions be compared, which is how you find out that a number moved because a document was restated rather than because a company changed.

Impact

A pile of dense filings became something an analyst can act on. Multi-week manual extraction and spreadsheet work became an automated pipeline with dashboards that refresh in hours, even at high document volume. And the governance layer made each metric verifiable rather than something the analyst had to accept on trust.

The reuse effect is the compounding one. Because uncertain figures are visibly flagged, teams stopped re-deriving prior work defensively — which is where most of the recovered time actually comes from.

  • Analysis cycles compressed from weeks to hours, with manual copy-paste replaced by an automated pipeline

  • Every answer citation-grounded and fact-checked, over genuinely hard filings rather than clean text

  • Governance by construction, with role-based access, audit trails and source lineage on each metric

Reading the numbers

What is directly demonstrable in the product: grounded research with citations, fact-checking against the existing knowledge base, custom evaluations, the generated analytical interfaces, per-field confidence, and human review of low-confidence extraction.

The weeks-to-hours cycle-time change is an internal working result from the build and has not been reconfirmed with the client on current volumes. We would also treat lineage as a claim to verify rather than assume — it holds in the product as designed, and "every metric, always" is the kind of universal that deserves checking against your own corpus before it is relied on.

For your own estimate: documents per cycle × average review time × cycles per year × loaded analyst cost, minus verification time. The term worth measuring separately is how much prior work your team currently re-derives because they cannot tell which figures were solid. In our experience that is the larger number and the one nobody tracks.

Read the engineering write-up: parsing, confidence, framework mapping and the claim-to-evidence path

Want the version of this built for you?

We can walk you through Institutional ESG Analytics Built an AI Filing Intelligence Platform live — the architecture, the failure modes, and what we would change for your constraints. Tell us what you are building and we will come back with a concrete plan.

Reply within 2h