How the scoring works

One question per topic: do the financial statements reflect what the outside report leads a reader to expect, and if not, is there a good reason? The model never sees both reports in the same call. It reads the outside report first and writes down what it found and what that implies for the financial statements; only then does it check the financial statements against that note. Code verifies every quote.

OUTSIDEsustainability / management reportINSIDEaudited FS notes, in fullSTAGE 1 · MODELOutside readone call per firm-year, 14 topicsdiscussed? · key elements · quotesone expected FS reflection, or nonecode: verify outside quotesTOPIC PROFILE · FROZENoutside_discussedkey_elements[]outside_quotes[] ✓expected_insideexpects_reflectionnot discussed → stops here(no FS check, no score)STAGE 2 · MODELFS checkone call per discussed topicreflection 0–3 + verbatim FS quotesoutcome: reflected · gap, reason stated· gap, reason inferred · gap, no reasonmechanisms · anchor pointscode: verify FS quotes · fix outcomeReviewerconfirms or corrects each output (below)the raw outside text never reaches Stage 2read onceone topic's blockfull FS text
Stage 1 sees only the outside text; Stage 2 sees only the financial statements plus the frozen profile of one topic. Each stage must quote the one text it saw, and code checks every quote.

Method

We score how outside climate disclosures connect to the audited financial statements with a reasoning LLM (gpt-5-mini-2025-08-07, long context window), in two separate stages, so the model never sees both disclosure venues in the same call.

  • Input: all inside and all outside paragraphs flagged as climate-related by DisclHits or DisclBERT.
  • Unit: each firm-year, assessed separately for the same 14 climate topics derived from regulatory guidance (Online Appendix B).

Stage 1 Outside read → topic profile

The model sees only the outside paragraphs. For each of the 14 topics it returns:

  • Discussed outside: whether the topic is substantively discussed, not merely mentioned in passing;
  • Summary and key elements: what the firm discloses, plus up to eight concrete specifics (named assets, amounts, targets, instruments, obligations);
  • Outside quotes: one to four literal quotations supporting the summary;
  • Expected FS reflection: the one reflection a reasonable, financially literate reader would expect in the audited financial statements given the outside disclosure alone, naming the FS area and the kind of item:
    • a recognised amount,
    • a measurement assumption,
    • an estimation-uncertainty or judgement disclosure,
    • a risk or sensitivity disclosure, or
    • an explicit statement of no material effect;
    or thatno reflection would be expected.

The expectation is formed before the model sees the financial statements, so it cannot be tailored to what they contain.

Stage 2 FS check → one outcome

The model is called once per topic discussed outside. It sees all inside paragraphs and that topic's Stage 1 profile (verified outside quotes only, never the raw outside text), and returns:

  • Reflection (0–3): how far the topic is reflected inside — no, limited, moderate or substantial — counting only inside content on this topic's subject matter, with literal inside quotes;
  • Form of reflection: general note, topical note, explicit no-effect conclusion, quantification, or recognised figure;
  • Inside areas: where the topic is reflected, from 14 pre-defined FS areas;
  • Outcome against the Stage 1 expectation, exactly one of reflected · gap, reason stated in the FS · gap, reason inferred · gap without good reason. For the two gap-with-reason outcomes the model names the boundary reason (Section 2) and justifies it.

Quote-grounding

Each stage must quote the one text it saw, and code verifies every quotation against that text (ignoring case, spacing and punctuation). Unverifiable evidence downgrades the result:

  • unverified outside quotes are dropped before Stage 2;
  • a reflection score ≥ 1 without a verified inside quote is treated as 0;
  • reflected whose evidence quote fails is kept when one of its reflection quotes verifies; with none it becomes gap without good reason;
  • reason stated without a verified quote becomes reason inferred.

As before, the prompts make clear that negative answers are valid and common: not discussed, not reflected, nothing expected.

The outcome

Stage 2 picks exactly one outcome against the Stage 1 expectation:

ReflectedThe financial statements deliver the expected reflection. Evidence: a verbatim FS quote, verified by code.
Gap · reason statedNot reflected, and the FS themselves say why (e.g. no material effect, recognition criteria not met, not yet a present obligation). Evidence: the FS sentence stating it, verified by code, plus the matching boundary reason.
Gap · reason inferredNot reflected, the FS are silent about why, but a boundary reason applies on the facts (forward-looking, non-monetary, value chain, immaterial, outside recognition criteria). The model names it and justifies it.
Gap · no good reasonNot reflected and no boundary reason applies: the matter is present, material and FS-relevant, the criteria appear met, yet the FS are silent or dismissive. This is the finding.
Nothing expectedStage 1 expected no FS reflection for this topic, so there is nothing to check.

Evidence is checked in code. "Reflected" without a verifiable quote becomes "gap, no good reason"; "reason stated" without a verifiable quote becomes "reason inferred". Boundary reasons (EFRAG Connectivity DP §2.2–2.3): Aggregated into larger FS items (EFRAG §2.3 a) · Longer time horizon than FS recognition (§2.3 b) · Forward-looking / not yet a present asset or liability (§2.3 c) · Non-monetary metric, not reliably translatable into money (§2.3 d) · Value-chain information beyond the FS boundary, e.g. Scope 3 (§2.3 e) · No FS line item / primary statement counterpart (§2.3 f) · Operational-control basis vs FS consolidation scope (§2.3 g) · Does not meet FS recognition / measurement / disclosure criteria · Below FS materiality (§2.2; IAS 1.7).

Reflection 0–3

Separately from the outcome, Stage 2 scores how far the financial statements reflect the topic at all. Only inside content about this topic counts; adjacent climate content does not.

0 · No meaningfulNo meaningful reflection: the financial statements do not address this topic at all, or only climate content that belongs to a different topic.
1 · LimitedLimited: a generic or boilerplate mention only (e.g. "climate risks are considered"), nothing specific to this topic.
2 · ModerateModerate: the topic is addressed specifically — a judgement, assumption, risk or estimation-uncertainty disclosure names it — but nothing is quantified.
3 · SubstantialSubstantial: specific AND quantified or recognised — an amount, sensitivity, useful life, provision, impairment or explicit no-effect conclusion attributable to this topic.

A reflection ≥ 1 must be backed by verbatim FS quotes, checked in code ignoring spacing and punctuation; if none is found the score is treated as 0. Mechanisms record how the topic connects: mentioned in a general note (significant judgments, accounting policies, estimation uncertainty); named explicitly in the relevant topical fs note (pp&e, impairment, provisions, …); explicit conclusion that the topic has no (material) accounting effect; the topic’s financial effects are quantified (amount, percentage, sensitivity); a recognised fs amount is attributed to the topic. Anchor points record which FS areas it appears in.

What the reviewer judges

Open a cell. The drawer shows the passages on the left, with every verified quote highlighted in place, and the two stages on the right. Five judgments per topic:

  1. Discussed outside. Is the topic substantively discussed in the outside passages? Confirm the model's yes/no or mark it wrong.
  2. Expectation. Is the expected reflection a reasonable thing to look for in the financial statements? Yes or no.
  3. Reflection. Set your own 0–3 for how the financial statements reflect the topic. Use the highlighted FS quotes.
  4. Outcome. Pick your own outcome from the four. The model's pick is marked; the evidence quote, the boundary reason and its justification are shown.
  5. Mechanisms and anchor points. Confirm the model's set as shown, or toggle chips to correct it.

Anything left untouched is recorded as unrated. Ratings are append-only and attributed to your account; saving again adds a new entry.

The grid

Rows are firm-years, columns the 14 outside climate topics. A cell shows the reflection score for topics discussed outside, or "–" when the topic was not discussed (no Stage 2 call). A red bar under a cell marks a gap without good reason.