How it works
Independent analysis · 39 mining majors · 2010–2026
Profiles
Side by side
01Methods · Strix

Where Kestrel reads the words, Strix reads the numbers.

Strix is a structured ESG-metric engine. It extracts the actual figures companies disclose — with full provenance — normalizes them, records the reporting basis behind each one, and tracks how they change over time.

0
ESG observations
0
metric families
0
restatements caught
0
targets quietly dropped
56 targets tracked153 disclosure events6 breakdown dimensions9 companies (deep coverage)
01
Extract

Pull every disclosed ESG metric from factbooks, data tables and reports — each tied to its source document, sheet, cell and page.

02
Normalize

Convert every figure to a canonical value and unit, so numbers compare cleanly across companies, years and formats.

03
Contextualize

Capture the reporting basis behind each number: consolidation boundary, GWP set, assurance, managed-JV and contractor inclusion, reporting standard.

04
Track

Follow targets and commitments year to year — and flag the ones that quietly disappear from later reports.

05
Detect

Surface restatements and disclosure events where a previously reported figure silently changes.

Every number, traceable

Each observation is pinned to the exact place it came from and converted to a comparable unit — so a figure can always be audited back to source.

Total Energy Consumption
Agnico Eagle · FY2018
2,018,129.274 GJ2,018.129 TJ
SourceAgnico-Eagle-2018-GRI-Table
Sheet · cellEnv - Energy · D23
Basis · confidenceOperational Control · 95%
The pipeline, one pass
strix · pipeline
$strix extract reports/*.pdf
✔ scanned 37 companies
✔ 25,107 observations · document · sheet · cell
$strix normalize --units canonical
✔ 58 metrics · 13 families → canonical units
$strix context --boundary --assurance --standard
✔ reporting basis captured (boundary · GWP · assurance)
$strix track --targets --restatements
✔ 1,051 targets · 542 quietly dropped
⚠ 1,192 restatements detected
$strix report --public
✔ 153 disclosure events logged
✔ strix_public.json ready
$
The basis behind the number

Two companies can report “emissions” on entirely different bases. Strix records the boundary, standard and assurance so figures are compared like for like.

Consolidation boundary
Operational Control544
Equity Share41
Financial Control40
Mixed32
Reporting standards captured
GRIWAF MCASASBGHG PROTOCOL CORPORATEGHG PROTOCOL SCOPE3TCFDICMM MINING PRINCIPLESUNGCICMMGHG PROTOCOL SCOPE2AA1000AS V3 2020WGC RGMPESRSMSHAISO 14064 1ISSBESRS E2GRI 14 MINING 2024 GRIRGMPIFRS S1
What the numbers hid

Because Strix holds every figure across years, it catches what a single read-through never could — figures silently restated, and targets that quietly vanish from later reports.

Anglo American · Total Energy Consumption 2021
83.7 → 62.8 million GJ
Anglo American · Total Energy Consumption 2022
83.3 → 64.4 million GJ
Anglo American · Total Energy Consumption 2023
89 → 68.4 million GJ
Anglo American · Deliver Net Positive Impact (NPI) for biodiversity across the organisation by 2030 (and where necessary to 2040 and beyond), using a 2018 baseline
dropped 2025
02Methods · Kestrel

How Kestrel reads the words — AI-extracted evidence, rules-based judgement.

Kestrel is the platform’s qualitative engine. It reads the narrative side of mining reporting — the sustainability, social-performance and governance prose — and turns thousands of pages of it into structured, comparable data: every passage scored against four explicit Creating Shared Value criteria, separating genuine, evidenced community and environmental benefit from PR. It is deliberately conservative — designed to err toward false negatives, never false positives — so only passages where real shared value is simultaneously and explicitly evidenced qualify as positive.

0
passages classified
0
mining majors
0
reports · 2018–2026
0.00%
rated true shared value
The pipelineone passage in · a model reads the evidence · rules assign the label
01PDF → passages
Ingest

Read the PDF and segment it into discrete passages of 100–220 words — one passage per request, so each judgement is independent.

02model
Extract

A language model extracts evidence from the passage — the four CSV definitional elements, the linkage claimed, and the supporting text. It does not assign the label.

03rules
Adjudicate

Deterministic, human-authored rules turn that evidence into a label. Positive only when all four elements are simultaneously and explicitly present; any ambiguity resolves downward.

04profiles
Aggregate

Passage labels roll up into company and sector profiles: share of true shared value, initiative types, CSV pillars and evidence strength.

A two-layer instrument. The model extracts evidence; human-authored rules — versioned and applied deterministically — turn that evidence into a label. No model is ever asked to decide what qualifies as shared value.

What Kestrel records on every passage

Five structured judgements per passage — here is how all 105,227 fall.

Primary label
0.62% positive · 4.20% ambiguous · 95.18% negative

Kestrel’s headline call on every passage — Positive: genuine, evidenced shared value; Ambiguous: shared value claimed but not clearly evidenced; Negative: no measurable shared value in the passage.

Initiative type

How far an initiative actually goes — Transactional (one-off giving or compliance), Transitional (improving existing operations), Transformational (reshaping the wider system). “No initiative” means the passage describes none.

Transformational0.06% · 61
Transitional11.3% · 11,859
Transactional15.2% · 16,023
No initiative73.4% · 77,284
CSV pillar

Which lever of shared value the activity pulls (the three ways business and society gain together) — Products & markets (serving unmet needs), Value-chain productivity (greener, fairer operations), Local cluster development (strengthening the surrounding economy).

Products & markets0.52% · 545
Value-chain productivity14.0% · 14,720
Local cluster development5.9% · 6,165
None79.6% · 83,797
Linkage strength

How clearly the passage ties the activity to a real outcome — Explicit (stated and evidenced), Implied (suggested but not shown), None (no link made).

Explicit1.9% · 2,012
Implied5.0% · 5,254
None93.1% · 97,961
Evidence strength

How hard the supporting evidence is, graded Tier 1–4 — Tier 4 is independently verified data, Tier 1 a bare assertion. “None” means there is nothing to check.

Tier 4 — strongest0.09% · 99
Tier 34.3% · 4,561
Tier 223.3% · 24,508
Tier 13.2% · 3,397
None69.1% · 72,662
Reliability · Kestrel V2.0 · GPT-5

We ran every passage 5 times. The labels hold up.

Inter-run agreement was measured with Fleiss’ κ across 5 independent classification runs per passage. Overall agreement is substantial; on the headline question — whether shared value is present at all — it is almost perfect.

Initiative type is the least stable dimension: the boundary between Transactional and Transitional is the main source of disagreement, and those definitions are the top candidates for refinement.

What “κ” means: Fleiss’ kappa scores how often the runs agree beyond what you’d expect by chance — 0 is no better than guessing, 1.0 is identical every time. Below 0.6 is moderate, 0.6–0.8 substantial, above 0.8 almost perfect.
Primary label — is shared value present at all?
almost perfect
0.822
κ
Overall — all diagnostic dimensions jointly
substantial
0.765
κ
Initiative type — transactional / transitional / transformational
substantial
0.702
κ
0.60
Almost perfect →
1.00
Full agreement
77.2%

of passages received the identical label across all 5 runs (5/5).

Majority agreement
99.4%

received the same label in at least a majority of runs — a clear consensus for nearly every passage.

Human agreement

Kestrel agrees with expert coders 85% of the time.

On a held-out sample, V2.0 matched expert human labels on the headline is this shared value? question 85% of the time. Exact three-category agreement is lower at 38% — and that gap is itself informative.

Humans are becoming the limiting factor. Raters routinely infer beyond what a report actually discloses and lean generous — awarding shared value where the evidence on the page doesn’t support it.

V3.0 targets >90% — with a refined rubric around Transactional vs. Transitional and a larger, multi-coder adjudicated sample.

Kestrel V2.0 · today
baseline
85%
agree
Kestrel V3.0 · target
targeted
>90%
agree
The full method
Kestrel · v4.0

A Conservative, Rule-Governed Classifier for Creating Shared Value in Corporate ESG Reporting

How Kestrel reads the narrative: the CSV classification schema, the conservative default, the text-only evidence rule, the two-layer decision architecture (the model extracts evidence; rules decide), and the validation approach — every design decision stated and defended.

Strix · v1.2

The Evidence Layer Behind Kestrel — Structured, Audit-Traceable ESG Extraction

How Strix reads the numbers: the four-layer data model, provenance to the source cell, additive unit normalisation, basis-aware comparability, and how restatements and quietly dropped targets are surfaced.

Methodology papers are free to download · the demo covers the Kestrel platform & data
Framework: Porter, M. E. & Kramer, M. R. (2011). “Creating Shared Value.” Harvard Business Review, 89(1), 62–77.
Reliability measured with Fleiss’ κ across 5 independent runs per passage. Fleiss, J. L. (1971). “Measuring nominal scale agreement among many raters.” Psychological Bulletin, 76(5), 378–382. doi:10.1037/h0031619.