Strix is a structured ESG-metric engine. It extracts the actual figures companies disclose — with full provenance — normalizes them, records the reporting basis behind each one, and tracks how they change over time.
Pull every disclosed ESG metric from factbooks, data tables and reports — each tied to its source document, sheet, cell and page.
Convert every figure to a canonical value and unit, so numbers compare cleanly across companies, years and formats.
Capture the reporting basis behind each number: consolidation boundary, GWP set, assurance, managed-JV and contractor inclusion, reporting standard.
Follow targets and commitments year to year — and flag the ones that quietly disappear from later reports.
Surface restatements and disclosure events where a previously reported figure silently changes.
Each observation is pinned to the exact place it came from and converted to a comparable unit — so a figure can always be audited back to source.
Two companies can report “emissions” on entirely different bases. Strix records the boundary, standard and assurance so figures are compared like for like.
Because Strix holds every figure across years, it catches what a single read-through never could — figures silently restated, and targets that quietly vanish from later reports.
Kestrel is the platform’s qualitative engine. It reads the narrative side of mining reporting — the sustainability, social-performance and governance prose — and turns thousands of pages of it into structured, comparable data: every passage scored against four explicit Creating Shared Value criteria, separating genuine, evidenced community and environmental benefit from PR. It is deliberately conservative — designed to err toward false negatives, never false positives — so only passages where real shared value is simultaneously and explicitly evidenced qualify as positive.
Read the PDF and segment it into discrete passages of 100–220 words — one passage per request, so each judgement is independent.
A language model extracts evidence from the passage — the four CSV definitional elements, the linkage claimed, and the supporting text. It does not assign the label.
Deterministic, human-authored rules turn that evidence into a label. Positive only when all four elements are simultaneously and explicitly present; any ambiguity resolves downward.
Passage labels roll up into company and sector profiles: share of true shared value, initiative types, CSV pillars and evidence strength.
A two-layer instrument. The model extracts evidence; human-authored rules — versioned and applied deterministically — turn that evidence into a label. No model is ever asked to decide what qualifies as shared value.
Kestrel’s headline call on every passage — Positive: genuine, evidenced shared value; Ambiguous: shared value claimed but not clearly evidenced; Negative: no measurable shared value in the passage.
How far an initiative actually goes — Transactional (one-off giving or compliance), Transitional (improving existing operations), Transformational (reshaping the wider system). “No initiative” means the passage describes none.
Which lever of shared value the activity pulls (the three ways business and society gain together) — Products & markets (serving unmet needs), Value-chain productivity (greener, fairer operations), Local cluster development (strengthening the surrounding economy).
How clearly the passage ties the activity to a real outcome — Explicit (stated and evidenced), Implied (suggested but not shown), None (no link made).
How hard the supporting evidence is, graded Tier 1–4 — Tier 4 is independently verified data, Tier 1 a bare assertion. “None” means there is nothing to check.
Inter-run agreement was measured with Fleiss’ κ across 5 independent classification runs per passage. Overall agreement is substantial; on the headline question — whether shared value is present at all — it is almost perfect.
Initiative type is the least stable dimension: the boundary between Transactional and Transitional is the main source of disagreement, and those definitions are the top candidates for refinement.
of passages received the identical label across all 5 runs (5/5).
received the same label in at least a majority of runs — a clear consensus for nearly every passage.
On a held-out sample, V2.0 matched expert human labels on the headline is this shared value? question 85% of the time. Exact three-category agreement is lower at 38% — and that gap is itself informative.
Humans are becoming the limiting factor. Raters routinely infer beyond what a report actually discloses and lean generous — awarding shared value where the evidence on the page doesn’t support it.
V3.0 targets >90% — with a refined rubric around Transactional vs. Transitional and a larger, multi-coder adjudicated sample.
How Kestrel reads the narrative: the CSV classification schema, the conservative default, the text-only evidence rule, the two-layer decision architecture (the model extracts evidence; rules decide), and the validation approach — every design decision stated and defended.
How Strix reads the numbers: the four-layer data model, provenance to the source cell, additive unit normalisation, basis-aware comparability, and how restatements and quietly dropped targets are surfaced.