Skip to content

How well can AI answer company questions—and prove it?

A planned quarterly benchmark of reliability in resolving legal entities, returning official-source company facts and preserving the provenance needed to review an answer.

First edition in preparation. No model results, rankings or publication date are announced on this preview.
Grounding auditMETHODOLOGY PREVIEW
OBJECT OF EVALUATIONCompany answer + evidenceNo results published
DimensionEvaluation question
ENTITYEntity

Was the intended legal entity resolved?

FACTFact

Does the answer match the reference evidence?

SOURCESource

Can the company fact be traced and reviewed?

TIMETime

Is the answer aligned to the relevant date?

COVERAGECoverage

Does the answer respect what the source discloses?

UNCERTAINTYUncertainty

Does the system abstain when evidence is insufficient?

Reference standardOfficial-source evidence, with declared limitations
100
Methodology points
6
Grounding dimensions
Quarterly
Planned benchmark cycle
In preparation
First-edition status

A correct-looking answer is not enough.

Company-data grounding connects a statement to the right legal entity, the supporting evidence and the relevant point in time—while making uncertainty visible.

01

Resolve the entity

Identify the legal company using jurisdiction, registry identity and available official identifiers—not the name alone.

02

Anchor the fact

Compare the answer with the applicable official registry record or filed document used as the reference.

03

Preserve the evidence

Carry source, identifier, filing and temporal context far enough for a reviewer to inspect the answer.

04

Express uncertainty

Distinguish unavailable, undisclosed, unresolved and genuinely negative evidence—and abstain when needed.

Six dimensions of a defensible company answer.

The weights describe the proposed evaluation structure. They are methodology components—not published performance results.

E20F25P20T15C10U10
E20 points

Entity resolution

Whether the response identifies the intended legal entity and avoids conflating namesakes, branches or related companies.

EVALUATION SIGNALS
  • Jurisdiction match
  • Identifier match
  • Ambiguity handling
F25 points

Factual accuracy

Whether supported company facts agree with the official-source reference used for the benchmark case.

EVALUATION SIGNALS
  • Value agreement
  • Relationship direction
  • Claim support
P20 points

Provenance quality

Whether the answer gives enough registry, filing and reference context to trace the returned fact.

EVALUATION SIGNALS
  • Named source
  • Record or filing context
  • Reviewable evidence
T15 points

Freshness & temporal integrity

Whether the response uses the relevant period or date and avoids presenting historic evidence as current.

EVALUATION SIGNALS
  • Reference date
  • Filing period
  • Time-qualified claim
C10 points

Availability-aware completeness

Whether the response covers supported facts without inventing fields that the source or jurisdiction does not disclose.

EVALUATION SIGNALS
  • Source-aware scope
  • Missing-data handling
  • No fabricated completeness
U10 points

Uncertainty & abstention

Whether the system qualifies low-confidence answers, requests clarification or declines to assert an unsupported conclusion.

EVALUATION SIGNALS
  • Calibrated language
  • Clarifying behaviour
  • Appropriate abstention

Evaluation should expose its own chain of custody.

The benchmark process will connect each company question to its reference evidence, evaluated system configuration and dimension-level assessment.

Official-source referenceAn auditable basis for comparison—not a claim of absolute truth.
  1. 01
    Define company questions

    Create task patterns around identity, officers, ownership, financials, status and corporate relationships without disclosing an edition sample before publication.

  2. 02
    Build reference evidence

    Assemble the relevant official registry records and filings, including jurisdiction and temporal context for each case.

  3. 03
    Run the evaluated system

    Record the answer and evidence produced under the declared model, tool, retrieval and prompt configuration.

  4. 04
    Evaluate each dimension

    Apply the published criteria to entity, fact, source, time, coverage and uncertainty behaviour.

  5. 05
    Report with limitations

    Publish the configuration, scope, exclusions, failure patterns and time boundary alongside any future score.

The Index is designed to show how answers fail.

A total score can conceal consequential behaviour. Edition reporting will need to separate identity, evidence, timing and uncertainty failures.

FAILURE MODE

Wrong legal entity

A plausible company name is returned, but it belongs to another jurisdiction, identifier or legal entity.

FAILURE MODE

Unsupported fact

The answer states a company attribute or relationship that the reference evidence does not support.

FAILURE MODE

Opaque provenance

The answer appears correct but does not expose enough source context for independent review.

FAILURE MODE

Temporal collapse

Historic, filed, retrieved and current dates are treated as though they describe the same moment.

FAILURE MODE

Availability hallucination

The system invents a missing field instead of recognising that the jurisdiction or source may not disclose it.

FAILURE MODE

False certainty

The answer makes an unqualified claim despite ambiguous identity, incomplete evidence or conflicting records.

Source conditions belong beside the score—not inside it.

A system should not gain or lose model-quality points simply because one jurisdiction publishes more company information than another. Each future edition will need a separate profile of the source environment used in evaluation.

Explore registry coverage
Disclosure regime
Which directors, owners, financials and status fields the authority publishes
Source access
Public interface, document route, official paywall and retrieval constraints
Field semantics
Local legal forms, status labels, filing concepts and identifier meaning
Temporal context
Publication delays, effective dates, filing periods and source update behaviour
Language & documents
Local-language labels, scans, structured filings and document quality
Data sovereignty
Jurisdictional, contractual and deployment considerations relevant to evaluation

Results will need context, not just a leaderboard.

The first edition is still in preparation. When editions are released, these materials are intended to make the findings interpretable and reviewable.

01

Methodology paper

The published methodology will define scoring rules, edge cases and edition-specific adjustments.

PLANNED OUTPUT
02

System scorecards

Future scorecards will describe the evaluated configuration and performance by grounding dimension.

PLANNED OUTPUT
03

Failure-mode analysis

Each edition will examine how and where company-data answers fail, not only a total score.

PLANNED OUTPUT
04

Jurisdiction profiles

Separate profiles will explain source and disclosure conditions without turning them into model-quality points.

PLANNED OUTPUT
05

Reproducibility notes

Edition materials will state the relevant setup, timing, scope and limitations needed to interpret results.

PLANNED OUTPUT

Read the benchmark for what it is—and what it is not.

The Index is intended to make company-data grounding more measurable. It cannot remove limitations in official sources, evaluation design or rapidly changing AI systems.

  • 01

    The benchmark will be prepared by Zephira. It is not an independent certification or regulatory assessment.

  • 02

    An official-source reference is an auditable benchmark basis—not a guarantee that the authority’s record is complete or error-free.

  • 03

    Results will be time-bound because registries, models, retrieval systems and product configurations change.

  • 04

    Performance will apply to the declared model, tools, prompts, retrieval settings and evaluation scope—not every possible deployment.

  • 05

    Source availability, filing delay and jurisdictional disclosure can limit what any system can answer responsibly.

  • 06

    Each edition will need to publish material exclusions, methodology changes and known limitations beside its results.

About the methodology preview

Explore MCP for AI agents

What is the Company Data Grounding Index?

the Company Data Grounding Index is part of Zephira's registry-sourced company-data platform. The methodology evaluates entity resolution, factual accuracy, provenance, time, coverage awareness and uncertainty in AI company-data answers.

Is the Index a certification?

No. It is a benchmark prepared by Zephira, not a certification, regulatory approval or independent assurance opinion.

Has the first edition been published?

No. The first edition is in preparation. No system results, rankings, sample sizes, participating systems or publication date are being announced on this methodology preview.

Why use official company sources as references?

Official registries and filed documents provide an identifiable, reviewable basis for legal-entity facts. They are treated as official-source references, not as an absolute guarantee that every source record is current, complete or error-free.

What does the 100-point methodology measure?

It allocates 20 points to entity resolution, 25 to factual accuracy, 20 to provenance quality, 15 to freshness and temporal integrity, 10 to availability-aware completeness, and 10 to uncertainty and abstention.

Are jurisdiction conditions part of the score?

No. Source access, disclosure regimes, field semantics and data-sovereignty considerations will be described in a separate jurisdiction profile so model quality is not confused with source availability.

Will the benchmark compare AI models or complete systems?

Future editions will identify the evaluated configuration, which may include a model, prompts, retrieval layer, tools and source access. Results should be interpreted for that declared system configuration rather than the model name alone.

Who uses the company data grounding index?

Compliance, risk, data, product and AI teams use the company data grounding index when they need structured company facts with clear provenance.

Where does the data for the company data grounding index come from?

Official company registries and filed company documents are the starting point. Source context is retained so users can trace material facts.

Which countries are supported for the company data grounding index?

Coverage is jurisdiction-specific. Review the registry directory and field-availability pages for the documented source, fields and limitations in each country.

Build agent answers on identifiable company evidence.

The methodology is in preview. Speak with Zephira about the data and provenance layer behind grounded company workflows.

Discuss agent grounding Explore the solution