Split-Perception: When Large Language Models Disagree About the Same Company

Split-Perception: The Unpriced Information Risk in AI-Assisted Diligence | VeritasLinks
VeritasLinks

Research → risk signal

Split-Perception: When Large Language Models Disagree About the Same Company

The same company. Identical questions. Materially different verdicts. The variation is currently unpriced.

Large language models are increasingly the first source consulted when an investor, analyst, credit officer, or KYB reviewer evaluates a company. The model returns a compression of everything it has absorbed — site, reviews, press, forums, structured data — synthesized into a short verbal judgment delivered at the moment first impressions form.

Yet different models, asked identical questions about the same firm on the same day, routinely return assessments that differ not merely in tone but in substance. One model may describe a defensible platform; another may treat it as a commodity tool; one may name the company among category leaders; another may omit it entirely. We call this split-perception.

Where traditional search returned a ranked list a reader could inspect, an AI answer is a single synthesized verdict. When verdicts diverge across models, the “fact of the matter” about a company’s standing becomes model-dependent. That dependence is currently invisible in most deal files, credit memos, and KYB records.

Live report surface

Perception affect field · split detected

Northline Analytics · illustrative fixture on the real UI. Not a screenshot.

Interactive

Perception affect field

Models split into two emotional camps

Polarization

64%

Division is material, not noise.

Dominant zone

Activated

The scale of the divergence

Measured magnitude, not anecdote

Drawing on a dataset of AI-generated assessments covering more than 5,000 companies and several million individual model responses, with quantitative claims resting on the fully scored current-vintage population, the following results hold:

  • In the fully instrumented subset, the median gap between the highest- and lowest-scoring models on a single measurement date is 100 points on the 300–870 VeritasScore scale.
  • The 90th percentile reaches 134 points.
  • Reporting bands on the scale are 200 points wide; a typical fully instrumented company therefore shows a gap of roughly half a reporting band between its most and least favorable model.
  • Among companies with category-query coverage from at least two models, 6.4% are simultaneously omitted entirely by one model and ranked in the top three by another. For these firms the split is not a matter of degree: the same company is a category leader in one AI’s answer and absent from another’s.

These figures are the first instrumented estimates rather than settled population statistics; per-model persistence now covers all new assessments and the estimates will tighten. Even at current sample sizes the magnitude is large enough to matter for any workflow that treats a single model’s output as a stable reading.

100 pts

Median cross-model gap on the 300–870 scale (fully instrumented subset).

134 pts

90th percentile gap. Structured disagreement, not one noisy prompt.

200+ pts

Retrieval-grounded model swing on one firm across repeated runs.

Two distinct regimes

Stable disagreement versus retrieval-imported volatility

Repeated runs on the same firm under a fixed methodology vintage separate the five production judge models into two regimes.

Stable disagreement

Four of the five models are individually stable. Their within-model score ranges across runs span 44–79 points, with standard deviations between roughly 15 and 24 — materially below the 100-point median between-model range. For these models, split-perception behaves as a stable property of the model–evidence pair: the same model keeps reaching approximately the same verdict, and the verdicts differ between models. Divergence is measurable precisely because the individual judges are consistent.

Retrieval volatility

The fifth model — retrieval-grounded — exhibits a different regime. Within-model range has been observed at 232 points with a standard deviation above 76, including runs scored at the scale floor of 300 and materially different scores from separate runs on the same day. A model that re-fetches live evidence on every run inherits the volatility of whatever it happens to retrieve. A firm with a thin citation-authority profile offers that retrieval little to stabilize on. Under this reading the volatility is not primarily a defect of the model; it is a measurement of the firm’s evidence base.

The two regimes carry distinct risks. Stable splits mean the verdict on a company depends on which model is consulted. Retrieval volatility means it also depends on when. An evaluator relying on a single retrieval-grounded query observes one draw from a wide distribution.

Live report surface

Perception affect field · split detected

Northline Analytics · illustrative fixture on the real UI. Not a screenshot.

Interactive

Perception affect field

Models split into two emotional camps

Polarization

64%

Division is material, not noise.

Dominant zone

Activated

Where divergence concentrates

Authority signals, not simple visibility

Models agree far more readily on visibility signals — whether a company is mentioned and with what sentiment — than on positioning signals: category assignment, leadership attribution, and defensibility. Divergence concentrates precisely where authority signals are weak or conflicting. A firm can show high mention rate and positive sentiment while still producing large composite splits because models disagree on what the company is and where it stands. Those are the dimensions a diligence, credit, or KYB process cares about most.

Why this constitutes an unpriced information risk

Asymmetric downside and no audit trail

The practical stakes are asymmetric. A company assessed favorably by the model its evaluator happens to use experiences no friction. The same company assessed unfavorably by a different model loses opportunities it never learns existed. The variation leaves no trail: the screening memo, credit note, or KYB file records conclusions, not the model that produced them or the spread across models.

For risk and compliance workflows that increasingly incorporate AI-assisted research — M&A screening, credit assessment, KYB review — the conclusion of a first-pass step can depend on an arbitrary tooling choice. That dependence is an input risk that is currently unpriced and usually unrecorded.

How measurement changes the situation

From invisible draw to explicit distribution

VeritasLinks executes a large structured battery (1,000–1,500 prompts) under a fixed orchestration layer, persists per-model scores, and reports both the composite VeritasScore and the explicit range. Additional outputs — affect field, polarization index, coded objections, component decomposition, citation profile — make the nature of the disagreement visible, not only its size.

The same measurement apparatus supports brand-scope assessment and narrative-scope assessment (pitch decks, business plans, sale materials). The separation allows deal and credit teams to see whether models disagree about the company itself, about the story the company tells, or both.

Procedural mitigations that require no new software — querying several models, recording the spread, treating wide splits as findings, noting model and version — become far more powerful once the distribution is measured rather than guessed.

Live report surface

VeritasScore · multi-model distribution

Northline Analytics · illustrative fixture on the real UI. Not a screenshot.

Interactive

Models

Score by LLM model

620Overall · 300–870

Research paper

Full methodology and findings

The definitions, measurement design, dataset description, divergence statistics, regime analysis, and longitudinal case study are presented in the working paper “Split-Perception: Measuring Divergence in How Large Language Models Assess Company Credibility” (Andrew Pomazkov, VeritasLinks, July 2026). The paper is the authoritative source for the quantitative claims used across VeritasLinks enterprise materials.

Read the paper →·PDF →

Measure the split before it becomes a silent decision input

Free analysis

Run it on any company

Public URL. Same funnel as the homepage. No card required.

Enter a public website

Multi-model · VeritasScore · ~5 minutes

See a sample dossier →

FAQ

Short clarifications — positioning and proof live in the sections above.

Is split-perception just ordinary model randomness?+

No. At the operating battery size, residual within-model variance is treated as the noise floor. Between-model divergence is judged against that floor. Four of five models show individual stability; the divergence between them is structured.

Does a wide split mean the company is low quality?+

Not necessarily. It usually means the public evidence base is thin or contested. That fact is itself diligence-relevant.

How is this different from standard GEO or AI-visibility tools?+

Standard GEO tools treat model output as a single surface to be optimized for mentions and sentiment. Split-perception measurement treats the models as a set of potentially disagreeing judges and quantifies the disagreement, its stability, and its concentration in positioning and authority signals.

Which workflows are most exposed?+

Any workflow in which AI-assisted research already shapes first-pass attention or risk notes — M&A sourcing and screening, credit assessment, KYB/KYC review, vendor selection, and investment screening.

Can the split change over time?+

Yes. New third-party content, changes in model behavior, or shifts in the competitive set can move both the level and the spread. The measurement system is designed so that absence of new signals produces relative stability while real signals appear in subsequent runs.