AI Perception Signals for Credit Risk and Underwriting

AI Perception Signals for Credit Risk & Underwriting | VeritasLinks
VeritasLinks

Credit risk · underwriting

AI Perception Signals for Credit Risk and Underwriting

Lend on stated data and history while models already form a market view of the same borrower.

Credit decisions rest on financial history, cash-flow analysis, collateral, covenants, and information provided by the borrower. Those sources remain primary and non-negotiable.

A second layer is already active inside many institutions. Analysts and internal tools consult large language models for first-pass framing of a name, competitive context, and risk language. The models synthesize public footprint, press, reviews, and third-party discussion into short verbal judgments. Different models reach different judgments. The divergence is rarely measured and almost never retained in the credit file.

VeritasLinks quantifies that layer so it can be inspected, decomposed, and placed beside traditional credit data rather than remaining an unrecorded influence on how a name is first understood.

620
Held back

Same 300–870 gauge used in every combined report.

Why perception is a forward-looking signal

Historical files describe the past. Model perception reflects the current synthesis.

Financial statements and historical performance describe what has already occurred. Large language models compress the current public evidence base — including how the market discusses the firm, which sources they treat as authoritative, and what risk language they attach — into a judgment that is available instantly.

A borrower can present clean historical numbers and still carry perception characteristics that traditional files do not contain:

  • Weak or conflicting authority signals (low citation quality, thin third-party discussion, absence from high-authority sources).
  • High cross-model disagreement on market position, defensibility, or category membership.
  • Recurring objections from evaluative panels around unit economics, platform dependency, competitive moat, or growth sustainability.
  • Material omission from the category and “best tools / best platforms” queries that customers, partners, and competitors already use.
  • Elevated polarization: models clustering into emotionally different camps about the same firm.

None of these replace cash-flow analysis or collateral valuation. They function as early indicators of how the market narrative is forming and how model-mediated attention is currently distributed. In workflows where AI-assisted research already occurs, those indicators are already influencing framing; the only question is whether they are visible.

Visibility

90

Brand Visibility component (max 150). Decomposable, monotonic.

Position

72

Market Position. Omission on category queries pulls this down.

Preference

22

Brand Preference from Focus Groups (max 40).

Cross-model gap

130 pts

ChatGPT 680 vs Claude/Grok 550. Near the paper’s 90th percentile (134).

Anatomy of the signal for credit teams

Decomposable components and explicit risk deduction

The VeritasScore sits on a 300–870 scale deliberately constructed with the same design principles as consumer credit scores: bounded, monotonic, and decomposable. Every report discloses the contribution of each component so that a credit officer can see what is carrying the score and what is subtracting from it.

Core components include:

  • Brand Visibility — whether and how prominently the firm is named across query families.
  • Market Position — category assignment and leadership attribution.
  • Customer Perception — sentiment and panel preference outcomes.
  • Content Presence — the footprint of answer-shaped content available for models to draw on.
  • Growth Potential — forward-looking language models attach to the firm.
  • Risk Deduction — an explicit subtractive term for objections, negative attributions, and instability.

A high visibility score paired with weak market position and a material risk deduction tells a different story from a balanced profile at the same overall number. The decomposition makes that difference legible.

In addition, the assessment produces:

  • Per-model score distribution (the visible split).
  • Valence–arousal affect field and polarization index.
  • Coded objections from synthetic buyer panels with retained transcripts.
  • Citation and authority profile.

When a business plan or credit application is supplied, the narrative-review module conditions models on a credit-officer role and scores the narrative against a fixed rubric. The reservations that surface are, in operational experience, frequently the same reservations human reviewers later raise. The narrative score is reported separately from brand-scope measurement so that the strength of the story can be distinguished from the strength of the public footprint.

Live report surface

VeritasScore · multi-model distribution

Northline Analytics · illustrative fixture on the real UI. Not a screenshot.

Interactive

Models

Score by LLM model

620Overall · 300–870

Stability, volatility, and monitoring implications

Two regimes of model behavior and what they mean for portfolio review

Longitudinal measurement under a fixed methodology vintage separates the model set into two regimes.

Four of the five production models used as judges are individually stable: their within-model score ranges across repeated runs remain materially below the typical between-model range. For these models, split-perception behaves as a stable property of the model–evidence pair. The same model keeps reaching approximately the same verdict; the verdicts differ between models. The practical implication is that the AI-assisted view of the borrower depends on which model is consulted.

The fifth model — retrieval-grounded — exhibits a different regime. Within-model ranges exceeding 200 points have been observed on the same firm, including runs at the scale floor and materially different scores from separate runs on the same day. A model that re-fetches live evidence inherits the volatility of whatever it retrieves. Firms with thin citation profiles offer little for that retrieval to stabilize on. Under this reading, the volatility is not primarily a defect of the model; it is a measurement of evidence thinness.

For credit monitoring both regimes matter. Stable splits mean that any single-model research note is tooling-dependent. High retrieval volatility is itself a flag that the public evidence base is thin. Periodic re-assessment of portfolio names creates a light time series of perception state that can be reviewed alongside financial monitoring. Material movements, especially when accompanied by rising polarization or new high-severity objections, warrant attention even if financial metrics remain stable.

Placement in the credit process

Where the layer adds value without displacing existing controls

The perception layer is most useful at three points:

Initial framing and screening

When a name first enters the pipeline and analysts or tools consult models for context. Recording the distribution and key objections prevents the first impression from remaining an invisible single-model draw.

Underwriting support

When the full assessment (including narrative review of the business plan) is placed beside financial analysis. Component decomposition and risk language become part of the credit discussion.

Ongoing portfolio monitoring

Periodic re-runs detect shifts in model perception that may precede or accompany changes in market standing.

In all three cases the traditional credit file remains primary. The perception layer is supplementary evidence that makes an already-active influence measurable.

Research foundation and key statistics

Evidence base

Quantitative claims rest on the methodology and dataset described in the working paper “Split-Perception: Measuring Divergence in How Large Language Models Assess Company Credibility.” Key measured results relevant to credit use include the median 100-point cross-model range, the 90th-percentile range of 134 points, the separation of stable versus retrieval-volatile regimes, and the concentration of divergence where authority signals are weak. The paper explicitly lists credit assessment among the workflows for which split-perception constitutes an unpriced information risk.

Read the paper →

Add a measurable perception layer to credit assessment

If AI-assisted research is already part of how names are framed inside the institution, the perception state at the time of underwriting or monitoring can be measured, decomposed, and retained.

Free analysis

Run it on any company

Public URL. Same funnel as the homepage. No card required.

Enter a public website

Multi-model · VeritasScore · ~5 minutes

See a sample dossier →

FAQ

Short clarifications — positioning and proof live in the sections above.

Is the VeritasScore a credit score or a substitute for traditional credit data?+

No. It measures AI-mediated reputation, positioning, and risk language. It is designed to sit beside financial history, cash-flow analysis, and collateral assessment as an independent layer.

How should a wide split be interpreted during underwriting?+

As a finding that the AI-assisted view of the company is highly model-dependent. It usually indicates a thin or contested public evidence base and supports additional human review of narrative strength and market positioning.

Can this be applied to existing portfolio companies?+

Yes. Periodic re-assessment produces a time series of perception state that can be reviewed alongside conventional financial monitoring.

What does high polarization indicate?+

Models are clustering into emotionally different camps about the same borrower. Simple averages hide the divergence. The polarization index makes the division visible.

How does narrative review differ from brand-scope assessment?+

Brand-scope measurement captures how models describe the firm across open queries. Narrative review places models in an explicit credit-officer or investment-reviewer role and scores a supplied business plan or deck against a fixed rubric. The two scores can diverge on the same date; the gap is informative.

Is the methodology stable enough for repeated use?+

The orchestration layer is fixed and the methodology is versioned. Within the same vintage, residual within-model variance at the operating battery size is treated as the noise floor. Between-model divergence is judged against that floor. Four of five judge models show individual stability across repeated runs.