Faro Index
Faro Index
Faro Index Research · Annual Benchmark

The State of AI Brand Intelligence

How ChatGPT, Perplexity, Gemini, and Claude describe the companies buyers ask them about. Measured across four signals, in their own words.

503
Companies
21
Industries
4
AI platforms
90.1%
Mean accuracy
July 2026 · Built with the Faro Method · faroindex.ai
Inside this report

Contents

  1. Executive summary
  2. The benchmark in numbers
  3. How scores are distributed
  4. Methodology, the four pillars & Six Signals
  5. Beyond accuracy: the other three signals
  6. How each platform fails
  7. What AI actually says: a field guide
  8. Industry rankings
  9. Failure mode 1 · Identity divergence
  10. Failure mode 2 · Erasure & fabrication
  11. Failure mode 3 · The homonym trap
  12. Recommendations by failure mode
  13. About Faro Index & pricing
  14. Glossary & method notes
01 · Executive summary

When AI is wrong about a brand, it is wrong in nameable ways.

For 24 of the 503 companies in this benchmark, AI gets more than one in five facts wrong. For 7 of them, more than three in ten. That is not a rounding problem. It is a pattern.

Across 503 companies in 21 industries, scanned on four assistants, the average Brand Accuracy Rate was 90.1% (median 90.9%). Roughly one in ten verifiable claims AI makes about the typical company is wrong or outdated. Eighty-two companies scored a perfect 100%.

Accuracy is the headline, but it is not the whole picture. The benchmark measures four signals for every company: accuracy, AI visibility, machine-readability, and how well a brand's own narrative survives being retold. Read together, they explain not just whether AI is wrong, but why.

The story is not that AI gets brands mostly right. The story is that when it fails, it fails predictably. Those failures concentrate in a small set of companies. They do not spread evenly across the field.

We measured accuracy directly. For every company we compared what the assistants say against what its own site and trusted sources state, then sorted each failure into one of three modes: identity divergence, erasure and fabrication, and the homonym trap. Every mode has a distinct cause and needs a different fix. Each one also reveals something about how these models build a picture of a business, working from its name, its structured data, and whatever third parties have written about it.

Read one way, a 90.1% average is reassuring. Read it the way a buyer does, though, and the math shifts. One assistant, one answer, one moment of decision. A single wrong claim about your category, your pricing, or your ownership becomes the whole impression. This report is about those single wrong claims. Where they cluster, why they happen, and what moves them.

02 · At a glance

The benchmark in numbers

90.1%
Mean Brand Accuracy Rate across 503 companies (median 90.9%)
24
companies below 80% BAR
7
companies below 70% BAR
82
companies at a perfect 100%
68.7
mean GEO Score (site machine-readability)
85.5
mean Leakage Protection
28.6%
lowest single BAR in the benchmark
Each company was scanned across ChatGPT, Perplexity, Gemini, and Claude. Roughly fifteen buyer-intent queries, four platforms, four repetitions each, up to 240 responses per company. That is the raw material behind every number in this report.
03 · Distribution

Accuracy is not evenly spread

Most companies score well. The problem lives in a concentrated tail. There is also a surprisingly large group sitting at a flawless 100%.

7
Below 70%more than 3 in 10 claims wrong
17
70–79%more than 1 in 5 wrong
397
80–99%the healthy majority
82
100%every claim correct

The 24 companies below 80% are where the story lives. They are not spread evenly across industries or platforms. Instead they cluster around three specific failure modes, and a company usually fails for exactly one of them. Identify the mode and the fix is obvious. Miss it, and you can pour months into the wrong remedy, adding schema when the real problem is that your name is also an English word.

04 · Methodology

What we measured, and how

Four pillars, one benchmark. Brand Accuracy Rate is the headline; the other three explain why a score lands where it does.

PILLAR 01

Brand Accuracy Rate (BAR)

The share of verifiable AI claims that match the company's own site or trusted sources: pricing, capabilities, leadership, founding date. The headline metric.

Scale 0–100%
PILLAR 02

AI Visibility

A 60/40 blend of mention rate and prominence across buyer-intent queries. Platform breadth was dropped this year because it behaved as a near-constant.

Scale 0–100 · supporting
PILLAR 03

GEO Score

Machine-readability of the company's own pages, built from schema coverage and content readability. Both are Six Signals in their own right, listed next.

Scale 0–100
PILLAR 04

Leakage Protection

How well a brand's own narrative holds up in AI answers, versus being reframed by competitors or misrepresented outright.

Scale 0–100
04 · Methodology

The Six Signals

The six cells of the Faro mark are not decoration. Each one is a signal we measure, and together they compose every score in this report.

AI Visibilitymention rate & prominence
Brand Accuracy Ratethe headline signal
GEO Scoremachine-readability
Leakage Protectionnarrative integrity
Schema Coveragestructured data
Content Readabilityanswer-block clarity
Coverage: 504 companies completed, 503 in the reporting set after deduplication. The full query library runs to about fifty questions; the benchmark scores the high-intent subset that maps to real buyer behavior. Identity divergence uses a strict publish gate. A company has to show both high divergence and low accuracy to appear, so a rich synonym vocabulary alone never lands anyone on the list.
05 · Beyond accuracy

The other three signals

Accuracy tells you whether AI is right about you. These three tell you why, and where the leverage sits.

A brand can be highly visible and confidently misdescribed at the same time. Visibility without accuracy is not an asset. It is amplified misinformation.
68.7

GEO Score

Mean machine-readability across the field, and the single biggest gap in the dataset. Most sites are only two-thirds legible to a model. Strong schema helps a model retrieve you correctly. It will not, on its own, fix a name that collides with a common word.

Necessary, not sufficient
85.5

Leakage Protection

Mean score for how well a brand's own narrative survives being retold. Generally strong. The tail is where it breaks. Competitors reframe you, or an assistant quietly swaps a rival's positioning in for yours, in an answer the buyer reads as neutral.

Strong center, exposed tail
22.8

AI Visibility

Mean across the field, with a standard deviation of just 2.8. That band is too narrow to separate companies meaningfully this year. Visibility tells you whether AI mentions you. Accuracy tells you whether what it says is true. Only one of them loses a deal.

Too narrow to rank on
06 · Platform behavior

Four assistants, four ways of being wrong

When the platforms fail, they do not fail identically. Each has a characteristic way of getting a company wrong, and knowing the tendency tells you where to look first.

Perplexity
Denial & homonym collapse

Most likely to say a company "does not exist" when retrieval comes up empty, or to open with the dictionary meaning of a brand name. When it is right, it is often the most precise of the four. When it misses, it misses absolutely.

ChatGPT
Confident fabrication

Rarely admits a gap. Where it lacks grounding it constructs a plausible-sounding profile: a category, a use case, sometimes even pricing. It arrives with full fluency and no hedging, which makes it the most convincing wrong answer of the set.

Gemini
Name-derived invention

Tends to build an identity from the letters of the name itself. It reads "assembly" as manufacturing, or invents a mapping product from a company whose name sounds geographic. The failure is lexical, not factual.

Claude
Honest refusal

The most likely to say plainly that it does not have reliable information. Safer for accuracy. For a buyer-intent query, though, silence reads as "not a real vendor," which is its own kind of cost.

These are tendencies across the failing cohort, not per-platform accuracy scores. The same company can draw denial from one assistant, fabrication from another, and a correct answer from a third. That split is exactly what identity divergence looks like on the ground.
07 · In their own words

What AI actually says: a field guide

Every score in this report traces back to real answers. Here are the recurring types. The same handful of shapes appears again and again across 503 companies.

08 · Industry rankings

Where accuracy breaks down by industry

Worst to best, by mean Brand Accuracy Rate across all 21 industries. Marketing intelligence sits far below the field; everything else is tightly packed.

Marketing intelligencen=4
54.4%
AI infrastructuren=23
88.0%
AdTechn=43
88.3%
Hospitalityn=25
88.5%
Health systemsn=25
89.5%
MarTechn=22
89.5%
E-commercen=23
89.9%
Healthcaren=22
90.0%
DTC brandsn=25
90.2%
SaaSn=41
90.3%
InsurTechn=22
90.4%
Wealth advisoryn=25
90.6%
Cybersecurityn=22
90.7%
Law firmsn=25
90.7%
Real estaten=22
90.8%
Legaln=19
90.9%
Fintechn=20
91.6%
EdTechn=22
91.9%
HRTechn=24
92.0%
Higher edn=25
92.6%
DevToolsn=24
92.7%

Lowest is marketing intelligence at 54.4%. That is a small cohort of four, pulled down by the erasure cases in Failure Mode 2, so read it as directional rather than definitive. Fintech, long assumed the hardest vertical, averaged 91.6%. The counts across all 21 industries sum to 503.

Checking access…