No probe run yet. Run one to document how each model handles the roadside-vs-insurer split, ownership, state footprint and contact numbers.
Misattribution history
Misattribution = answers graded as conflated or tagged with an entity/product/ownership/contact/footprint error, divided by graded answers, counting each answer once. AI grading requires review; legacy runs have no pinned reference or model revisions, so changes are descriptive rather than controlled remediation proof.
No completed probe history.
| Run | Misattribution rate | Conflation rate | Graded answers |
|---|
This checks published crawl policy, not server/CDN logs, visit frequency or actual WAF blocking. Verify agents against provider-published IP ranges or documented forward-confirmed DNS, not geography alone. Re-check historical audits to apply the current bot taxonomy.
No check recorded yet.
Mentioned vs recommended
The paper's "Brand Consideration" gap, measured on this platform's own prompts. Not a share metric — a conversion-signal metric.
Re-run- Gemini 3.8 Flash0% / 50%
- GPT-5.6 Luna0% / 100%
- Claude Sonnet 50% / 100%
Dark bar = recommended · light bar = named. Ran 9/16/2026 across 12 prompts.
Evidence confidence
The paper's own caveats, and how this platform treats each one.
- Vendor-sourced data
- Granular visibility figures come from Somantra, an AEO vendor. They measure mentions/citations, not sales, and monthly counts swing with sampling (ChatGPT share 30.8% Jan vs 1.5% Feb 2026). Directional, not audited market share.
- Inference vs documented
- Entity confusion in actual LLM outputs was strong inference in the paper. The Entity Integrity Monitor on this page converts it into documented, graded evidence per model.
- Unconfirmed technicals
- robots.txt AI-bot rules and llms.txt presence needed a direct check. The Surface Hygiene audit performs it and reports anything its edge cannot reach as unconfirmed — never guessed.
- Moving target
- RG 234 was updated June 2026 to cover AI advertising; ChatGPT began citing brand homepages far more from mid-2026. Re-benchmark quarterly.