/ai-search593 words

Burson Study Names Visibility-Believability Gap in Generative Search

Burson's new Generative Engine Optimization research argues that brand visibility in AI-generated answers and the believability of those mentions are not the same thing — and that marketers tracking only one are missing the point.

Burson Research Highlights Gap Between Visibility and Believability in Generative Engine Optimization - Burson | Global
Burson Research Highlights Gap Between Visibility and Believability in Generative Engine Optimization - Burson | GlobalAI-generated
  • Burson, a WPP-owned global communications agency, released new research on Generative Engine Optimization.
  • The study separates two outcomes marketers typically conflate: visibility (mention frequency in AI answers) and believability (trust and accuracy of those mentions).
  • GEO operates on a different output surface than classical SEO, with AI products such as ChatGPT, Google AI Overviews, Perplexity and Claude collapsing results into single-paragraph answers rather than ranked URLs.
  • Three recommended monitoring signals are branded mention frequency, sentiment and accuracy of mentions, and the share of citations linking to first-party or high-authority third-party sources.
  • No standardized believability benchmark exists yet from a major platform or vendor, leaving brands to combine volume tracking with manual answer audits.

Burson, the WPP-owned global communications agency, has released research examining how brands appear inside AI-generated search answers. The study's headline finding, embedded in its title, separates two outcomes that marketers frequently treat as identical: visibility — being mentioned in a generative-engine response — and believability — being mentioned in a way that readers actually act on.

What is Generative Engine Optimization, and why is it under review now?

Generative Engine Optimization, abbreviated GEO, names the practice of shaping content so large language models surface a brand inside synthesized answers rather than inside a list of blue links. The discipline sits next to classical SEO but operates on a fundamentally different output surface. Where Google historically returned ten ranked URLs with visible titles, meta descriptions and source domains, products such as ChatGPT, Google's AI Overviews, Perplexity and Claude collapse the decision into a single paragraph — often with no source attribution visible to the reader at all.

That shift is what makes Burson's framing consequential. A brand can be cited on every relevant prompt and still lose the conversion when the model's description is generic, inaccurate, or stripped of the differentiators that justify a purchase.

What does the visibility-believability gap look like in practice?

The gap, as the report frames it, sits between mention volume and mention quality. The first metric is straightforward to count: how often a brand surfaces across a panel of prompts run against multiple models. The second is harder — it asks whether the AI's framing matches the brand's positioning, whether third-party corroboration shows up in the model's training data, and whether the reader walks away with the impression the brand intended to project.

For search marketers, this reframes how to report results. A GEO dashboard that only tracks presence is reporting on the wrong axis. The metrics that map to revenue — citation accuracy, sentiment, contextual specificity, and share-of-model across the answer — are the ones that determine whether visibility becomes a click, a store visit, or a qualified lead.

How does this reshape measurement and strategy?

Traditional SEO tooling measures rank, impressions and click-through rate against a relatively stable SERP. GEO tooling has to sample answers across prompts, models and geographies — outputs that change week to week, sometimes by the hour. Burson's argument is that believability cannot be inferred from volume alone; it requires qualitative review of how the model represents the brand, not just whether the brand name appears.

That distinction puts weight back on earned media, Wikipedia accuracy, analyst coverage and authoritative third-party reviews. These inputs shape model representations, and they have always functioned as GEO work — long before anyone used the acronym.

What should teams monitor next?

Three signals offer the clearest read on whether a brand is closing or widening the visibility-believability gap:

  • Branded mention frequency inside AI answers across the major generative platforms
  • Sentiment and accuracy of those mentions, scored against the brand's positioning
  • Share of citations that link back to first-party or high-authority third-party sources

Until tooling vendors ship a standardized believability score, brands auditing AI-search presence will need to combine volume tracking with manual answer reviews, sentiment checks, and source-quality audits. Burson's research is the first major agency-side study to name this gap as a structural problem rather than a tactical edge case. The next milestone to monitor is whether any major platform — a model lab, an SEO vendor, or a trade publisher — produces the first benchmark specifically designed to measure it.

via Google News: generative engine optimization (Source)

More from Tom Whitfield

Tom Whitfield

Show full bio

News editor covering marketplaces and e-commerce at SERP Journal.

62 articles