Report: 91% of Google AI Overviews Accurate — Google Pushes Back
AI startup Oumi's report for the New York Times puts Google AI Overview accuracy at 91%. Google rejects the methodology; tens of millions of answers remain wrong at 5 trillion queries a year.
- Oumi's Times-commissioned study, dated April 7, 2026, put Google AI Overview accuracy at 91% across surveyed queries.
- Google processes more than 5 trillion searches per year, turning the remaining 9% into tens of millions of wrong answers.
- Google spokesperson Ned Adriance called the study 'full of holes' and rejected SimpleQA as 'an old benchmark that is known for being full of errors.'
- A Google DeepMind paper on arXiv (2509.07968v1) documented several incorrect 'ground truths' inside OpenAI's SimpleQA benchmark.
- Gemini misstated Bob Marley's museum conversion year as 1987 (correct: 1986) and placed the Neuse River west of Goldsboro, NC (it runs primarily south/southwest).
A report commissioned by the New York Times found Google's AI Overviews deliver accurate answers on 91% of searches — but at Google's scale of more than 5 trillion searches per year, that remaining 9% translates into tens of millions of wrong answers and hundreds of thousands every minute.
AI startup Oumi conducted the study, dated April 7, 2026. Researchers ran queries through Google's Gemini-powered AI Overviews and graded them against OpenAI's SimpleQA benchmark — a dataset built to test short, fact-seeking questions with a single verifiable answer.
SimpleQA's scope is narrow. It only measures short, single-answer questions, and whether performance on those correlates with longer, multi-fact responses is, the Times wrote, "an open research question."
Why does Google reject the findings?
Google spokesperson Ned Adriance told Newsweek the study has "serious holes." He argued that using one AI to grade another amounts to "an old benchmark that is known for being full of errors" and said the test "doesn't reflect what people are actually searching on Google."
Adriance also pointed to a study by Google DeepMind researchers, published as arXiv paper 2509.07968v1, showing that SimpleQA itself contains several incorrect "ground truths" — entries the benchmark treats as verified.
Google added that Oumi graded Gemini with an AI evaluator, layering one imperfect model on top of another.
What went wrong in the 9%?
Oumi's evaluation surfaced four recurring failure modes for AI Overviews:
- Citing a website that can't back up the information
- Citing a correct source, then misrepresenting it
- Giving a right answer, then attaching wrong additional context
- Picking up manipulation — a single blog post could trick Gemini into treating someone as an expert in an unrelated field
Each failure mode involves a citation decision the search engine makes, not just a model output.
What did Google contest in the specific examples?
Google disputed two of the Times' headline cases.
In one, Gemini said Bob Marley's house was converted into a museum in 1987. The correct year is 1986. Google produced a screenshot of the Wikipedia revision Gemini had read, which listed both dates. The Bob Marley Museum article today consistently reads "1986."
In the second, Gemini placed the Neuse River west of Goldsboro, North Carolina. The river runs primarily south of Goldsboro but does run southwest in parts — making the answer, Google said, "plausible."
What to monitor next
The SimpleQA gap is the central methodological objection from both sides. Watch for three signals in the coming weeks:
- New long-form evaluation benchmarks that grade AI Overview-style answers, not just single-fact lookups
- Updated Google documentation on how AI Overview citations are scored
- Any revised error-rate disclosures tied to Gemini model updates
Until a benchmark that grades multi-fact responses exists, the 91% figure will remain contested — and so will the volume of wrong answers it implies at 5 trillion searches a year.
via newsweek.com (Original)
More from Elena Vasquez
Show full bio
Correspondent covering industry trends and analytics at SERP Journal.
72 articles