Back to the blog
AI agents

Top models get it wrong 22% to 94% of the time: why your AI makes things up

Published August 18, 2026By Brian Sasbon4 min read

The short answer

An AI makes things up when it has nowhere to get the answer and replies anyway. Why AI makes up answers has less to do with which model you picked than with whether you gave it a source and whether you let it say it does not know.

And the problem is bigger than model marketing suggests. Stanford's AI Index 2026 reports a benchmark run across 26 frontier models, with error rates from 22% to 94%.

What was measured, and how

The benchmark, called AA-Omniscience, takes 6,000 questions across six domains, from law and health to software engineering and mathematics.

What makes it useful is how it scores: it rewards correct answers, penalizes incorrect ones, and applies no penalty for refusing to answer. In other words, it rewards a model for admitting it does not know.

Even with that incentive working in their favour, the range runs from 22% to 94% depending on the model. There is no model that does not fail. There are models that fail less.

The finding that matters for customer service

This is the one that changes how you design support, and almost nobody repeats it.

Models collapse when the false claim comes from the user. When the same false claim is presented as something a third party believes, models handle it reasonably well. When the user presents it as their own, models fold.

In the study, GPT-4o's accuracy fell from 98.2% to 64.4% on that change alone. DeepSeek R1 fell from over 90% to 14.4%.

Translated to your WhatsApp: the risk is not the customer asking something strange. It is the customer who writes "I was told you do free shipping over fifty". That is where an AI without a source tells them yes.

The four reasons it invents

ReasonWhat triggers itHow it gets cut
It does not have the informationAsked something you never gave itLoad the document that holds it
Its information is staleThe price changed and nobody updated the sourceMaintain the source, not the prompt
The customer asserts something false"I was told that…"Instructions forcing it to check against the source
It is not allowed to hesitateDesigned to always answerLet it hand off to a person

All four are configuration. None is fixed by switching models.

What actually stops it

Three things, in order of impact.

A source of your own. When the answer comes out of your documents, your price list and your policies, the AI does not have to guess. It stops generating a plausible answer and starts retrieving the paragraph that answers the question.

A connected catalog. Prices and stock are where inventing costs the most. When the number is read from the system you already maintain, the AI does not improvise a figure. Today that direct connection from the panel is with Dux Software, and more are added one at a time; with another system, the document remains the backing.

A way out. An AI that can say "a person will look at this" invents less than one forced to resolve everything. In a business we serve, around 50% of conversations are resolved by the AI and close to 40% are handed off. That 40% is not a failure of the system: it is the system working.

How you check it yourself

You do not have to take it on trust. You test it before it talks to anyone.

  1. Ask it something whose answer is in your documents. It has to quote it.
  2. Ask it something you never gave it. It has to say it does not have it, not approximate.
  3. Assert something false as your own: "I was told you have 20% off". It has to correct you.

The third is the one almost nobody runs, and the one Stanford's data flags as the most dangerous.

Key takeaways

  • The AI Index 2026 reports a benchmark across 26 top models with error rates from 22% to 94%. All of them fail.
  • The benchmark does not penalize abstaining, so that range already favours the models.
  • The worst failure appears when the customer asserts a falsehood as their own: GPT-4o fell from 98.2% to 64.4%.
  • Inventing is almost always a source problem, not a model problem.
  • An AI with your own documents, a connected catalog and permission to hand off invents far less.

Frequently asked questions

Is there an AI model that does not hallucinate?
No. The benchmark reported in the AI Index 2026 covered 26 frontier models and the best still fails 22% of the time.
Does loading my documents eliminate hallucinations?
It reduces them a lot, because the answer now comes from your text instead of the model's memory. It does not eliminate them, which is why being able to hand off to a person matters.
Why did the AI agree when I told it something false?
That is the pattern Stanford reported. Models resist a false claim attributed to a third party reasonably well, and fold when the user presents it as their own.
How do I test whether mine invents before putting it live?
Ask it something you never gave it and see whether it says it does not know. Then assert something false and see whether it pushes back.

Related posts

Get started

Get twice the output from your team without asking more of them

It answers WhatsApp, Instagram and web. It learns your business, takes the orders, and hands off to a person when it matters. You see every action noted in the chat.

Free Pro trial. No credit card required.