Is ChatGPT accurate? OpenAI answered that question itself in August 2025, in more detail than any outside critic has managed, and the answer was published in a table. Then the company stopped publishing the table.

Both halves of that are worth knowing before you rely on anything it tells you.

The Numbers OpenAI Published

The GPT-5 system card gives error rates for two very different conditions, and the gap between them explains most of the confusion in this argument.

On SimpleQA, a benchmark of short factual questions with browsing switched off, the strongest model was accurate 55 percent of the time and hallucinated 40 percent of the time. GPT-4o, a model millions of people used daily for a year, was accurate 44 percent of the time and hallucinated 52 percent.

On real ChatGPT traffic with browsing on, the same model got 4.5 percent of its factual claims wrong, and 12.7 percent of its responses contained at least one major error.

Those two sets of numbers describe the same system. Ask it a hard fact from memory and it is close to a coin flip. Let it look things up in the ordinary flow of use and roughly one response in eight carries a significant error. Neither figure is small enough to trust blind, and the second is the one that matters for work.

Here is the part nobody reports. Every OpenAI system card since has dropped SimpleQA reporting in favour of an internal benchmark of user-flagged errors with no absolute numbers in it. The 2026 cards say the current models are less likely to reproduce hallucinations than the previous ones. They do not say how often they hallucinate. The last comparable published figure is from August 2025.

A developer walking a colleague through where the model got it wrong

Why It Guesses Instead of Saying It Does Not Know

In September 2025 OpenAI published a paper called Why Language Models Hallucinate, and its central claim is the most useful thing written on the subject: models hallucinate because standard training and evaluation reward guessing over admitting uncertainty.

The analogy in the paper is a multiple-choice exam that gives no credit for a blank answer. Under that scoring, guessing always beats abstaining, so a system optimised against it learns to produce a confident answer every time.

The worked example is stark. One small model abstained on 52 percent of questions and was wrong on 26 percent. Another abstained on 1 percent, scored two points higher on accuracy, and was wrong on 75 percent. Slightly better accuracy, three times the errors, and on a leaderboard the second model looks superior.

The practical consequence is that the confidence of an answer carries no information about whether it is right. The system was trained to sound certain. That is the product working as designed.

Is ChatGPT Accurate Enough for Professional Work?

Domain by domain, the measured answers are worse than the headline benchmarks suggest.

Stanford researchers ran more than 800,000 queries about real federal court cases across four models. On directly verifiable questions, hallucination rates ran from 58 percent for GPT-4 to 88 percent for the worst model tested. A follow-up study of purpose-built legal research tools, the expensive ones with retrieval attached, found 17 to 33 percent hallucination rates depending on the product.

On bibliographic citations, a study in Scientific Reports found GPT-4 fabricated 18 percent of references outright, and of the ones that were real, 24 percent contained substantive errors. For book chapters the fabrication rate was 70 percent.

The largest audit of news answers, run across 22 public service broadcasters in 18 countries by the European Broadcasting Union and the BBC, found 45 percent of AI responses contained at least one significant issue, most often a sourcing problem. That figure covers four assistants rather than ChatGPT alone.

The pattern across all of them is consistent: the more checkable the claim, the more often it is wrong. Citations, case names, quotes and figures are precisely where the model is weakest and precisely where a reader assumes it is strong.

The machines behind the answer, which have no idea whether it was right

What It Costs When Nobody Checks

A New York lawyer filed a brief in 2023 citing six cases that did not exist. The court sanctioned him and his firm $5,000 and ordered them to write to every judge whose name had appeared on a fabricated opinion.

That was the famous one. As of today there are 2,046 recorded cases worldwide in which a court has found that a party relied on hallucinated material, across 76 jurisdictions. The database only counts findings, never allegations, so the real figure is higher.

Companies carry the liability too. A Canadian tribunal ordered Air Canada to pay a customer misled by its own website chatbot, rejecting the argument that the chatbot was somehow a separate entity. If your business puts a model in front of customers, its confident wrong answers are yours.

The quietest finding is the one about experienced people. In a randomised trial of 44 physicians, all of whom had completed AI literacy training, diagnostic reasoning accuracy fell from 84.9 percent to 73.3 percent when the AI suggestion contained an error. Training did not protect them. It is a preprint with a small sample, so treat it as a signal rather than a settled result, but the direction matches everything else: people accept a fluent wrong answer from a machine more readily than from a colleague.

Where It Is Reliable

Accuracy is the wrong single axis, because the error rate depends almost entirely on the kind of task.

Transformation is safe. Summarising a document you supply, restructuring your own argument, translating, turning rough notes into prose, explaining a concept you can already sanity-check. The material is in front of it and the failure mode is a bad edit rather than an invented fact.

Recall is not safe. Anything the model has to retrieve from training rather than from something you gave it: statistics, citations, case names, dates, prices, who said what. That is the category every benchmark above is measuring, and it is where the 40 to 50 percent figures live.

The working rule is the difference between asking it to handle information and asking it to supply it. Handling is what the tool is genuinely good at. Supplying is where you become the fact checker, and most people move between the two mid-conversation without noticing they have crossed a line.

The Check That Takes a Minute

Three habits cover most of the risk, and none of them require a better model.

Ask for the source and open it. Not the citation, the link. Fabricated references look perfect until you click. This single step catches nearly every error that would embarrass you.

Ask the same question twice in separate conversations. Where the model knows, you get the same answer. Where it is filling a gap, the two versions diverge, and the divergence tells you which claims to verify.

Decide in advance what class of work never ships unchecked. Anything with a number, a name, a date or a legal consequence. OpenAI’s own terms of use tell you to do exactly this: you must evaluate output for accuracy, including human review, before using or sharing it. The company that built it does not claim it is reliable.

There is a cost to the checking, and it is worth being honest about it, because the checking is what the time saving is spent on. A tool that returns an answer in five seconds and needs four minutes of verification is still useful. It is just not the four-minute saving it appeared to be, which is the same arithmetic error that makes most productivity purchases look better than they are.

One Detail Worth Noticing

OpenAI’s own help page explaining what ChatGPT is currently states that the product has limited knowledge of events after 2021 and is not connected to the internet. Both statements stopped being true years ago.

That is a small thing and it points at a real one. Documentation drifts, capabilities change monthly, and the public’s model of what these systems can do is built out of whatever they read at some point in the past. Confidence about how accurate the tool is tends to be as stale as the page it came from.

Trust reflects this. Only about 20 percent of people globally say they trust AI chatbot answers about the news, and among those who use chatbots regularly the figure is 44 percent against 17 percent for everyone else. The more people use it, the more they trust it. Whether the accuracy justifies that gradient is the open question.

What the Honest Answer Sounds Like

Is ChatGPT accurate? On ordinary questions with browsing on, roughly one response in eight from the best model tested contained a major error, and that was the last time the number was published. On hard factual recall it is close to a coin toss. In law, medicine and citations it is worse than its general benchmarks imply.

It is still an extraordinary tool, and none of the above argues for going back to doing the work by hand. It argues for treating fluent output the way you would treat a confident stranger: useful, often right, and never the last word on anything that carries a consequence.

Public trust already reflects this. Among Americans who use chatbots for health information, 48 percent call them highly convenient and only 18 percent call them highly accurate. People have worked out what the product is. The risk is not that anyone believes it blindly, it is the quiet drift that happens when an answer arrives fast, sounds right and nothing forces you to look. The same drift decides whether a policy gets kept for thirty years or quietly dropped, which is the mechanism behind why most whole life policies never pay off: the failure is never the decision you made deliberately.