ChatGPT hallucinates more than ever: 48% failure rate on person test
In mid-April 2025, as the industry continued selling the arrival of general artificial intelligence as an imminent fact, OpenAI published the results of an uncomfortable internal test. In PersonQA, an evaluation designed to measure accuracy regarding specific people, the o3 model failed 33% of the questions. The o4-mini, its latest lightweight version, reached 48%. Error rates double or triple those of the previous generation.
The test that dismantles ChatGPT's reliability
The data, collected by TechCrunch, compare four models from the o family. The o1 failed 16%, the o3-mini 14.8%, the o3 33%, and the o4-mini 48%. This is not just any regression: it is a worsening trend as new versions are released. Company executives may call them "hallucinations," but in practice, this means the model invents people, data, and sources with total naturalness.
- o1: 16% error rate on PersonQA
- o3-mini: 14.8%
- o3: 33%
- o4-mini: 48%
Someone consulting ChatGPT to document an article, a report, or a file is citing an interlocutor that cannot distinguish between truth and verisimilitude. And that, applied at scale, is an informational health problem.
"With total certainty, but invented": what the user checks at home
The statistics have their domestic reflection. A real case is that of the orange picker in Valencia in 1970: ChatGPT's response placed the salary between 40 and 60 pesetas daily, when reality was around 200. Grok, asked the same question, gave an approximate range of 100 to 300 pesetas. Closer, but without precision.
Another example illustrates the mechanism: to the question of a Spanish word with the five vowels, ChatGPT responded with "abstemious," which has four. The tool does not check; it completes. It is a kind of tweeter that never admits it does not know.
Lite models, paid versions, and the AGI smoke
There are nuances that qualify the headline. o3 and o4 are not the most powerful models, but lightweight versions, and on them the disaster has been measured. Furthermore, the difference between the free and paid versions is abysmal: the same question about vowels is solved well in the 240 euros per month model, although the free access model fails it in seconds. The problem is that most common people use the free version, and it is the one generating poisoned trust.
Parallel to this, the landscape has fragmented. Grok, which last summer was the carefree alternative, has fallen in quality and incorporated censorship. Claude maintains better conversation but limits the free plan. DeepSeek, within its pro-Chinese bias, convinces some. The sense of stagnation is general: image generators do what they want, text models respond with overwhelming confidence and false data, and AGI moves further away.
And yet, the bubble remains inflated. Those selling the imminent general artificial intelligence speak of a future that is not supported by the PersonQA numbers. Current models are, at best, supervised assistants that save time on specific tasks. At worst, statistical parrots flooding the internet with plausible garbage. The question is when it will be reflected in the balances of those who have based their valuation on the AGI promise.