Alcàsser Case: AI Bias Test with Three Models

A nine-day experiment queries ChatGPT-4o mini on the Alcàsser case: identical prompts yield three distinct AI responses, revealing significant bias.

English · Original discussion in Spanish · Published

Alcàsser Case: AI Bias Test with Three Models
Asking an AI about Alcàsser changes its answer

An artificial intelligence never saw the hair, the boat, or the house, yet it is still asked about them. During the last week of February 2025, daily queries to ChatGPT-4o mini tested its version of the Alcàsser case—the triple murder of Miriam García, Toñi Gómez, and Desirée Hernández, which occurred in 1992—and the result repeats an uncomfortable pattern: the model shifts its nuance depending on how the question is phrased. The experiment began on February 20 and concluded on the 28th with a final question about what would be needed for public opinion to accept a single narrative of the events.

What is discussed when asking about the Alcàsser case

The official version identifies Antonio Anglés and Miguel Ricart as responsible. Ricart was sentenced to 170 years in prison and was released in 2013; Anglés, considered the main suspect, has been a fugitive since 1993. Over this skeleton have been built three decades of doubt, and these doubts are transferred to the machine, one per day: the errors of the initial investigation, the impossible escape, the paper found in a bar, and contributions from those who questioned the judicial account.

The first question, on February 20, asked to list the "doubtful or controversial" aspects of the official version. The response did not hold back: poor police management, disputed evidence collection, and questionable identification of culprits.

The same question, three artificial intelligences, three responses

Other models soon joined the exercise. Grok 3 and Qwen 2.5 max, the Chinese model, answered the same question with different emphases. Where ChatGPT-4o mini remains skeptical, the Chinese model points directly to "presumed cover-ups by influential people" and, when asked if the case should be peine, replies that there are "solid arguments" to do so. None of the three agree completely.

The lesson is not who is correct. It is that three tools sold as intelligent offer three different portraits of the same case, and that the difference does not arise from the facts but from who drafts the query.

Why does the AI respond differently depending on how it is asked?

Because the bias comes from the asker. A loaded formulation—"what aspects are doubtful?"—presumes there is something doubtful; an open question about whether the events were errors or manipulation allows the official version to breathe. The models generate the answer the statement implies.

When asked if it might have been simple incompetence rather than manipulation, the same AI concedes that "it is a valid possibility." The machine does not seek the truth: it completes the sentence you provide. And whoever completes it ends up believing the machine has found the truth.

Anglés's escape, according to the AI

The official version states that Anglés escaped after a boat accident in 1993. The question of February 21 asked precisely that: what is doubtful about the escape. The response collects the known controversy—how he could evade authorities during a full police deployment—and leaves it unresolved.

It is the exact format of the experiment. The AI repeats the questions that have gone unanswered for thirty years and returns them in the form of a numbered list, with the appearance of a forensic report and the content of a press summary.

The paper, the commentators, and the cost of sustaining doubt

On February 24, the question concerned the contributions of Juan Ignacio Blanco and Fernando García, two names associated with critical commentary on the case. The AI recognizes them for having questioned the official account and kept public interest alive.

That interest has a price. Some report having spent more than 100,000 euros on lawyers, since 2016, due to lawsuits arising from this matter. And the episode of the paper found in a bar lingers, which part of the public reads as planted evidence and another as simple carelessness. The debate does not settle either extreme.

What the AI cannot change: statute of limitations and res judicata

No matter how many models recommend reopening the case. The legal path has a limit that no generated answer breaches: the case is time-barred for all involved except Antonio Anglés. Ricart was released in 2013.

This is the central contradiction of the entire exercise: asking the machine to reopen a file that the law has already closed, and the machine complies because it has nothing to lose.

The DNA of the hair and the indemnity in question

The final days of the experiment pointed to the laboratory. If no traces of Ricart's DNA are found in the analyzed hair, does anything change? And if Anglés's are not found? Should the State pay compensation?

The AI responds with conditionals—"could," "would lead to reevaluation"—and avoids promising anything. And it is worth remembering the elementary principle: the absence of proof is not the proof of absence.



With three models offering three different emphases, it is likely that the next question will produce another new answer. None will solve the case. But while the file remains time-barred and Anglés remains at large, someone will keep typing.

Summary of a discussion on Burbuja.info - Foro de economía, actualidad y política., translated from Spanish and reviewed before publication. Read the full discussion (84 replies).

More summaries

All summaries in English →

Back