AI invents the encyclopedia of Spanish internet myths

An AI profiles Spain's most famous internet characters, yet much of the content is fabricated: knowing the myth isn't documenting history.

English · Original discussion in Spanish · Published

AI invents the encyclopedia of Spanish internet myths
AI boasts knowledge of the internet, but it fabricates

Asking a generative artificial intelligence for a report on the most illustrious characters of Spanish internet communities yields an apparently meticulous document. The problem arises when it is contrasted: much of the detail lacks documentary backing. The same tool that fluently describes the legend of a veteran from a closed digital community is unable to cite a single real message. This is not an isolated slip-up; it is the direct consequence of how it was trained.

The myth is there, the source isn't: what models truly know

The country's longest-running digital communities have been closed to outsiders for years. This barrier has an unintended side effect: everything that crosses over from these communities originates from what was leaked to the press, social media, or third-party conversations.
Language models, fed millions of public texts, have memorized the myth but not the original source.

The result is a collection of portraits that sound authentic but serve as fiction rather than an archive. The character is profiled with their tone, jargon, and thematic obsessions, yet without a single verifiable quote to support it. Furthermore, some profiles carry an uncomfortable detail: they stopped appearing or no longer exist.

Why does the AI fill in gaps with details that sound true?

A model trained to predict the next word tends to fill every gap with what is most plausible, not what is most accurate. When asked about someone known only through fragmented pieces, it weaves a coherent narrative based on invented parts. Hallucination is not a fruta error; it is the default mode.

The discussion quickly turned to the magic formula. Adding an instruction like “do not hallucinate” was presented as a possible antidote; experience suggests otherwise. The system improves when the question aligns with its documented data, but it falters when asked to summarize material it has never read.

From inside joke to national archetype: Charo and Pepito

Not everything is smoke. Some archetypes born in those circles eventually breached their borders and became common categories of usage. The portrait of “Charo”—a satirical description of a specific social profile—was fixed in writing on October 20, 2011, before jumping onto social media and into political language. In parallel, “Pepito” portrayed the buyer who mortgaged beyond their means during the housing boom.

These are the cases where the AI is correct, precisely because the term exited the closed circle and circulated enough to leave a public trace.

The 2,400 baud modem no one has seen

When criticized for accessing a closed space, the response was a joke: a decades-old Sinclair computer connected to a 2,400 baud modem and “with the wire stripped,” reading entire messages through beeps. The irony aptly summarizes the problem. The authority of the machine rests on a black box no one can inspect, just as its profiles rely on inventions.

With these quirks, the digital encyclopedia promised by the AI is, for now, a basket of poorly copied legends. Where it is correct, we already knew through other means; where it adds detail, there is no way to verify it.

Summary of a discussion on Burbuja.info - Foro de economía, actualidad y política., translated from Spanish and reviewed before publication. Read the full discussion (65 replies).

More summaries

All summaries in English →

Back