ALIA, Sánchez's state AI: open source and doubts over its corpus

Government presents ALIA as Spanish-language AI, but a user calculates corpus at 39.31% English and 16.12% Spanish

English · Original discussion in Spanish · Published

ALIA, Sánchez's state AI: open source and doubts over its corpus
Sánchez launches ALIA: more corpus in English than in Spanish

The Government presented ALIA as the artificial intelligence "in Spanish" promoted by the State. Public models, open source and a vocation to foster research "based on our language", in the words of Pedro Sánchez. The event was named "HispanIA 2040". The promise clashed almost immediately with a fact that a forum user claims to have checked in the model's public profile: 39.31% of material in English and only 16.12% in Spanish.

In short, only a sixth of the training would be in the language the project boasts about.

What is ALIA and what does the Government promise

The official proposal is summarized as trinc: an open language model, developed with public money, intended to foster research in Spanish. Open source, free access, state investment. On paper, technological sovereignty.

So far the narrative. The discomfort appears in the same speech: according to several messages in the thread, Sánchez warns about the "technocaste" and what type of AI should be developed. That warning becomes slippery when the State enters to compete, with public budget, against tools that have cost those same companies years and billions of euros.

The corpus cited in the thread: 39% in English, 16% in Spanish

The technical test deflates the headline. According to the links shared in the thread, the model is distributed on the industry's reference platform and there the real training weights are listed. English dominates comfortably; Spanish remains at less than a sixth of the total. For a project sold as a bet on the language, the proportion is, at the very least, questionable.

To this is added usability. A forum user who presents himself as a software engineer described that he spent five minutes trying a query and didn't succeed. Others pointed out that the public version returns errors when tested and that the page is a labyrinth where access to the model cannot be found. The underlying suspicion: that a good part of the work consists of fine-tuning, with public money, an already existing open model.

Why the language of the corpus is not a minor detail

A model learns from what it reads. If 39.31% of its material is in English and only 16.12% in Spanish, its default behavior will tend to reproduce the logic of the dominant language: turns of phrase, references and even biases. For a tool that is sold as Spanish, the proportion is not cosmetic, it is the core of the product.

The public cost and the shadow of the chiringuito (a shady public body)

Where criticism is most fierce is in the bill. The dominant reading is that it is another public body —with its staff of advisors, experts and middle managers— rather than a technological product. In the thread the comparison with Irene Montero's application appears.

The argument has a documented basis. Spain accumulates dozens of publicly funded technology parks whose weight on employment is uneven. Employment in these enclosures stands at 2% in the Basque Country, 1.8% in Catalonia and 1.7% in Andalusia, and drops to 1.6% in Asturias and Cantabria. In Málaga, the technology park concentrates 3.1% of the province's workers. Innovation incubators for some; façades with payroll for others.

For Hacienda (tax agency) and Sanidad (health system): the antiestéticared use

The most repeated antiestéticar is not that the AI works badly, but that it works well in the wrong direction. A hypothesis launched in the thread: that a state model ends up cross-referencing bank movements, instant transfers and small payments in search of the slightest tax irregularity. If the tool detects patterns, the first to take advantage would be the tax collector. For Hacienda, according to this reasoning, any investment pays for itself.

At the other extreme there are those who imagine the opposite: an indoctrinated AI, which responds with bias, dodges certain topics and treats anyone who asks too much as suspicious. The two suspicions —the tax eye and the party handbook— coexist and feed each other. And neither is ruled out with an announcement.

From innovation to control: how the conversation drifted

As the hours passed, the matter migrated from the technical to the geopolitical terrain. It was no longer discussed whether the model is good, but who should control the next critical infrastructure. There are those who maintain that the State does not compete against China, but against the citizen.

Even the name of Elon Musk appears as a counterexample: it is assured that Spain would be better off with half a dozen profiles like that than with a state plan like ALIA. The paradox is served: a Government that warns against the concentration of technological power presents its own AI, with its corpus, its biases and its bill.

Does anyone really believe that the objective is research?

Summary of a discussion on Burbuja.info - Foro de economía, actualidad y política., translated from Spanish and reviewed before publication. Read the full discussion (232 replies).

More summaries

All summaries in English →

Back