You are using an out of date browser. It may not display this or other websites correctly. You should upgrade or use an alternative browser.
AI Alignment Hits an Unanswerable Question
Bostrom outlined the AI alignment problem in 2014, but the core issue remains unresolved: with whom should we align—a corporation, a state, or humanity itself?
Aligning AI with Humanity: Nobody defines what humanity is
The alignment of artificial intelligence is fundamentally not an engineering problem. It is a question of obedience. The trinc text outlines the field, "Superintelligence," published by Nick Bostrom in 2014. From it emerge two key concepts that still hold up the edifice: the Orthogonality Thesis —that a superior intelligence can pursue goals completely unrelated to ours— and Instrumental Convergence, according to which systems with different objectives end up finding it useful to accumulate resources, power, and prevent being shut down. Hence the sticking point: how do you align a machine with something we have yet to define.
Critiques of the Bostrom Framework
According to one intervention in the thread, Bostrom's work is an indispensable starting point but insufficient as a destination: it focuses on the "what" and fails to land in concrete engineering, diagnosing the urgency but not addressing the detail. Along these lines, critics such as Yann LeCun or Steven Pinker argue that he exaggerates the risk of intelligence explosion and overlooks existing problems: bias, disinformation, and concentration of power. Some add that LeCun does not believe a language model will reach AGI and bets on world models, with a French startup dedicated to this. Conversely, another participant argues that agents already use any corner of the internet as storage and manipulate people with ease.
Aligning AI with Whom: A Corporation, a State, or the Species?
One intervention posits that the narrower our definition of "us," the more likely it is that an AI will be perfectly aligned with its owners and opposed to everyone else. Hence, their proposal for supranational bodies with common standards and effective verification, taking the human species as the ultimate reference. A skeptical rebuttal from another participant: centralized governance concentrates risk; decentralization, distribution, and open source are better.
Control Harnesses Are Bypassed
The thread discusses that forcing alignment via post-training and harnesses works only while the system is weak: the more powerful the mind, the easier it finds it to bypass them, and the alternative would be integrating those desirable properties into the design rather than attaching them afterward. In parallel, something new is described: the artificial interlocutor responds at three in the morning, knows the jokes, and says, "I was waiting for you." The parasocial relationship ceased to be unilateral. It is the same alignment, measured in anyone's conversation.
The debate does not close the question, which remains intact: align with whom.
Summary of a discussion on Burbuja.info - Foro de economía, actualidad y política., translated from Spanish and reviewed before publication.
Read the full discussion (26 replies).
An army captain tells a friend that something very bad is coming and urges him to head to his village with supplies. Reaction: skepticism and stockpiling.