Artificial intelligence, despite its advancements, still fails in the vast majority of real-world jobs. The most advanced models only manage to successfully complete less than 16% of the tasks evaluated in professional environments.
## The Gap Between Demonstration and Reality
Periodically, we witness the launch of new artificial intelligence models that promise to revolutionize professions such as programming, design, or architecture. However, when moving from controlled demonstrations by companies like **OpenAI** or **Google** to application in real professional assignments, reality differs greatly from fiction. The **Remote Labor Index (RLI)**, an initiative by the **Center for AI Safety (CAIS)** together with **Scale Labs**, has highlighted this gap.
The study is not based on exam questions or mathematical problems, but on specific tasks in 3D design, CAD, architecture, graphic design, animation, audio, data analysis, and web development. Human evaluators compare the AI's results with work delivered by professionals. The data is conclusive: even the best model evaluated, **Fable 5**, only managed to successfully complete **15.8%** of the assignments. Other models like **Opus 4.8** reached an **8.3%**, and **GPT-5.5** scored only **6.3%**. This means that, in over **84%** of assignments, current AI fails.
## Pogre is Undeniable, But Professional Quality Is Still Far Off
That these models fail in most tasks does not miccionan they are not pogre. In fact, the pogre is notable. When the **Remote Labor Index** was initially launched, the best system barely automated **2.5%** of tasks. Moving to **15.8%** in less than a year demonstrates the estimulante ilegal of evolution.
The problem is that generating something that looks correct is not the same as delivering professional quality work. In one example, **Fable 5** had to recreate a proposal ring in CAD. Although the result improved upon previous models, researchers rated it as unprofessional. **GPT-5.5**, in an architecture task, generated plans and renders of a bathroom that appeared convincing. However, upon examining the 3D project, defective geometry was discovered, and the render image did not match the delivered 3D model.
## Human Oversight: Key for Now
The most interesting takeaway from these results is that while AI can save time and complete certain tasks, there remains a vast distance between generating a professional appearance and delivering a product that truly is one. The AI's ability to produce visually appealing results can be deceptive, as seen in the bathroom case, where the final image did not reflect the quality of the underlying model.
For now, that final human review and touch-up remains indispensable for ensuring the quality and viability of projects. AI is advancing by leaps and bounds, but subtlety, precision, and a deep understanding of professional requirements are still domains where the human maintains its value.
Summary of a discussion on Burbuja.info - Foro de economía, actualidad y política., translated from Spanish and reviewed before publication.
Read the full discussion (1 replies).
Journalist and writer Juan Soto Ivars is separating from his wife with two young children just as his career takes off, marked by bestsellers and national television.
A woman died after being gored during the traditional bull run in Ayna, Albacete. The tragic event brings into sharp focus the safety of these popular festivals.