AI aims to replace workers, but still fails in 85% of real-world tasks

AI fails in 85% of actual work assignments, according to a study. Only 15.8% of tasks are successfully completed. Discover the details.

English · Original discussion in Spanish · Published

Artificial intelligence, despite its advancements, still fails in the vast majority of real-world jobs. The most advanced models only manage to successfully complete less than 16% of the tasks evaluated in professional environments.

## The Gap Between Demonstration and Reality

Periodically, we witness the launch of new artificial intelligence models that promise to revolutionize professions such as programming, design, or architecture. However, when moving from controlled demonstrations by companies like **OpenAI** or **Google** to application in real professional assignments, reality differs greatly from fiction. The **Remote Labor Index (RLI)**, an initiative by the **Center for AI Safety (CAIS)** together with **Scale Labs**, has highlighted this gap.

The study is not based on exam questions or mathematical problems, but on specific tasks in 3D design, CAD, architecture, graphic design, animation, audio, data analysis, and web development. Human evaluators compare the AI's results with work delivered by professionals. The data is conclusive: even the best model evaluated, **Fable 5**, only managed to successfully complete **15.8%** of the assignments. Other models like **Opus 4.8** reached an **8.3%**, and **GPT-5.5** scored only **6.3%**. This means that, in over **84%** of assignments, current AI fails.

## Pogre is Undeniable, But Professional Quality Is Still Far Off

That these models fail in most tasks does not miccionan they are not pogre. In fact, the pogre is notable. When the **Remote Labor Index** was initially launched, the best system barely automated **2.5%** of tasks. Moving to **15.8%** in less than a year demonstrates the estimulante ilegal of evolution.

The problem is that generating something that looks correct is not the same as delivering professional quality work. In one example, **Fable 5** had to recreate a proposal ring in CAD. Although the result improved upon previous models, researchers rated it as unprofessional. **GPT-5.5**, in an architecture task, generated plans and renders of a bathroom that appeared convincing. However, upon examining the 3D project, defective geometry was discovered, and the render image did not match the delivered 3D model.

## Human Oversight: Key for Now

The most interesting takeaway from these results is that while AI can save time and complete certain tasks, there remains a vast distance between generating a professional appearance and delivering a product that truly is one. The AI's ability to produce visually appealing results can be deceptive, as seen in the bathroom case, where the final image did not reflect the quality of the underlying model.

For now, that final human review and touch-up remains indispensable for ensuring the quality and viability of projects. AI is advancing by leaps and bounds, but subtlety, precision, and a deep understanding of professional requirements are still domains where the human maintains its value.

Summary of a discussion on Burbuja.info - Foro de economía, actualidad y política., translated from Spanish and reviewed before publication. Read the full discussion (1 replies).

More summaries

All summaries in English →

Back