Remote Labor Index Update: Fable 5 Leads AI Automation at Just 16.1% of Real Freelance Work
TL;DR
The latest Remote Labor Index results show Claude Fable 5 automating 16.1% of 240 real freelance projects, the highest score yet. Is that a breakthrough, or proof AI still cannot replace human freelancers? We ran the unit economics.
I have not found a scoring shortcut that reliably predicts what RLI’s human raters will accept without just paying for human raters. The benchmark’s own numbers show automated judges overrating GPT-5.5 by roughly 2.9x and Opus 4.8 by roughly 2.3x compared to the human panel, while still preserving the ranking order. If you have run a calibration study that gets an LLM judge within, say, 10 percentage points of human acceptance rates on creative or technical deliverables, not just code correctness, I would like to see the setup.
The Center for AI Safety and Scale AI published the latest Remote Labor Index results on July 1. The benchmark runs 240 real projects sourced from freelance marketplaces, spanning 23 professional domains: 3D modeling and CAD, architecture, graphic design, video and audio editing, data analysis, web development, and more. Every deliverable gets scored by human raters against the actual paid deliverable a real client received. Anthropic’s Fable 5 leads at 16.1% automation, Opus 4.8 sits second at 8.3%, and OpenAI’s GPT-5.5 is third at 6.3%. This round only completed 218 of the 240 projects, since the U.S. government restricted access to Fable 5 mid-evaluation, the same restriction that took the model offline for 19 days in June, a story I covered in a separate post.
The index launched in October 2025 with GPT-5.2 leading at 2.5%. Nine months later: Opus 4.6 hit 4.17%, Opus 4.8 hit 8.3%, and now Fable 5 sits at 16.1%, each generation roughly doubling the last. CAIS’s announcement calls it a significant increase in digital labor automation.
240 Real Freelance Projects, and AI Clears Only 16.1% of Them
RLI’s scoring bar is different from most benchmarks. It does not ask whether a model produces a correct answer, it asks whether a paying client would accept the deliverable. Each model gets up to 24 hours of wall-clock time, access to an A100 GPU, and more than 30 professional applications including Blender, FreeCAD, GIMP, Kdenlive, and Audacity. Models run inside agent scaffolds like Claude Code or Codex CLI, and every deliverable passes through an independent critic agent before submission, a worker-critic loop. Budget is $50 per task by default, bumped to $150 for Fable 5 because its per-task inference cost runs roughly 6x higher than Opus 4.8.
The failure patterns are worth noting. Of the failed attempts, 45.6% fell short of professional quality, 35.7% were incomplete or malformed, and the rest split between technical errors and internal inconsistencies across project files. Models do best on generative work built from scratch, audio and image tasks especially, and worst on projects that require matching an existing spec or editing existing files incrementally.
What the Numbers Actually Say
16.1% itself is not in dispute. It is Scale AI and CAIS’s own data, the methodology is public in the arXiv paper, and anyone can rerun it. What deserves more scrutiny is using that number to argue AI is about to replace white-collar freelance work.
Start with cost. Fable 5 burns $150 per attempted task against a 16.1% success rate, which works out to roughly $930 in compute per deliverable that actually clears the bar. The human freelancer doing the same job charges a median of $200. That gap holds even before counting the labor cost of someone reviewing the AI’s output, since, as the researchers themselves note, judging whether an RLI deliverable meets professional standard is itself a demanding agentic task, not something you can offload to a quick automated check.
That points to a second problem: how much to trust automated scoring at all. When the same test set gets graded by an LLM judge instead of a human panel, GPT-5.5’s score comes out inflated by roughly 2.9x, and Opus 4.8 by roughly 2.3x. The rankings hold, the absolute numbers do not. Any headline claiming a model hit some benchmark percentage should be discounted hard whenever the grading was done by another model instead of a human panel.
Last, the growth curve. Going from 2.5% to 16.1% looks exponential, but it is not one architecture steadily improving, it is four separate model generations layered on top of an agent toolchain that keeps getting more capable. Budget went from $50 to $150, tool access went from a handful of apps to more than 30. Some of that curve belongs to the engineering harness, not to the model getting smarter.
Metrics Worth Watching Next
First, whether the next RLI update keeps pace with roughly doubling every few months. If it does, automation could approach 30% by year end, a specific number worth writing down now and checking later, rather than a vague forecast either way.
Second, whether other benchmarks that rely on LLM judges get caught with the same inflation gap RLI found. If that spreads, the way the industry cites benchmark scores needs a broad recalibration.
Third, whether Fable 5’s per-task cost converges toward Opus 4.8’s as inference gets cheaper. That is the number that decides when automated freelance work turns from a research curiosity into a viable business.
Fourth, whether U.S. government access restrictions interrupt RLI evaluation again. This is the second time a policy action has cut a test run short, and the timing of the next one is itself a signal worth tracking.
If this was useful, subscribe to the newsletter for weekly AI PM insights and GenAI case studies.
Related reading:
Related Articles
Google Launches Gemini 3.7 Flash: Half-Price Until Year-End, With Agent Costs Still Tied to Retry Rates
Google positions Gemini 3.7 Flash as a workhorse for coding and AI agents, with higher vendor benchmarks and temporary half-price access, while architecture, training methods, and production retry rates remain undisclosed.
Databricks Raises $5 Billion at a $190 Billion Valuation to Fund Enterprise AI Agents
Databricks closed a $5 billion round at a $190 billion valuation to fund Lakebase, Genie, and Unity AI Gateway, while its operating figures remain company-reported.