← Back to Insights

ChatGPT Work Launches as OpenAI Folds Codex Into Desktop App to Rival Claude Cowork

Nils Liu
OpenAI ChatGPT Work Codex GPT-5.6 Anthropic Claude Cowork AI 代理人 News

TL;DR

OpenAI launched ChatGPT Work on July 9, an autonomous office agent built on GPT-5.6, and folded Codex into a new three-tab desktop app alongside Chat and Work. The launch lands three days after Anthropic's Claude Cowork expanded to mobile and web.

ChatGPT Work Launches as OpenAI Folds Codex Into Desktop App to Rival Claude Cowork

OpenAI launched ChatGPT Work on July 9, an autonomous agent that turns office tasks straight into finished documents, and at the same time folded the standalone Codex app into a rebuilt ChatGPT desktop app split into three tabs: Chat, Work, and Codex. The agent runs on the freshly released GPT-5.6 model family, and its pitch is that it can stay on a task for hours, break a goal into steps, connect to Slack, Gmail, and Google Calendar, and hand back a finished spreadsheet, deck, document, or even an interactive web app. The timing is pointed: it landed three days after Anthropic expanded Claude Cowork from desktop to mobile and web.

Here’s my read, and I’d genuinely like to see it challenged: this office-agent arms race won’t produce a clear winner within the next six months. The deciding factor won’t be benchmark scores, it’ll be whichever company first drives down the friction of “does a human need to approve this step” while still keeping users confident the output isn’t garbage. If you’ve already run the same batch of tasks through both ChatGPT Work and Claude Cowork inside your own company, does the approval frequency and error rate you’re seeing match that read?

What Happened

ChatGPT Work is live today for Pro, Enterprise, and Edu plans, with OpenAI saying it will reach Plus and Business within days. Free users don’t get the Work tab yet, but the new desktop app’s Chat and Codex tabs are open to every plan, free included. The old ChatGPT desktop app has been renamed ChatGPT Classic, and anyone who had the standalone Codex app installed needs to manually update to the merged app to get the three-tab layout.

The agent runs on GPT-5.6, released just three days earlier in three tiers: Sol, Terra, and Luna. Sam Altman’s line at launch was a practical one: every enterprise is now thinking about spend versus the value they’re getting for it. The number he attached was a claimed 54% reduction in tokens spent on agentic coding tasks compared to GPT-5.5, which works out to more than half the cost for the same work.

ChatGPT Work’s operating logic is to take a goal, connect to whatever tools the user has authorized, break the job into smaller steps, and execute them one at a time, pausing to ask for human approval before anything sensitive. That plan-then-execute-then-check-in pattern is close to the same script Anthropic ran a week earlier when Claude Cowork expanded to mobile and web. Both companies’ product copy circles the same idea: shipping finished work, not just conversation.

What the Numbers Actually Tell You

Start with that 54% figure. It’s OpenAI comparing its new model against its own previous generation, GPT-5.5, not against Claude Opus 4.8 or Gemini on the same batch of tasks. No independent benchmark group has re-run that number against real enterprise workflows yet, and that gap hasn’t been filled.

Working from first principles, a token-efficiency gain like this usually comes from engineering work on the training or inference side, trimming redundant reasoning output, tuning routing, compressing intermediate steps, not necessarily an architectural breakthrough. The token savings themselves are the less interesting part of this story. What’s worth watching is that OpenAI and Anthropic shipped nearly identical product shapes in the same week: an agent that stays on a task for hours, plugs into enterprise tools, and hands back finished documents. That convergence isn’t a coincidence. Both companies have concluded the chat window is tapped out for growth, and the next battleground is finishing the user’s work for them.

Some scale math helps here. Say a mid-size company runs 3,000 agent tasks a month, each averaging 50,000 output tokens. On GPT-5.5 pricing, that’s roughly 3,000 times 50,000 times $30 divided by a million, or about $4,500. If the 54% savings claim holds, the same workload on GPT-5.6 runs closer to $2,070, saving nearly $30,000 a year. That number is attractive on its face, but what enterprises actually care about is task completion rate more than token price. An agent that saves tokens but frequently gets things wrong and needs a rerun ends up costing more in practice, and there’s no independent data yet to confirm or deny that either way.

There’s an engineering-reality detail worth flagging too. ChatGPT Work advertises hours of independent work on complex projects, but OpenAI’s own documentation says it pauses for human approval before sensitive actions, which puts real distance between the marketing and full autonomy. Behind the “ships finished work” pitch is a semi-automated workflow that still leans on a human checking in at the key moments. Folding Codex into the ChatGPT desktop app also reads as OpenAI consolidating a product line that had sprawled across Codex CLI, the standalone Codex app, and ChatGPT itself. That’s good news for maintenance overhead, but it also means users have to relearn an interface.

Metrics Worth Watching Next

Three things should give an answer within three to six months. First, whether an independent group runs ChatGPT Work and Claude Cowork through the same batch of enterprise tasks and measures actual completion rates and token savings, instead of each company citing its own numbers. Second, whether Google or Microsoft ship a comparable office-agent product in this wave. If nothing shows up within three months, that tells you those companies think enterprise customers aren’t ready to hand workflows to an autonomous agent yet. Third, whether paid conversion and real usage frequency hold up once ChatGPT Work expands from Pro, Enterprise, and Edu to Plus and Business, a number that will say more about actual enterprise appetite than any benchmark.

Also worth tracking: how often these agents ask for approval on sensitive actions. If OpenAI or Anthropic quietly loosen that approval threshold, it signals trust is building faster than expected, and that shift will matter more than any single model upgrade.

If this was useful, subscribe to the newsletter for weekly AI PM insights and GenAI case studies.

Sources: MacRumors, Forbes


Related reading:

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.