NeuralTrust has been recognized by Gartner → Read more
Back

Codex vs Claude Code (2026): Benchmarks

Roger Howroyd September 29, 2026
Share
Codex vs Claude Code (2026): Benchmarks

Last updated: September 2026

Is Codex better than Claude Code in 2026?

In the Codex vs Claude Code decision, neither tool wins every test. On an independent Terminal-Bench 4.0 run, GPT-6 Astra from OpenAI's Codex costs less than half as much per task. Claude Opus 5.5, the default in Anthropic's Claude Code, scores higher. For enterprises, the deciding factors are usually cost per task, where the agent runs, and how much it may execute without a human.

Both vendors shipped new models this month. OpenAI released GPT-6 Astra on September 3 and GPT-6 Sol and Luna on September 22. Anthropic released Claude Opus 5.5 on September 22 and Claude Sonnet 5.5 on September 28. This guide compares them using only dated, published sources.

TL;DR: Key Takeaways

  • Independent benchmark: On Vals AI's Terminal-Bench 4.0 run, Claude Opus 5.5 scores 61.62% at $19.07 per task, GPT-6 Astra 57.07% at $8.21, and Claude Sonnet 5.5 53.03% at $19.33.
  • Product head to head: On the official Terminal-Bench 4.0 leaderboard, Codex with GPT-6 Astra scores 58.2% and Claude Code with Claude Fable 5.1 scores 57.9%, a statistical tie.
  • Vendor claims: On Terminal-Bench 4.0, Anthropic reports 70.6% for Sonnet 5.5 and OpenAI reports 57.9% for GPT-6 Astra.
  • Pricing: Both start at $20 per month, and both team seats cost $25 per user billed monthly, per the Codex and Claude pricing pages. Only Codex has a free tier.
  • Adoption: Claude Code's workplace use reached 39% and Codex's 16% in mid-2026, up from 18% and 3% in January, according to JetBrains Research.
  • Security: Both agents have shipped critical fixes for repository-borne attacks, including Codex's CVE-2025-61260 (CVSS 9.8) and Claude Code's CVE-2026-39861 (CVSS 10.0).

At a Glance: Key Differences in 2026

OpenAI CodexAnthropic Claude Code
Product typeCoding agent in the ChatGPT app, CLI, IDE extension and cloudAgentic coding tool in terminal, IDEs, desktop, web and mobile
Current modelsGPT-6 Astra, GPT-6 Sol, GPT-6 LunaClaude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1
Context windowGPT-6 Sol: 1,050,000 tokensOpus 5.5 and Sonnet 5.5: 1M tokens
Top list price (per 1M tokens)GPT-6 Astra: $10 input / $50 outputOpus 5.5: $4 input / $20 output
Individual plansFree, Go $8, Plus $20, Pro from $100Pro $20, Max from $100
Team seatBusiness: $25 monthly, $20 annualTeam Standard: $25 monthly, $20 annual
AI code review@codex review on GitHub, GitLab in betaCode Review on GitHub (research preview)
Models also served byMicrosoft Azure, Amazon BedrockAmazon Bedrock, Google Cloud, Microsoft Foundry
Best forCost-sensitive teams, parallel cloud tasks, GPT-6 usersTerminal-heavy engineers, CI automation, Google Cloud buyers

Try our AI Gateway today for free

What Are OpenAI Codex and Claude Code in 2026?

Codex and Claude Code are AI coding agents that read a repository, edit files and run commands for a developer. Codex is OpenAI's agent, built into the ChatGPT app and shipped as an open source CLI. Claude Code is Anthropic's agent, running one engine across terminal, IDEs, desktop, web and CI.

OpenAI Codex: the agent inside ChatGPT

Codex runs in the ChatGPT app, an IDE extension, the CLI and Codex cloud. The cloud version runs tasks "in isolated cloud environments" started from the web, GitHub, GitLab, Linear or Slack, per the Codex cloud docs. The CLI is Apache-2.0 licensed, with about 127,000 stars on GitHub.

The Codex models page lists three GPT-6 models: Astra, "our most capable model for complex work", Sol for coding and agentic workflows, and Luna for high-volume tasks. CLI 0.157.0, released September 25, added Sol and Luna with Amazon Bedrock support, per the Codex changelog. OpenAI's Tibo Sottiaux reported 8 million active users across Codex and ChatGPT Work in July, according to The New Stack.

Claude Code: Anthropic's agentic coding tool

Anthropic's documentation calls Claude Code "an agentic coding tool that reads your codebase, edits files, runs commands" and connects to your tools. It also runs in Slack, GitHub Actions and GitLab CI/CD.

Version 2.1.280 made Claude Opus 5.5 the default Opus model, and version 2.1.284 made Claude Sonnet 5.5 the default Sonnet model. Both have a 1M-token context window, per the Claude Code changelog. Anthropic says Sonnet 5.5 "generates outputs 30%+ faster than Sonnet 5" at the same $2 input and $10 output price per million tokens.

What it means for buyers: choosing Codex or Claude Code is also choosing GPT-6 or Claude 5.5 as your model family.

Codex vs Claude Code Benchmarks: What the Numbers Show

On independent tests, Claude Opus 5.5 posts the highest Terminal-Bench 4.0 score and GPT-6 Astra the lowest cost per task among frontier models. The only public product-level comparison, which tests each agent with its own model, shows Codex and Claude Code statistically tied.

Terminal-Bench 4.0: vendor claims vs independent runs

Terminal-Bench 4.0 measures "complex, multi-step professional tasks within a command-line interface". Both vendors now lead with it, so it is the best common yardstick.

ModelVendor-reportedVals AI independentVals cost per task
Claude Opus 5.566.4% (Anthropic)61.62%$19.07
GPT-6 Astra57.9% (OpenAI)57.07%$8.21
Claude Sonnet 5.570.6% (Anthropic)53.03%$19.33
GPT-6 SolNot reported34.34%Not listed
GPT-5.6 Sol (previous)37.3% (OpenAI)Not listedNot listed

Sources: Anthropic, OpenAI, Vals AI (updated September 27, 2026).

Vals runs every model in one neutral agent, Terminus 2, over three runs. Vals flags provider fallback on Anthropic rows: 30 of Opus 5.5's 198 attempts were served by older Claude models, and adjusting for them gives 53.54%. Sonnet 5.5's gap between 70.6% reported and 53.03% measured is the widest, so treat vendor numbers as upper bounds.

Codex vs Claude Code as products

The official Terminal-Bench 4.0 leaderboard, hosted by Harbor and the Laude Institute with Snorkel AI, ranks agent and model pairs. Codex with GPT-6 Astra leads at 58.2% (±2.8), ahead of Claude Code with Claude Fable 5.1 at 57.9% (±3.8). No Claude Code entry with a 5.5 model was listed on September 29.

FrontierCode, DeepSWE and CursorBench

Benchmark (who reports it)Codex-side modelsClaude Code-side models
FrontierCode 1.1 (Anthropic)GPT-6 Astra 53.3%, GPT-6 Sol 49.3%Opus 5.5 54.4%, Sonnet 5.5 52.1% (Xhigh)
DeepSWE v1.1 (OpenAI)GPT-6 Astra 74.1%, GPT-6 Sol 68.8%, GPT-6 Luna 66.6%Claude Fable 5 69.9%, Opus 5 66.0% (no 5.5 score)
CursorBench 4.0 (Cursor)GPT-5.6 Sol 41.7% (no GPT-6 entry yet)Opus 5.5 57.8%, Sonnet 5.5 55.5%

Sources: Anthropic, OpenAI, Vellum, CursorBench 4.0 (September 10, 2026).

Each vendor picks favorable benchmarks: OpenAI's DeepSWE table compares GPT-6 with older Claude models, and CursorBench had not scored GPT-6 at its last update. For the model-level view, see our Claude Opus 5.5 enterprise guide and our GPT-6 Astra analysis for CISOs.

What happened to SWE-bench?

OpenAI stopped reporting SWE-bench Verified in February 2026, finding that "at least 59.4% of the audited problems have flawed test cases", per its SWE-bench Verified post. It recommends SWE-bench Pro, but Scale's public SWE-Bench Pro leaderboard had not scored GPT-6 or Claude 5.5 models by September 29.

How we compared Codex and Claude Code

We used only published, dated sources: vendor pages, product documentation, independent leaderboards, surveys and security advisories. Vendor-reported scores are labeled, and prices are US list prices in September 2026. We ran no private tests, so treat these figures as a starting point for a pilot on your own repositories.

Cost per Task and Token Efficiency

GPT-6 Astra is the cheaper frontier option per task, while Claude's list prices are lower per token. On Vals AI's Terminal-Bench 4.0 run, Astra costs $8.21 per task against $19.07 for Opus 5.5, because a cheaper token still costs more when a model uses more of them.

Per 1M tokensInputCached inputOutputWhere it runs
GPT-6 Astra$10Separate rate$50Codex
GPT-6 Sol$2$0.20$10Codex
GPT-6 Luna$0.10$0.01$0.50Codex
Claude Opus 5.5$4$0.20$20Claude Code
Claude Sonnet 5.5$2$0.20$10Claude Code

Sources: OpenAI, OpenAI API, Vellum, Anthropic.

On Vals's numbers, Astra costs about $0.14 per Terminal-Bench point and Opus 5.5 about $0.31. Sonnet 5.5 and GPT-6 Sol share a list price, yet Sonnet 5.5 scored 53.03% against 34.34% for Sol. On CursorBench 4.0, Sonnet 5.5 at High effort scores 47.8% for $1.67 per task.

Effort levels also drive cost: Claude Code defaults to Medium, and GPT-6 Sol offers six levels. Gartner (2026) predicts that "by 2028, AI coding costs will overtake the average developer's salary", in a June 2026 press release. Set effort defaults and model routing in your rollout policy.

Codex vs Claude Code Pricing: Plans and Limits

Codex and Claude Code have nearly identical paid tiers: $20 for an individual plan, from $100 for power users and $25 per team seat billed monthly. Codex adds free and $8 tiers and publishes estimated message ranges, while Claude Code starts at Pro.

Plan levelCodex (ChatGPT plans)Claude Code (Claude plans)
FreeFree: GPT-6 Luna in the desktop app, subject to rolloutFree plan does not include Claude Code
EntryGo: $8/month, GPT-6 LunaNot offered
IndividualPlus: $20/month, GPT-6 Sol and Luna, code review, SlackPro: $20/month ($17 billed annually)
Power userPro: from $100/month, 5x or 20x Plus limitsMax: from $100/month, 5x or 20x Pro usage
TeamBusiness: $25/user monthly, $20 annualTeam Standard: $25/seat monthly, $20 annual
Team, heavy useCredit extensionTeam Premium: $125/seat monthly, $100 annual
EnterpriseContact sales$20/seat/month annual plus usage at API rates

Sources: ChatGPT Learn, Codex pricing; Claude pricing. US list prices, September 29, 2026.

OpenAI estimates 15 to 150 GPT-6 Sol messages per five-hour window on Plus, and says these are "not fixed message limits". Anthropic lists usage multiples but no message counts. Both accept API keys at standard rates.

Workflows: CLI, IDE, Desktop and Cloud Agents

Both agents now cover the terminal, IDE, desktop, browser, mobile and pull request review. Codex centers on the ChatGPT app and parallel cloud tasks. Claude Code centers on a terminal-first engine that also runs in CI pipelines.

FeatureCodexClaude Code
Main surfacesChatGPT desktop app, CLI, IDE extension, web, iOSTerminal, VS Code, JetBrains, desktop, web, iOS and Android
Project instructionsAGENTS.mdCLAUDE.md, and it can read AGENTS.md
ExtensibilityMCP, skills, plugins, Codex SDKMCP, skills, hooks, subagents, Agent SDK
Background and cloud workCodex cloud tasks, scheduled automations, worktreesCloud sessions, Routines, background agents, /loop
AI code review@codex review and @codex security review on GitHub; GitLab in betaCode Review for GitHub (research preview, Team and Enterprise)

Sources: Codex GitHub code review; Codex cloud; Claude Code overview; Claude Code Review.

Where Codex is stronger

Codex suits teams that run many tasks in parallel without tying up laptops. Cloud tasks start from GitHub, GitLab, Linear or Slack, and @codex security review adds a security pass on pull requests. Security teams can audit the open source CLI.

Where Claude Code is stronger

Claude Code suits teams that automate engineering work in pipelines, with native GitHub Actions and GitLab CI/CD plus hooks and subagents. Its Code Review sends "a fleet of specialized agents" over each pull request, averaging $15 to $25 per review and 20 minutes per run. The Claude Code vs Cursor comparison covers its IDE story.

Developer Adoption: What the Surveys Say

Claude Code leads adoption, but Codex grew fastest in 2026. JetBrains found that Codex's workplace use grew about five times in six months, and its awareness rose from 27% to 65% of developers.

SurveyCodexClaude Code
JetBrains, use at work (May to July 2026, 15,000+ developers)16% (up from 3% in January)39% (up from 18% in January)
Pragmatic Engineer (906 respondents, early 2026)"60% of Cursor's usage"Most used; "most loved" by 46%

Sources: JetBrains Research (August 2026); The Pragmatic Engineer (March 2026).

According to JetBrains (2026), 31% of developers name Claude Code as their most-used AI tool. Both agents often arrive on personal accounts before any procurement review.

Enterprise Controls and Data Governance

Both vendors offer SSO, SCIM, audit logs and no training on business data by default. OpenAI lists more certifications and data residency regions, while Anthropic offers more cloud deployment routes. Retention rules differ by plan, so check them against your policy.

ControlCodex (OpenAI)Claude Code (Anthropic)
SSO and SCIMSAML SSO on Business; SCIM and EKM on EnterpriseSSO on Team; SCIM, audit logs and Compliance API on Enterprise
Managed policyrequirements.toml managed configurationManaged settings; option to turn auto mode off
Training on customer dataNot by default for Business, Enterprise and APINot under commercial terms
RetentionConfigurable for qualifying orgs; zero data retention on the API platform30 days standard; zero data retention per organization on eligible Enterprise accounts
CertificationsSOC 2 Type 2, ISO/IEC 27001, 27017, 27018, 27701SOC 2 Type 2, ISO 27001
Data residency10 regions, including the US, Europe and the UKNot listed on the pricing page

Sources: OpenAI, Codex admin, ChatGPT Work FAQ, Claude pricing, Claude Code data usage, Claude Code security.

Anthropic says zero data retention "is not included in the standard Enterprise plan", and Claude Code keeps transcripts "locally in plaintext" for 30 days. OpenAI's Compliance Logs Platform keeps data for 30 days, so export it continuously.

Security and Governance: Running Coding Agents Safely

Codex and Claude Code share one core risk: they read untrusted content from repositories, config files, MCP servers and the web, then act with a developer's permissions. Both have fixed critical flaws where a malicious repository could run commands, so defaults and runtime controls matter more than brand.

Sandboxes and approval defaults

ControlCodexClaude Code
Local defaultworkspace-write sandbox, network offAuto mode classifier when no mode is set (v2.1.284)
Strictest optionread-only sandboxManual mode, read-only start
Riskiest optiondanger-full-access, approvals neverbypassPermissions mode (org policy can disable it)
Automated approvalauto_review reviewer agentAuto mode classifier model
OS sandboxSeatbelt on macOS, bubblewrap on LinuxSandboxed bash with filesystem and network isolation
Cloud executionOpenAI containers, agent phase offline by defaultIsolated VMs; GitHub credentials never enter the VM

Sources: Codex sandboxing; Codex approvals and security; Claude Code security; Claude Code changelog.

One change deserves a policy review now. Claude Code 2.1.284 makes terminal and VS Code sessions "start in auto mode when no permission mode is configured", so a classifier approves most actions instead of the developer. Trust verification "is disabled when running non-interactively with the -p flag", which is how CI jobs run.

A timeline of notable vulnerabilities

DateToolIssueSeverityStatus
Sep 2025Claude CodeProject hooks consent bypassCVSS 8.7Fixed in 1.0.87
Oct 2025Claude CodeCVE-2025-59536: project code runs before the startup trust dialogCVSS 8.8 (NVD)Fixed in 1.0.111
Jan 2026Claude CodeCVE-2026-21852: API key exfiltration via ANTHROPIC_BASE_URLCVSS 5.3Fixed in 2.0.65
Feb 2026CodexGitHub token theft via branch name command injectionNo CVE listedPatched February 5, 2026
Apr 2026CodexCVE-2025-61260: project .env and .codex/config.toml run MCP commandsCVSS 9.8Affects CLI 0.23.0 and earlier
Apr 2026Claude CodeCVE-2026-39861: symlink sandbox escapeCVSS 10.0 (NVD)Fixed in 2.1.64
May 2026CodexWindows RCE via web search injection and node.bat hijackNo CVEOpenAI closed as "not reproducible"
May 2026Claude CodeDeeplink RCE via claude-cli:// linksNo CVSS publishedFixed in 2.1.118
Sep 2026CodexCVE-2026-19591: PowerShell approval bypassCVSS 8.8Affects CLI 0.72.0 to 0.130.0

Sources: The Hacker News on Claude Code; The Hacker News on Codex; GitHub Advisory; GitLab Advisory; NVD on CVE-2025-59536 and CVE-2026-39861; Cymulate; Strix.

The pattern repeats: project config files that execute before a human approves anything. The deeplink flaw, found by joernchen of 0day.click, is explained in our Claude Code RCE analysis. In the Codex token case, BeyondTrust warned that agent access can become "a scalable attack path" into enterprises.

Mapping both agents to OWASP risks

OWASP ranks prompt injection as LLM01, including injection from "external sources, such as websites or files". The OWASP Top 10 for Agentic Applications maps these incidents:

  • ASI01 Agent Goal Hijack: poisoned web results or READMEs, as in the Codex Windows chain. See our indirect prompt injection guide.
  • ASI04 Agentic Supply Chain Vulnerabilities: malicious MCP entries in .codex/config.toml or .claude/settings.json. See MCP Security 101.
  • ASI05 Unexpected Code Execution: hooks, approval bypasses and sandbox escapes such as CVE-2026-19591 and CVE-2026-39861.

The generated code needs checks too: the Veracode 2026 GenAI Code Security Report found an average security pass rate of 56%, "virtually unchanged since last year's report".

How NeuralTrust secures Codex and Claude Code

NeuralTrust adds one runtime security layer across both agents:

  • Agent Gateway (TrustGate): routes agent traffic to models, MCP servers and APIs through one policy point, with identity-based access and budgets. TrustGate ships with 200+ pre-built MCP servers.
  • Agent Runtime Security (TrustGuard): inspects prompts, tool calls and responses in real time to block injection, secret leakage and destructive commands before they run. It enforces least privilege against excessive agency.
  • AI Red Teaming (TrustTest): tests your coding agent setup against poisoned repositories and MCP servers before rollout.

Which Should You Choose? Codex or Claude Code

Choose Codex if cost per task, parallel cloud tasks or an open source client matter most. Choose Claude Code for the top independent Terminal-Bench score, deep CI automation or a Google Cloud route. Many teams run both, which makes a shared policy layer essential.

If you are...ChooseWhy
A team optimizing cost per completed taskCodexGPT-6 Astra at $8.21 vs $19.07 per task on Vals Terminal-Bench 4.0
An engineer who wants the highest independent scoreClaude CodeOpus 5.5 at 61.62% on Vals Terminal-Bench 4.0
A platform team automating CIClaude CodeGitHub Actions, GitLab CI/CD, hooks and the Agent SDK
A team running many tasks in parallelCodexCloud tasks from GitHub, GitLab, Linear and Slack
A buyer standardized on Google CloudClaude CodeRuns through Google Cloud, Amazon Bedrock or Microsoft Foundry
A CISO approving agents at scaleEither, behind one gatewayOne policy, one audit trail, runtime checks on every tool call

Conclusion

The Codex vs Claude Code choice in 2026 is close on capability and different on cost. Claude Opus 5.5 posts the top independent Terminal-Bench 4.0 score, while GPT-6 Astra comes close at under half the cost per task. Codex brings a free tier, an open source CLI and wider certifications. Claude Code brings stronger adoption, CI integration and a route through Google Cloud. Both fixed critical repository-borne flaws this year, so set approval defaults and runtime controls before scaling either.

Secure Codex and Claude Code in Production with NeuralTrust

Run Codex, Claude Code or both behind one policy layer, with real-time protection for every prompt, tool call and MCP connection.

Try our AI Gateway today for free

Related Comparisons

FAQs about Codex vs Claude Code

1. Is Codex better than Claude Code?

Neither wins outright. On Vals AI's Terminal-Bench 4.0 run, Claude Opus 5.5 scores 61.62% and GPT-6 Astra 57.07%, but Astra costs $8.21 per task against $19.07. On the official leaderboard, Codex with Astra (58.2%) and Claude Code with Fable 5.1 (57.9%) are within error margins.

2. Which is cheaper, Codex or Claude Code?

List prices match: $20 for an individual plan, from $100 for power tiers and $25 per team seat billed monthly. Per task, GPT-6 Astra is cheaper than Opus 5.5 on current independent data ($8.21 vs $19.07 on Terminal-Bench 4.0). Per token, Claude is cheaper at the top end: Opus 5.5 costs $4/$20 against Astra's $10/$50.

3. Is Codex free to use?

Partly. OpenAI's Free plan includes GPT-6 Luna in the desktop app, subject to rollout, and Go costs $8 per month. GPT-6 Sol, code review and Slack start with Plus at $20. Claude Code is not in Claude's Free plan and starts with Pro at $20 per month.

4. Can you use Codex and Claude Code together?

Yes. Both read an AGENTS.md file, so one set of project instructions can serve both agents. The trade-off is governance: two agents mean two permission systems, two audit trails and two bills. A shared gateway keeps policy and logging consistent across both.

5. What models do Codex and Claude Code use in 2026?

Codex uses OpenAI's GPT-6 Astra, GPT-6 Sol and GPT-6 Luna, and picks a recommended model when none is set. Claude Code uses Anthropic models, with Claude Opus 5.5 and Claude Sonnet 5.5 as the default Opus and Sonnet models since September 2026.

6. Which is more secure for enterprise use, Codex or Claude Code?

Neither is secure by default. Both have fixed critical flaws where a malicious repository could run commands, including Codex's CVE-2025-61260 (CVSS 9.8) and Claude Code's CVE-2026-39861 (CVSS 10.0). OpenAI lists more certifications and residency regions. Add least privilege, approval gates and runtime monitoring either way.

7. What is Terminal-Bench 4.0?

Terminal-Bench 4.0 tests AI agents on complex, multi-step tasks in a command-line interface, and it now leads both vendors' launch pages. The official leaderboard ranks agent and model pairs, while Vals AI runs every model in one neutral agent for a like-for-like view.

About the Author

Roger Howroyd is Head of Global SEO and AI at NeuralTrust, where he leads the company's search strategy across SEO, AEO, GEO, and LLM optimization. He specializes in AI-powered search, content strategy, and SEM. Connect on LinkedIn.

NeuralTrust is the leading platform for securing and scaling AI agents. Named a Pioneer in the Gartner Emerging Market Quadrant for AI Application Security 2026, recognized across four Gartner Hype Cycle reports in 2026, and featured in the Gartner Market Guide for Guardian Agents 2026, the Gartner Market Guide for AI Gateways 2025 and the KuppingerCole Leadership Compass for Generative AI Defense 2025. Headquartered in Barcelona with offices in London and New York. ISO 27001 certified.

Sources

  1. Anthropic: Introducing Claude Sonnet 5.5 (September 2026)
  2. Anthropic: Introducing Claude Opus 5.5 (September 2026)
  3. OpenAI: GPT-6 Astra (September 2026)
  4. OpenAI: Introducing GPT-6 Sol and Luna (September 2026)
  5. OpenAI API Docs: GPT-6 Sol (September 2026)
  6. ChatGPT Learn: Codex changelog (September 2026)
  7. ChatGPT Learn: Codex models (September 2026)
  8. ChatGPT Learn: Codex cloud (September 2026)
  9. GitHub: openai/codex (September 2026)
  10. Claude Code Docs: Overview (September 2026)

Subscribe to our newsletter

Share

Join the leaders securing the agent ecosystem

Get a Demo