Same model, same bugs, same fixes: a UC Berkeley and Arena study found Claude Code cost about 2x what Pi, a tiny open-source coding agent, cost. The HarnessTax study ran 7 models through three coding agent harnesses (Claude Code, Codex CLI and Pi) on real GitHub bugs from SWE-bench Lite. The 2x is the average across those 7 models, with token volume priced at one 2026-09-01 API list price, not anyone's billed spend. The biggest visible difference is setup. Claude Code's first call carried 27,011 tokens with 23 tools by default, against Pi's 1,972 tokens and 4 tools, and that setup is resent every turn (the authors note that caching and output tokens matter too). For Claude Fable 5 that meant $1.33 vs $0.67 per attempt at about the same success rate, but Sonnet 4.6 and Haiku 4.5 showed no clear gap, and model choice still moves cost far more than harness choice. ABOUT THE NUMBERS - Study: HarnessTax (2026-09-16) by Melissa Z. Pan, Shuo Yang, Negar Arabzadeh, Ion Stoica, Matei Zaharia (UC Berkeley) and Wei-Lin Chiang (Arena). - Setup: Claude Code 2.1.224, Codex 0.146.0, Pi 0.85.1. 30 SWE-bench Lite and 30 Terminal-Bench 2.0 tasks, 3 attempts each, 100-turn cap, high effort. - Pricing: the same list for every harness, with cached tokens almost certainly priced at cache rates (our inference). Pi used subscription keys, and the key type for Claude Code and Codex is unstated. - Claude Code vs Pi on SWE-bench Lite: Fable 5 2.00x, Opus 4.8 2.06x, Kimi K3 1.72x, GPT-5.6 Sol 3.50x, GPT-5.6 Luna 5.09x. Sonnet 4.6 (0.99x) and Haiku 4.5 (1.14x) showed no clear gap, and the study doesn't say why. Vs Codex: 1.6x. Terminal-Bench vs Pi: about 1.5x. - Success moved within about 2% on average (about 5% on Terminal-Bench). Claude Code vs Pi, Fable 5 scored 97.8% vs 96.7% and Opus 4.8 86.7% vs 82.2%, both gaps "not clear" per the authors. - The first-call figure (27,011 vs 1,972 tokens) is the mean across 7 models and includes the task prompt and cache tokens. Codex averaged 11,308. Contents and traces were unpublished as of 2026-09-28. - Not from this study: Systima (~33k vs ~7k, Claude Code vs OpenCode, July 2026) and Portkey (~27k vs ~2.6k, trivial task, April 2026). Different versions, models and metrics. - Companion paper (arXiv 2609.20804, other authors): four models on the authors' own lightweight harness, tested on SWE-Bench Verified and Terminal-Bench 2.1. It compares bash-only with predefined tools, not Claude Code with Pi. - Plans: since April 2026, Claude subscriptions don't cover third-party harnesses like Pi. VentureBeat reports Agent SDK credits (Pro $20, Max 5x $100, Max 20x $200); the start date is unconfirmed. - Not covered: Opus 5.5 or prices after 2026-09-22 (the tested Claude Code can't run Opus 5.5). We didn't rerun the study; as of 2026-09-28 we found no independent rerun or Anthropic response. TIMESTAMPS 0:00 - Same fix, two bills 0:25 - Why agent bills feel mysterious 0:56 - What a harness is (the harness tax) 1:40 - How Berkeley and Arena tested it 2:18 - The results: 2x Pi, same success 3:12 - Where the extra tokens go 4:14 - The twist: not every model pays 5:15 - The catch 6:02 - Who actually pays it 6:41 - Verdict LINKS & RESOURCES - HarnessTax study: https://harnesstax.github.io/ - Arena post: https://arena.ai/blog/coding-agents-harness-tax - Data, cost ratios: https://harnesstax.github.io/data/charts/system-harness-effect-swe.f233b4a6e8.json - Data, first call: https://harnesstax.github.io/data/charts/system-agent-context-swe.06f521bdb1.json - Data, tokens and turns: https://harnesstax.github.io/data/charts/system-cumulative-accrual-swe.2b6a6206a5.json - Data, cost by model: https://harnesstax.github.io/data/charts/system-frontier-swe.8fd44a1e0f.json - Data, Terminal-Bench: https://harnesstax.github.io/data/charts/system-harness-effect-tb.4d5a1affd4.json - Pi, by Mario Zechner: https://mariozechner.at/posts/2025-11-30-pi-coding-agent/ - How Claude Code works: https://code.claude.com/docs/en/how-claude-code-works - Claude Code costs: https://code.claude.com/docs/en/costs - Claude API pricing: https://platform.claude.com/docs/en/about-claude/pricing - Claude Code changelog: https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md - BleepingComputer: https://www.bleepingcomputer.com/news/artificial-intelligence/anthropic-is-cutting-claude-codes-current-weekly-limits-by-17-percent/ - TechCrunch: https://techcrunch.com/2026/04/04/anthropic-says-claude-code-subscribers-will-need-to-pay-extra-for-openclaw-support/ - VentureBeat: https://venturebeat.com/technology/anthropic-reinstates-openclaw-and-third-party-agent-usage-on-claude-subscriptions-with-a-catch - Simon Willison: https://simonwillison.net/2026/Sep/22/opus-and-sol-and-luna/ - Companion paper: https://arxiv.org/abs/2609.20804 - Systima: https://systima.ai/blog/claude-code-vs-opencode-token-overhead - Portkey: https://portkey.ai/blog/the-harness-tax/ TAGS #ClaudeCode #HarnessTax #CodingAgents #AIAgents
The information provided is not trading advice. kdj.com does not assume any responsibility for any investments made based on the information provided in this article. Cryptocurrencies are highly volatile and it is highly recommended that you invest with caution after thorough research!
If you believe that the content used on this website infringes your copyright, please contact us immediately (info@kdj.com) and we will delete it promptly.