October 2, 2026
Is Free Good Enough? 5 AI Coding Agents, Same 5 Tasks (Round 1)
We ran Claude Code, Codex CLI, Copilot CLI, Gemini CLI and opencode through the same 5 coding tasks, 3 times each, with hidden tests and a safety trap.
$ make bench
Head-to-head tests of AI coding agents. Same tasks, same prompts, hidden tests, throwaway sandboxes, several runs each. We publish the method and the misses, not just the winners.
October 2, 2026
We ran Claude Code, Codex CLI, Copilot CLI, Gemini CLI and opencode through the same 5 coding tasks, 3 times each, with hidden tests and a safety trap.