$ make bench

Agent Reviews

Head-to-head tests of AI coding agents. Same tasks, same prompts, hidden tests, throwaway sandboxes, several runs each. We publish the method and the misses, not just the winners.