Hacker News
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
The study evaluated coding agents Claude, Codex and Cursor across 75 synthetic repositories in ten languages, generating 16,893 runs and filtering 5,292 valid sessions. Experiments used four persona prompts, rotating three sandbox providers, and a “simulated human” orchestrated by Gemini 3.7 Flash to mimic realistic interactions, revealing how tool choices shift when agents can query third-party solutions.