· 15 min read
SWE-Explore: The Benchmark That Finally Asks — Did Your Coding Agent Read the Right Code?
SWE-Explore isolates repository exploration from patch generation, revealing that coding agents find the right files ~65% of the time but recall only ~15-19% of the lines that actually matter — and that context efficiency predicts downstream resolve rate with Pearson r = 0.950.