AI evaluationreal-world benchmarkspublished Aug 13, 2026 · arXiv:2608.11341
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
program repair agentsvalidation evidencepublished Aug 3, 2026 · arXiv:2607.28871
Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?
deletion avoidanceautomated software engineeringpublished Aug 3, 2026 · arXiv:2607.28887
To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
coding agentsexecutable world modelspublished Jul 20, 2026 · arXiv:2607.15439
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?