arxivcs.AI2026-07-07
Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1
Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on…