Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning.
The creators of the ARC-AGI benchmark say Anthropic's Claude Opus 5 owes its massive lead on ARC-AGI-3 to genuinely better reasoning.
The model scored 30.2 percent on ARC-AGI-3, making it the new leader. The previous record was 7.8 percent, set by Op... [3300 chars]

