GPT-5.6 Sol scores 7.8% on ARC-AGI-3, first AI to complete a benchmark game task
OpenAI's GPT-5.6 Sol model scored 7.8% on the ARC-AGI-3 benchmark and became the first verified frontier model ever to complete one of the games included in the test, according to results published by the ARC Prize foundation on 9 July. Previously, no model had surpassed roughly 2% on this benchmark, making the result a significant milestone. ARC-AGI-3 is considered one of the most demanding general intelligence tests for AI systems.
Comments
No comments yet
Comments
No comments yet โ be the first to weigh in ๐
No comments yet. Be the first!