A fresh wave of large language models is being pushed to its limits, and the results are drawing attention from mathematicians and AI watchers alike.
According to a widely shared post on Reddit's r/math, amplified on Hacker News, GPT-5.6 used a prompt to close a 30-year gap in convex optimization. The discussion followed what the post describes as OpenAI's "CDC proof announcement," and involved a variant referred to as GPT-5.6 Sol. The thread drew 172 points and 87 comments on Hacker News, a sign of how much interest the claim generated.
Separately, a blog post at charlesazam.com pitted Fable 5 against GPT-5.6 Sol on an NP-hard problem—a class of computational challenges known to be extremely hard to solve efficiently. The writeup specifically examines whether a "/goal" prompt helps the models perform better.
The head-to-head framing continues in the consumer press. Decrypt, surfaced via Google News, published a review titled "GPT-5.6 vs Fable 5" concluding that which model you should pick depends on several factors, rather than one model being universally best.
Not every headline is a victory lap. BankInfoSecurity reports that Kimi K3 highlights the limits of AI benchmark leaderboards—a reminder that ranking tables don't always capture how models behave on real tasks.
Taken together, the items sketch a fast-moving field where multiple frontier models—GPT-5.6, Fable 5, and Kimi K3 among them—are being stress-tested against one another on difficult problems, with prompting techniques like "/goal" emerging as a variable in performance.
Why it matters: as AI models increasingly tackle genuine research-grade problems, the way we measure and compare them is becoming as consequential as the models themselves.