The Illusion of Thinking
Apple argues that frontier large reasoning models improve on some reasoning benchmarks, but their accuracy collapses past certain problem complexities, their reasoning effort can fall as complexity rises, and they show limits in exact computation.