↓ Ir para o conteúdo principal

← todas as notas

📎 Webclip

GPT-5: Overdue, overhyped and underwhelming. And that’s not the worst of it.

Marcus says GPT-5 landed late and failed to meet the expectations OpenAI itself had built. The launch drew disappointment, reports of errors and hallucinations, and a fast drop in confidence that OpenAI still leads the field.

He argues the model remains only an incremental advance, still weak on chess, vision, reasoning, reading, summarization, and generalization outside training distributions. He uses the release and a new Arizona State University study to support his long-standing view that scaling alone will not reach AGI, and that systems need explicit world models and a different approach.

Reading notes
#

  • GPT-5’s debut is described as late, overhyped, and widely disappointing.
  • OpenAI’s launch messaging is presented as far stronger than the model’s actual performance.
  • Users reported errors, hallucinations, bad routing, weak benchmarks, and frustration with the removal of older models.
  • The reaction is framed as a drop in OpenAI’s credibility and in confidence that the company still has a clear technical lead.
  • The post says GPT-5 is not terrible, but it is not a radical advance over prior models.
  • Marcus says GPT-5 still fails on chess, visual comprehension, image generation, reading, and summarization.
  • A new Arizona State University study is presented as confirming that chain-of-thought reasoning breaks down outside training distributions.
  • The core claim is that LLMs do not generalize broadly enough, and that this is a principled limitation, not a temporary bug.
  • Marcus argues that pure scaling has had more funding and benefit of the doubt than any other hypothesis, but still has not produced AGI.
  • He calls for neurosymbolic AI and explicit world models instead of scaling alone.