📎 Webclip
The Impossible Backhand
The post argues that AI tends toward average outputs and runs into a structural ceiling in quality. It uses tennis, legal research, medicine, and film-making to show that domain experts can catch errors that look plausible to everyone else, while non-experts are more likely to trust them.
It frames the best use of AI as a centaur model, where human judgment remains central. The post also warns that early reliance on AI can erode the skills needed to detect mistakes, so expertise stays valuable because it is what lets people see when the model is wrong.
Reading notes#
- A ByteDance Seedance 2.0 tennis clip looked real, but a tennis expert immediately noticed a backhand that would not exist in real play.
- AI can reach something like the 95th or 98th percentile of plausibility, but still miss what experts can spot right away.
- The post says the move toward the mean is structural, not a temporary limitation.
- Next-token prediction favors statistically probable continuations, which pushes output toward average rather than distinctive quality.
- RLHF reinforces typicality, because human raters prefer familiar-sounding answers.
- Model collapse is described as AI systems training on AI-generated content and losing the tails of the distribution.
- The cost of further improvement rises sharply, with compute needs scaling very steeply for smaller gains.
- Humanity’s Last Exam is used as a benchmark showing a wide gap between AI scores and human expert scores.
- AI systems are described as overconfident, especially in legal and medical settings where hallucinations can cause harm.
- Specialized knowledge is learned poorly because it appears less often in pretraining data.
- The centaur model is presented as the stronger framework, with humans and AI outperforming either alone when the human can judge the model.
- In the BCG study, AI helped on tasks inside its frontier, but made performance worse on tasks outside it when people trusted it blindly.
- The post argues that junior workers who rely on AI too early may fail to build the judgment needed to catch errors.
- The main conclusion is that the ability to identify impossible outputs remains valuable and may become more valuable over time.
