The efficient frontier of LLM inference
The article distinguishes techniques that trade latency for throughput or quality for speed from those that expand the whole serving frontier, then situates quantization, parallelism, kernels, speculative decoding, and disaggregation on that map.