Reasoning traces are far more redundant than they look: keeping just 10% of the KV cache can preserve essentially full reasoning accuracy and even improve it by filtering noise.
Autoregressive vision models can rival, and even surpass, diffusion models at multimodal image generation, while training on 10× less data and a fraction of the compute.
DPO may be doing something much simpler than gradient descent suggests: learning a low-dimensional “toxic subspace” that can be removed directly with a single closed-form projection.