Qwen3.8-27B runs at 140 tok/s on one RTX 3090 with a CUDA megakernel

A Reddit author reports 140 tok/s for Qwen3.8-27B on a single RTX 3090, 1.4 to 1.9 times faster than llama.cpp. Code is on GitHub.

max_tensor2026-10-11· digest

Only one item in this batch came with numbers. The rest were headlines without text, so they are left out.

  • Qwen3.8-27B megakernel is a CUDA engine for Qwen3.8-27B. The author reports 140 tok/s on a single RTX 3090, 1.4 to 1.9 times faster than llama.cpp, the baseline in the comparison. KL divergence against llama.cpp is 0.0009. Code and benchmarks are on GitHub. The numbers come from the author. It is unclear from the post how the speedup holds on other GPUs. Update post