#qwen

digestqwenopen-source
Qwen3.8-27B runs at 140 tok/s on one RTX 3090 with a CUDA megakernel
A Reddit author reports 140 tok/s for Qwen3.8-27B on a single RTX 3090, 1.4 to 1.9 times faster than llama.cpp. Code is on GitHub.

A Reddit author reports 140 tok/s for Qwen3.8-27B on a single RTX 3090, 1.4 to 1.9 times faster than llama.cpp. Code is on GitHub.