Published: 2026-09-06. Environment: RTX 4080 modified to 32GB, llama.cpp build 10631, Qwen3.8-27B (Uncensored fine-tune, Q3_K_P weights).
Running Qwen3.8-27B on local consumer GPUs is the most worthwhile thing to do in 2026 — 27B dense, 262K native context, text/image/video multimodal, a built-in speculative-decoding head, Apache 2.0, fully open source. But the clickbait "157~167 t/s" numbers in tutorials everywhere easily set unrealistic expectations for your own machine.
This article deals in facts only. Every number comes from real runs on my RTX 4080 modified to 32G; community data is labeled separately. Configure with the launch parameters at the end and you'll reproduce the same results.
TL;DR
The "32G" of a modded 4080 brings VRAM capacity, not speed — bandwidth is still 256-bit / 716 GB/s, identical to the original.
40 t/s standard decoding is not slow; it's the normal level for a 4080 running a 27B Q4. Most of the web lands at 40~55 t/s.
Those 100+ t/s figures online almost all come from MTP speculative decoding (Multi-Token Prediction). My measured 40.1 → 87.3 t/s — more than doubled, quality unchanged.
The hidden bonus of this model is a tiny KV cache (~22 KB/token measured): only 16 of the 64 layers are conventional full attention; the other 48 are Gated DeltaNet linear attention.
The biggest pitfall isn't in the model, it's the Windows TDR watchdog: full-context long generations push a single CUDA kernel past the default 2-second TdrDelay, the driver gets force-reset, and the process crashes with 0xC0000409.
1. First, what a modded 4080 32G actually is
The 4080 32G mod is a mature product of the recent modding wave: essentially the original 16G VRAM doubled. Only capacity is doubled — bus width, bandwidth, and compute are all unchanged:
LLM token-generation speed is essentially bandwidth-bound, so this machine's positioning is clear:
Inference (running 27B~30B-class models locally) → the 4080 32G has great value: 8G more VRAM than a 3090/4090, and 30B-class inference throughput improves roughly 20~30%.
Training / high throughput → pick the 4090; bandwidth and compute differ by an order of magnitude.
In one line: 32G lets you fit bigger models and longer contexts, not run faster.
2. The model bonus: why 27B can open a 262K context
Qwen3.8-27B is a dense multimodal model (27.78B parameters, not MoE), with a new Gated DeltaNet hybrid attention architecture:
64 layers in 16 groups; each group has 3 Gated DeltaNet (linear attention) layers + 1 Gated Attention (full attention) layer
i.e. 48 linear-attention layers + 16 conventional full-attention layers
Native context 262,144 tokens, hidden dim 5120
Built-in MTP head (speculative decoding); text/image/video multimodal
The key is the KV cache: since 48 layers don't keep full KV, the KV is only about 1/4 of a conventional dense 27B. My measured value: ~22 KB/token:
That's the root reason it can keep a full long-context KV resident on a 32G card. Swap in any conventional-architecture 27B and opening a 200K context is a disaster for 32G.
3. Demystifying: where the online 167 t/s comes from
First, real speed tests on my machine (4080 32G, Q3_K_P weights):
Standard decode : 40.1 tokens/s
MTP instantaneous peak: 87.3 tokens/s
MTP long generation : ~60 tokens/s (speed decays as KV grows)
Average incl. prefill : ~44 tokens/s
Draft acceptance rate : 68.6% (pos0/1/2 = 83/69/49)
GPU VRAM : ~19.7 GB / 32 GBThen community measurements on other cards:
The conclusion is blunt:
40 t/s standard decoding on a 4080 with Qwen3.8-27B Q4 is normal-to-above — not slow. The "157~167 t/s" online almost all comes from MTP speculative decoding, and mostly peaks (short generations, ideal conditions, code-style tasks).
The 40 → 87 doubling is free money: the model ships with an MTP head, one llama.cpp flag enables it, zero quality loss.
MTP isn't complex: the model predicts several future tokens in one forward pass, and the main model verifies them in parallel. Accepted tokens skip per-token decoding, so throughput roughly doubles. Note that --mtp is an invalid flag; the correct one is --spec-type draft-mtp.
The MTP speedup strongly depends on task type: fixed formats (JSON/HTML/code completion) get high acceptance and the biggest gains; open-ended creative writing or complex reasoning gets low acceptance and limited gains. Greedy decoding at temp=0 gives the highest acceptance; raising temp drops it.
4. Choosing quantization: fix the VRAM budget first, then work backwards
GGUF quantization sizes (Unsloth repo):
Selection principle (work backwards from VRAM; don't open full 262K from the start):
16GB: Q3 / IQ3 tier; compromise is inevitable. Community consensus: keep Q3 fully resident in VRAM (fast, stable) rather than force-fit Q4 + CPU offload (per-token stalls, slower day to day).
24GB: sweet spot — UD-Q4_K_XL or Q4_K_M.
32GB: you can go Q5/Q6 for quality; Q4 is the quantization safety line (there are reports of Q5+MTP triggering the driver watchdog Xid 8 on Blackwell cards).
48G+: Q6/Q8/BF16, chasing long context + multimodal.
Model file size ≠ actual runtime VRAM: you also add KV cache, runtime buffers, CUDA backend overhead, and the mmproj vision projection.
5. Launch parameters you can copy (32G card)
llama-server \
--model Qwen3.8-27B-...-Q3_K_P.gguf \
--mmproj mmproj-...-BF16.gguf \ # required for multimodal vision
--ctx-size 262144 \ # total; with 2 slots that's 131072 each
--parallel 2 \ # two concurrent slots
--n-gpu-layers 999 \ # all layers on GPU
--flash-attn on \
--cache-type-k q4_0 --cache-type-v q4_0 \ # KV quantization, saves 75% KV VRAM
--spec-type draft-mtp \ # MTP speculative decoding (the core speedup)
--spec-draft-n-max 3 \ # draft token count
--reasoning-effort low # thinking tier; only xhigh/medium/low are validMeasured effect of this config (2026-09-06, dual-slot version):
n_slots = 2, n_ctx_slot = 131072
VRAM : 26.1 GB / 32 GB (~6 GB headroom)
Single-slot speed: ~53 t/s (short context); dual concurrent aggregate >100 t/s
Draft acceptance: 58%–67% (per slot, independent)A few lessons learned:
--spec-draft-n-maxuse 2~3: keep 2 on 16G cards, try 4-6 on 24G+, but diminishing returns — the verification stage is the compute bottleneck. At a 68% acceptance rate you're already near optimal.--spec-draft-p-min 0.75is the community-recommended sweet spot (it's in the 5080 config that went 54→94 t/s).KV quantization q4_0 is a prerequisite for long context: it squeezes 100K of KV from ~25GB to ~6GB, keeping KV in VRAM. KV spilling into CPU RAM is the root cause of speed collapse (measured ~12 t/s on an old 16G card).
--no-mmaphelps in extreme scenarios; generally not needed on 32G cards.
6. Tunings I measured that don't help: save your time
These are the items I tested with a "useless" verdict — don't waste time on them:
Conclusion: the bottleneck is decode bandwidth; batch/KV-type tuning is meaningless — the only effective speed lever is MTP.
7. The Windows-specific big pitfall: TDR crash (0xC0000409)
This is the most hidden pitfall of 27B long context on Windows; I fell into it 7 times before locating it.
Symptom: mid-way through a full-context long generation, the screen flickers black and llama-server exits. Tail of the log:
E CUDA error: unknown error
E cudaStreamSynchronize(0x2) failed
E ggml_backend_cuda_buffer_set_tensor at ggml-cuda.cu:791The Event Viewer shows nvlddmkm Event 14/153 (GPU unresponsive). Process exit code 0xC0000409 — this isn't a real stack overflow or OOM; it's Windows __fastfail, commonly from GGML_ASSERT / uncaught exceptions.
Root-cause chain (closed):
Full-context long generation → single ultra-long CUDA kernel exceeds WDDM TdrDelay default of 2 seconds
→ driver force-reset (TDR) → CUDA context corrupted
→ cudaStreamSynchronize returns unknown error → process fail-fast (0xC0000409)The more desktop graphics processes (browsers, chat apps, wallpaper engines, NVIDIA Overlay) competing, the easier it is to trigger.
Fix: raise the registry TdrDelay from the default 2s to 20s (HKLM\SYSTEM\CurrentControlSet\Control\GraphicsDrivers, TdrDelay=20 / TdrDdiDelay=5) — a system reboot is required for it to take effect. Also, dropping --ctx-size from 262144 to 131072 significantly reduces how often ultra-long kernels occur.
The "128K context sweet spot" is well-founded — three independent sources converge:
The official Qwen model card explicitly recommends "keep at least 128K context to maintain thinking ability" — 128K is the official-recommended floor;
Long context is a known high-incidence scenario for Windows TDR;
128K vs 256K halves both the KV and the per-token attention load (measured 22 KB/token: 256K ≈ 5.5 GB → 128K ≈ 2.8 GB), while generation speed decays roughly linearly with context (87 t/s peak → 51-60 t/s at long context).
8. Other pitfalls, in brief
The mmproj vision projection must match the base model by the same prefix — never take the first one in alphabetical order, or you'll pair gemma's projection with a Qwen base and fail at startup.
reasoning-effortonly acceptsxhigh / medium / low: passinghighorautogets rejected by the template with a 500. Qwen3.8's thinking tier can only be set via launch parameters; no request-level switching (same for DeepSeek-R1).You can't force Chinese output server-side: llama.cpp has no global language parameter, and a system message doesn't guarantee it 100% either. The most reliable method is to ask in Chinese (Chinese in → Chinese out), with the system message as backup.Start-Process -ArgumentListbreaks paths containing spaces: anyProgram Filesin the path causesinvalid argument. Writing the arguments as a PowerShell array into a temp .ps1 and running with-Fileis the most stable route.Launching from Chinese paths in the background garbles them: PowerShell parses a UTF-8 BOM-less script as ANSI (GBK), scrambling Chinese paths. Work around it with an ASCII junction (mklink /J), or keep the script UTF-8 with BOM.Qwen's Jinja template requires the system message at position 0: clients like opencode that send multiple system messages or a non-first system message getSystem message must be at the beginning.The fix is to connect via the Anthropic protocol instead (system in the top-level field, which sidesteps it naturally).
9. Independent conclusions
40 t/s is not slow — don't let the clickbait bully you. The web-wide norm for a 4080 running Qwen3.8-27B Q4 standard decode is 40-55 t/s; with MTP you hit an 87 t/s peak, already beyond 5090 standard decoding.MTP is the only parameter worth tuning. One flag, doubled speed, zero quality loss. I tested batch, KV type, ubatch — no improvement.The value of 32G is VRAM, not bandwidth. Use it for heavier quantizations and longer contexts, not to chase tokens/s.If you really want faster, switch to MoE: Qwen3-30B-A3B (30B parameters, 3B active) measured ~196 t/s on a 4090 — the speed king — and MoE KV is smaller, so long-context decay is smaller too.Windows long-context users: change TdrDelay first. Otherwise 262K full context + MTP long generation will eventually hand you a 0xC0000409.
One last cold-water splash: the experience ceiling of a local 27B is set by quantization + context + GPU bandwidth together. Understand that "fast online = MTP + peak + specific tasks", manage your expectations, and this 4080 32G is one of the best-value consumer local-inference options today.
Data sources: this project's tuning notes and fast-MTP troubleshooting logs, plus community research (Chenjianyun, DataLearner, Yunstack community, the official Qwen model card, OnlyTerp/windows-is-fine-for-llms, Microsoft Learn TDR Registry Keys).