Published: 2026-09-06. Environment: RTX 4080 modified to 32GB, llama.cpp build 10631, Qwen3.8-27B (Uncensored fine-tune, Q3_K_P weights).

Running Qwen3.8-27B on local consumer GPUs is the most worthwhile thing to do in 2026 — 27B dense, 262K native context, text/image/video multimodal, a built-in speculative-decoding head, Apache 2.0, fully open source. But the clickbait "157~167 t/s" numbers in tutorials everywhere easily set unrealistic expectations for your own machine.

This article deals in facts only. Every number comes from real runs on my RTX 4080 modified to 32G; community data is labeled separately. Configure with the launch parameters at the end and you'll reproduce the same results.

TL;DR

  • The "32G" of a modded 4080 brings VRAM capacity, not speed — bandwidth is still 256-bit / 716 GB/s, identical to the original.

  • 40 t/s standard decoding is not slow; it's the normal level for a 4080 running a 27B Q4. Most of the web lands at 40~55 t/s.

  • Those 100+ t/s figures online almost all come from MTP speculative decoding (Multi-Token Prediction). My measured 40.1 → 87.3 t/s — more than doubled, quality unchanged.

  • The hidden bonus of this model is a tiny KV cache (~22 KB/token measured): only 16 of the 64 layers are conventional full attention; the other 48 are Gated DeltaNet linear attention.

  • The biggest pitfall isn't in the model, it's the Windows TDR watchdog: full-context long generations push a single CUDA kernel past the default 2-second TdrDelay, the driver gets force-reset, and the process crashes with 0xC0000409.


1. First, what a modded 4080 32G actually is

The 4080 32G mod is a mature product of the recent modding wave: essentially the original 16G VRAM doubled. Only capacity is doubled — bus width, bandwidth, and compute are all unchanged:

Spec

4080 original

4080 modded 32G

RTX 3090

RTX 4090

VRAM

16GB

32GB

24GB

24GB

Bus width

256-bit

256-bit

384-bit

384-bit

Memory bandwidth

~716 GB/s

~716 GB/s

~936 GB/s

~1008 GB/s

CUDA cores

9728

9728

10496

16384

LLM token-generation speed is essentially bandwidth-bound, so this machine's positioning is clear:

  • Inference (running 27B~30B-class models locally) → the 4080 32G has great value: 8G more VRAM than a 3090/4090, and 30B-class inference throughput improves roughly 20~30%.

  • Training / high throughput → pick the 4090; bandwidth and compute differ by an order of magnitude.

In one line: 32G lets you fit bigger models and longer contexts, not run faster.

2. The model bonus: why 27B can open a 262K context

Qwen3.8-27B is a dense multimodal model (27.78B parameters, not MoE), with a new Gated DeltaNet hybrid attention architecture:

  • 64 layers in 16 groups; each group has 3 Gated DeltaNet (linear attention) layers + 1 Gated Attention (full attention) layer

  • i.e. 48 linear-attention layers + 16 conventional full-attention layers

  • Native context 262,144 tokens, hidden dim 5120

  • Built-in MTP head (speculative decoding); text/image/video multimodal

The key is the KV cache: since 48 layers don't keep full KV, the KV is only about 1/4 of a conventional dense 27B. My measured value: ~22 KB/token:

Context

KV cache usage

32K

~2 GB

131K

~8-10 GB (~2.9 GB measured after q4_0 quantization)

262K

~16-20 GB (quantization required; ~5.8 GB at q4_0)

That's the root reason it can keep a full long-context KV resident on a 32G card. Swap in any conventional-architecture 27B and opening a 200K context is a disaster for 32G.

3. Demystifying: where the online 167 t/s comes from

First, real speed tests on my machine (4080 32G, Q3_K_P weights):

Standard decode      : 40.1 tokens/s

MTP instantaneous peak: 87.3 tokens/s

MTP long generation   : ~60 tokens/s (speed decays as KV grows)

Average incl. prefill : ~44 tokens/s

Draft acceptance rate : 68.6% (pos0/1/2 = 83/69/49)

GPU VRAM              : ~19.7 GB / 32 GB

Then community measurements on other cards:

GPU

Standard decode

With MTP

Uplift

RTX 5090 32G

~74 t/s

~133 t/s

+80%

RTX 5080 16G

54.3 t/s

93.9 t/s

+73%

RTX 3090 24G

~41 t/s

~55 t/s

+35%

My 4080 32G

40.1 t/s

87.3 t/s

+118%

The conclusion is blunt:

  1. 40 t/s standard decoding on a 4080 with Qwen3.8-27B Q4 is normal-to-above — not slow. The "157~167 t/s" online almost all comes from MTP speculative decoding, and mostly peaks (short generations, ideal conditions, code-style tasks).

  2. The 40 → 87 doubling is free money: the model ships with an MTP head, one llama.cpp flag enables it, zero quality loss.

MTP isn't complex: the model predicts several future tokens in one forward pass, and the main model verifies them in parallel. Accepted tokens skip per-token decoding, so throughput roughly doubles. Note that --mtp is an invalid flag; the correct one is --spec-type draft-mtp.

The MTP speedup strongly depends on task type: fixed formats (JSON/HTML/code completion) get high acceptance and the biggest gains; open-ended creative writing or complex reasoning gets low acceptance and limited gains. Greedy decoding at temp=0 gives the highest acceptance; raising temp drops it.

4. Choosing quantization: fix the VRAM budget first, then work backwards

GGUF quantization sizes (Unsloth repo):

Quant

File size

Positioning

BF16

54.7 GB

Original reference; needs 48G+

Q8_0

29.0 GB

Near-original precision

Q6_K

22.9 GB

High quality

Q5_K_M

19.8 GB

Quality/resource compromise

UD-Q4_K_XL

17.9 GB

Competitive 4-bit

Q4_K_M

17.1 GB

Most general 4-bit

UD-Q3_K_XL

13.4 GB

For 16G

UD-IQ2_XXS

9.0 GB

Extreme low VRAM

Selection principle (work backwards from VRAM; don't open full 262K from the start):

  • 16GB: Q3 / IQ3 tier; compromise is inevitable. Community consensus: keep Q3 fully resident in VRAM (fast, stable) rather than force-fit Q4 + CPU offload (per-token stalls, slower day to day).

  • 24GB: sweet spot — UD-Q4_K_XL or Q4_K_M.

  • 32GB: you can go Q5/Q6 for quality; Q4 is the quantization safety line (there are reports of Q5+MTP triggering the driver watchdog Xid 8 on Blackwell cards).

  • 48G+: Q6/Q8/BF16, chasing long context + multimodal.

Model file size ≠ actual runtime VRAM: you also add KV cache, runtime buffers, CUDA backend overhead, and the mmproj vision projection.

5. Launch parameters you can copy (32G card)

llama-server \
  --model Qwen3.8-27B-...-Q3_K_P.gguf \
  --mmproj mmproj-...-BF16.gguf \      # required for multimodal vision
  --ctx-size 262144 \                  # total; with 2 slots that's 131072 each
  --parallel 2 \                       # two concurrent slots
  --n-gpu-layers 999 \                 # all layers on GPU
  --flash-attn on \
  --cache-type-k q4_0 --cache-type-v q4_0 \   # KV quantization, saves 75% KV VRAM
  --spec-type draft-mtp \              # MTP speculative decoding (the core speedup)
  --spec-draft-n-max 3 \               # draft token count
  --reasoning-effort low               # thinking tier; only xhigh/medium/low are valid

Measured effect of this config (2026-09-06, dual-slot version):

n_slots = 2, n_ctx_slot = 131072

VRAM     : 26.1 GB / 32 GB (~6 GB headroom)

Single-slot speed: ~53 t/s (short context); dual concurrent aggregate >100 t/s

Draft acceptance: 58%–67% (per slot, independent)

A few lessons learned:

  • --spec-draft-n-max use 2~3: keep 2 on 16G cards, try 4-6 on 24G+, but diminishing returns — the verification stage is the compute bottleneck. At a 68% acceptance rate you're already near optimal.

  • --spec-draft-p-min 0.75 is the community-recommended sweet spot (it's in the 5080 config that went 54→94 t/s).

  • KV quantization q4_0 is a prerequisite for long context: it squeezes 100K of KV from ~25GB to ~6GB, keeping KV in VRAM. KV spilling into CPU RAM is the root cause of speed collapse (measured ~12 t/s on an old 16G card).

  • --no-mmap helps in extreme scenarios; generally not needed on 32G cards.

6. Tunings I measured that don't help: save your time

These are the items I tested with a "useless" verdict — don't waste time on them:

Trial

Result

Verdict

--batch-size 4096 --ubatch-size 2048

40.8 t/s vs 40.1

No decode improvement

ubatch 512→2048 (prefill-only A/B, 35k tokens)

1457/1503 vs 1515/1424 t/s

Within ±3% noise; no gain

KV cache switched to q8_0

39.3 t/s, and ~5GB more usage

Actually slightly slower

Conclusion: the bottleneck is decode bandwidth; batch/KV-type tuning is meaningless — the only effective speed lever is MTP.

7. The Windows-specific big pitfall: TDR crash (0xC0000409)

This is the most hidden pitfall of 27B long context on Windows; I fell into it 7 times before locating it.

Symptom: mid-way through a full-context long generation, the screen flickers black and llama-server exits. Tail of the log:

E CUDA error: unknown error

E   cudaStreamSynchronize(0x2)  failed

E ggml_backend_cuda_buffer_set_tensor  at ggml-cuda.cu:791

The Event Viewer shows nvlddmkm Event 14/153 (GPU unresponsive). Process exit code 0xC0000409 — this isn't a real stack overflow or OOM; it's Windows __fastfail, commonly from GGML_ASSERT / uncaught exceptions.

Root-cause chain (closed):

Full-context long generation → single ultra-long CUDA kernel exceeds WDDM TdrDelay default of 2 seconds

→ driver force-reset (TDR) → CUDA context corrupted

→ cudaStreamSynchronize returns unknown error → process fail-fast (0xC0000409)

The more desktop graphics processes (browsers, chat apps, wallpaper engines, NVIDIA Overlay) competing, the easier it is to trigger.

Fix: raise the registry TdrDelay from the default 2s to 20s (HKLM\SYSTEM\CurrentControlSet\Control\GraphicsDrivers, TdrDelay=20 / TdrDdiDelay=5) — a system reboot is required for it to take effect. Also, dropping --ctx-size from 262144 to 131072 significantly reduces how often ultra-long kernels occur.

The "128K context sweet spot" is well-founded — three independent sources converge:

  1. The official Qwen model card explicitly recommends "keep at least 128K context to maintain thinking ability" — 128K is the official-recommended floor;

  2. Long context is a known high-incidence scenario for Windows TDR;

  3. 128K vs 256K halves both the KV and the per-token attention load (measured 22 KB/token: 256K ≈ 5.5 GB → 128K ≈ 2.8 GB), while generation speed decays roughly linearly with context (87 t/s peak → 51-60 t/s at long context).

8. Other pitfalls, in brief

  • The mmproj vision projection must match the base model by the same prefix — never take the first one in alphabetical order, or you'll pair gemma's projection with a Qwen base and fail at startup.

  • reasoning-effort only accepts xhigh / medium / low: passing high or auto gets rejected by the template with a 500. Qwen3.8's thinking tier can only be set via launch parameters; no request-level switching (same for DeepSeek-R1).

  • You can't force Chinese output server-side: llama.cpp has no global language parameter, and a system message doesn't guarantee it 100% either. The most reliable method is to ask in Chinese (Chinese in → Chinese out), with the system message as backup.

  • Start-Process -ArgumentList breaks paths containing spaces: any Program Files in the path causes invalid argument. Writing the arguments as a PowerShell array into a temp .ps1 and running with -File is the most stable route.

  • Launching from Chinese paths in the background garbles them: PowerShell parses a UTF-8 BOM-less script as ANSI (GBK), scrambling Chinese paths. Work around it with an ASCII junction (mklink /J), or keep the script UTF-8 with BOM.

  • Qwen's Jinja template requires the system message at position 0: clients like opencode that send multiple system messages or a non-first system message get System message must be at the beginning. The fix is to connect via the Anthropic protocol instead (system in the top-level field, which sidesteps it naturally).

9. Independent conclusions

  1. 40 t/s is not slow — don't let the clickbait bully you. The web-wide norm for a 4080 running Qwen3.8-27B Q4 standard decode is 40-55 t/s; with MTP you hit an 87 t/s peak, already beyond 5090 standard decoding.

  2. MTP is the only parameter worth tuning. One flag, doubled speed, zero quality loss. I tested batch, KV type, ubatch — no improvement.

  3. The value of 32G is VRAM, not bandwidth. Use it for heavier quantizations and longer contexts, not to chase tokens/s.

  4. If you really want faster, switch to MoE: Qwen3-30B-A3B (30B parameters, 3B active) measured ~196 t/s on a 4090 — the speed king — and MoE KV is smaller, so long-context decay is smaller too.

  5. Windows long-context users: change TdrDelay first. Otherwise 262K full context + MTP long generation will eventually hand you a 0xC0000409.

One last cold-water splash: the experience ceiling of a local 27B is set by quantization + context + GPU bandwidth together. Understand that "fast online = MTP + peak + specific tasks", manage your expectations, and this 4080 32G is one of the best-value consumer local-inference options today.


Data sources: this project's tuning notes and fast-MTP troubleshooting logs, plus community research (Chenjianyun, DataLearner, Yunstack community, the official Qwen model card, OnlyTerp/windows-is-fine-for-llms, Microsoft Learn TDR Registry Keys).