Retranslated from Reddit's r/LocalAIServers, author u/UnDadFeated. Original: "Got Qwen3.8 27B running locally with DSpark on my RTX 5090 — ~208 tok/s": https://www.reddit.com/r/LocalAIServers/comments/1vv3mzj/ . All rights to the original author.

I took a weekend to get my local LLM deployment to a "no lag" state. Here's the full write-up. The headline in one sentence: Qwen3.8-27B (gittensor's NVFP4 build made for the 5090) with a DSpark speculative-decoding draft model, served by SGLang in Docker, hits about 208 tok/s (230+ peak) on a single RTX 5090 32GB.

Measured Speed

Hitting the local OpenAI-compatible API directly, 256 tokens per turn, thinking on, 5-run benchmark:

  • run1: 230.9 tok/s

  • run2: 148.3 tok/s (first run after a cold start, i.e. warmup)

  • run3: 232.2 tok/s

  • run4: 232.4 tok/s

  • run5: 229.1 tok/s

Average 208.2 tok/s. The first dip was just warmup; once settled it holds steady just above 220.

Why Docker (and the gotchas)

The DSpark draft model only loads on one specific runtime: a tagged private SGLang branch image lmsysorg/sglang:qwen38-27b. The author first tried the normal route — official PyPI SGLang crashes the moment it loads the draft model (weight shape mismatch, dies at startup). The branch source is not public, so there's no "pip install" fix; a Docker image was the only way.

CachyOS gotchas (can save you an evening)

If you're on CachyOS (or an Arch system with this kernel), the stock Docker daemon refuses to start until you write /etc/docker/daemon.json as:

{"iptables": false, "bridge": "none", "storage-driver": "vfs"}
  • iptables:false + bridge:none fixes the nf_tables ... chain PREROUTING error (Docker can't create the bridge network).

  • storage-driver:vfs fixes the overlay mount "no such device" error (this kernel has no overlay module compiled in). A bit slower, but you only run one container, so it doesn't matter.

  • With the bridge network off, the container must run with --network host instead of -p port mapping.

Deploy commands (safe to hand straight to an AI agent)

The author handed the whole setup to a coding agent (Hermes/Claude/Codex all work; no private info, just adjust paths):

docker run --gpus all --ipc=host --shm-size 32g --network host \
  -v <HF_HOME>:/root/.cache/huggingface \
  lmsysorg/sglang:qwen38-27b sglang serve \
    --model-path gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
    --speculative-algorithm DSPARK \
    --speculative-draft-model-path gittensor-model-hub/Qwen3.8-27B-DSpark-NVFP4 \
    --speculative-draft-model-quantization modelopt_fp4 \
    --speculative-dspark-block-size 7 \
    --trust-remote-code --tp-size 1 \
    --context-length 122880 --kv-cache-dtype fp8_e4m3 \
    --attention-backend flashinfer --chunked-prefill-size 1024 \
    --mamba-radix-cache-strategy extra_buffer_lazy \
    --mamba-ssm-dtype bfloat16 --max-mamba-cache-size 8 \
    --mm-feature-transport cpu --cuda-graph-max-bs-decode 1 \
    --mem-fraction-static 0.86 --max-running-requests 1 \
    --served-model-name qwen3.8 \
    --reasoning-parser qwen3 --tool-call-parser qwen3_coder \
    --host 0.0.0.0 --port 8000
  • Only this branch is DSpark-compatible; native PyPI SGLang 0.5.18 refuses to load with a weight-shape error — don't install the native version.

  • DeepSeek Harness (dsh) Web UI runs alongside on :3080, pointing at :8000/v1.

  • Measured operation: ai-sglang brings everything up in one shot (pull image, run container, start dsh, report tok/s), ready in ~58 seconds; ai-stop shuts it all down.

Appendix: Qwen3.8 tool-calling benchmark (5 configs, RTX 5090)

The author compared configs side by side with a 20-prompt tool-calling suite (pass = a valid tool_calls array with valid JSON arguments, not a natural-language answer; temp=0). The configs span three llama.cpp GGUF builds, vLLM, and SGLang+DSpark:

Config

Tool-call accuracy

Avg decode speed

llama.cpp A: utautako Q8attn+MTP (GGUF)

19/20 (95%)

~120 tok/s

llama.cpp B: unsloth UD-Q4_K_XL (GGUF, no MTP)

19/20 (95%)

~119 tok/s

llama.cpp C: Jackrong Q4_K_M+MTP (GGUF)

19/20 (95%)

~120 tok/s

vLLM: NVFP4 base

18/20 (90%)

~76 tok/s

SGLang + DSpark: NVFP4

19/20 (95%)

~211 tok/s

Conclusions:

  • All five configs land at 90%–95%, a narrow spread; every failure was the model answering in natural language instead of calling a tool — no JSON-format or wrong-function-name errors in any case.

  • The compounding-interest-calculator case was the usual miss (failed on 4 of 5 configs; only vLLM failed on something else).

  • The NVFP4 builds (vLLM and SGLang+DSpark) are not worse than GGUF on tool calling (vLLM is actually the lowest at 90%).

  • Speed differences are huge: DSpark (~211) is ~2.8× the vLLM base (~76) and ~1.8× the GGUF builds (~120), at on-par accuracy.

  • Note: single turn, 20 cases, temp 0 — directional results, not a rigorous evaluation; if tool calling is your core need, scale up the test volume.