RULE OF THUMB
Parameters | Tier | Representative models | Decode speed | Notes |
|---|
~1B | Natural language processing (NLP) | Qwen 3.5 0.8B Qwen 3.5 2B (minimum Q8 quantization) | 100+ tok/s | On-thread language understanding; mostly summarization, named-entity recognition (NER), sentiment analysis. |
1B–9B | Q&A | Qwen 9B (E4B) Ling Tiny | 100+ tok/s | Simple user-facing question answering, usually paired with RAG. |
20B–35B (MoE) | Personal assistant (PA) | Gemma 4 26B A4B Qwen 3.6 35B A3B North Mini Code | 75+ tok/s | Personal assistant. The floor is reliable tool calling (e.g., passing the BFCL benchmark). |
25–30B (dense) | Coding core | Qwen 3.8 27B | 45+ tok/s | Interactive coding agents. Strong at core coding tasks, weaker at whole-codebase-level understanding. |
70B (dense) / 100–200B (MoE) | Code base | DeepSeek V4 Flash IQ2 Qwen 3.8 Next Flash IQ4 | 8+ tok/s | The "let it run overnight" tier: whole-codebase understanding, deep research and deep-analysis workloads. |
200T+ (sparse) | Human (Homo Sapiens) | — | 0.5 tok/s | I put humans here as a baseline to show how slow a real person (me) is at producing content. |
The numbers above are normalized to my system (1× RTX 3090 + 96GB RAM). Of course, some systems can run GLM 5 or MiMo, and even 122B+ models may not hit these tiers.
Terminology
T/S: tokens/s, tokens generated per second
MoE: Mixture-of-Experts model
DENSE: dense model (every parameter participates in every forward pass, unlike MoE)
Q8 / IQ2 / IQ4: quantization precision tiers (fewer bits = fewer resources, more quality loss)
BFCL: Berkeley Function-Calling Leaderboard, a tool-calling benchmark