RULE OF THUMB

Parameters

Tier

Representative models

Decode speed

Notes

~1B

Natural language processing (NLP)

Qwen 3.5 0.8B
Qwen 3.5 2B (minimum Q8 quantization)

100+ tok/s

On-thread language understanding; mostly summarization, named-entity recognition (NER), sentiment analysis.

1B–9B

Q&A

Qwen 9B (E4B)
Ling Tiny

100+ tok/s

Simple user-facing question answering, usually paired with RAG.

20B–35B (MoE)

Personal assistant (PA)

Gemma 4 26B A4B
Qwen 3.6 35B A3B
North Mini Code

75+ tok/s

Personal assistant. The floor is reliable tool calling (e.g., passing the BFCL benchmark).

25–30B (dense)

Coding core

Qwen 3.8 27B

45+ tok/s

Interactive coding agents. Strong at core coding tasks, weaker at whole-codebase-level understanding.

70B (dense) / 100–200B (MoE)

Code base

DeepSeek V4 Flash IQ2
Qwen 3.8 Next Flash IQ4

8+ tok/s

The "let it run overnight" tier: whole-codebase understanding, deep research and deep-analysis workloads.

200T+ (sparse)

Human (Homo Sapiens)

—

0.5 tok/s

I put humans here as a baseline to show how slow a real person (me) is at producing content.

The numbers above are normalized to my system (1× RTX 3090 + 96GB RAM). Of course, some systems can run GLM 5 or MiMo, and even 122B+ models may not hit these tiers.

Terminology

  • T/S: tokens/s, tokens generated per second

  • MoE: Mixture-of-Experts model

  • DENSE: dense model (every parameter participates in every forward pass, unlike MoE)

  • Q8 / IQ2 / IQ4: quantization precision tiers (fewer bits = fewer resources, more quality loss)

  • BFCL: Berkeley Function-Calling Leaderboard, a tool-calling benchmark