I Trained GPT-2 From Scratch on a Single 3090 in Two Days: How Far Away Is "Alchemy" for Ordinary People?
Treat this experiment as a mirror: it reflects both the astonishing power consumer hardware gives an individual, and the gap that still can't be crossed from "it runs" to "it's useful".
1. Where This Started
After finishing Sebastian Raschka's Build a Large Language Model From Scratch, one thought wouldn't go away:
Without the cloud, without a cluster — with just a consumer GPU at home — can an ordinary person actually "brew" a decent foundation model from scratch?
The gut instinct says no. But Andrej Karpathy already proved it once with nanoChat — a 561M-parameter model trained in 4 hours on 8 H100s for $100. He knocked the barrier to the ground.
So let's take one step back: 163M parameters, GPT-2 small scale, a single 24GB gaming GPU. Any chance?
2. Nailing Down the Target First
What I'm reproducing is the 2019 OpenAI GPT-2 small, configuration copied strictly:
- 50,257-token vocabulary (GPT-2's BPE tokenizer)
- 1024-token context length
- 768 embedding dimension, 12 attention heads, 12 Transformer layers
- dropout 0.1
- One modern-practice tradeoff: the bias term in the QKV projections is removed
The final parameter count lands at 163M — not one more, not one less — right on the sweet spot between "fits in single-GPU memory" and "can learn something".
3. The Data Hurdle Is Brutal Even for the Hardware
The original GPT-2 used roughly 10B tokens of WebText, long since unmaintained. Hugging Face's ready-made FineWeb dataset now covers that scale, and the more refined FineWeb-Edu keeps only the subset of "educational" web pages.
The really nasty detail: pull down FineWeb and look at the text-length distribution, and a large share of documents already exceeds 1024 tokens on its own. Naively truncating per document and padding throws away 29% of the tokens.
Two options:
- Truncate and pad — simple, but wasteful;
- Concatenate all the text into one long stream with separators, then chunk it into 1024-token pieces — more complex, but closer to how GPT-2 was actually trained.
I chose the second.
4. 24GB of VRAM Is a Hard Line
The RTX 3090's memory caps the batch size, which caps throughput. I wrote a test script to measure at different precisions:
| Precision | Max batch | Throughput |
|---|---|---|
| FP32 | 5 | 12,599 tokens/s |
| TF32 (tensor cores on) | 5 | 15,402 tokens/s |
| AMP (automatic mixed precision) | 6 | ~20,000 tokens/s |
The last row is the sweet spot: throughput near 20K tokens/s, and the batch can push up one more notch. This is the life-or-death line for whether the whole project can finish over a weekend.
5. How Much to Train? The Chinchilla Answer
The original GPT-2 paper is vague; outside estimates put it at roughly 40 epochs over that 10B tokens — clearly out of reach for a single GPU.
DeepMind's Chinchilla scaling law gives a compute-optimal rule of thumb: training tokens ≈ 20× parameter count.
163M × 20 ≈ 3.26B tokens. FineWeb's 10B samples cover it with room to spare — one epoch is enough, no need to loop.
Do the math: 3.26B ÷ 20,000 ≈ 163,000 seconds ≈ 45 hours. One weekend, exactly.
6. The Infra Is More Grinding Than the Training Itself
Before actually starting, three things had to be in place:
Data preprocessing: extract 3.26B training tokens + 19.6M validation tokens from FineWeb, concatenate, chunk, and save in tensor format. The training set is about 13GB.
Checkpointing: this is a two-day run; if it crashes, you can't start over. Each checkpoint stores model weights, optimizer state, gradient-scaler state, and training progress — nearly 2GB each. One every 30 minutes. The disk is expensive, but an interrupted restart is more expensive.
Validation cadence: validate every 7,020 steps; a single validation pass eats 19.6M tokens and takes about 5 minutes. That's a reasonable share of total runtime — but no more frequent than this.
7. 48 Hours Later, the First Result
After roughly 48 hours of continuous training, validation loss settled at 3.94.
A generation test, with the prompt "Every effort moves you":
- Untrained model: "…ISIS Keectar handling holistic Supply query…" — pure gibberish;
- My trained model: "…towards a sustainable and holistic diet of water, protein, vitamins, and protein" — coherent semantics, but with the small-model habit of repeating (protein twice);
- Original GPT-2 weights: "…as far as the hand can go until the end of your turn unless something interrupts your control flow…" — clearly a tier higher in both logic and imagination.
The numbers sting more: on the same validation set, original GPT-2 small has loss 3.50 and perplexity ≈33.1; mine has loss 3.94 and perplexity 51.4. The gap is visible to the naked eye.
8. "Cleaner Data" Turns Out to Be Worse?
The first instinct: FineWeb-Edu — the "most educational" data — would that make a better model?
So another 48-hour run. The result is counterintuitive:
- On FineWeb-Edu's own validation set, loss drops to 3.693 — looks better;
- But back on the original FineWeb validation set, loss is 4.16 — significantly worse.
Then I fine-tuned both on Alpaca-format instruction data and had GPT-5.1 score them: the FineWeb version scored 16.14 after fine-tuning; the FineWeb-Edu version only 15.18.
"Cleaner" didn't turn directly into "stronger". The reasonable guess: the educational subset cut away a lot of real-world corpus diversity, and that locked in the generalization.
9. So What About Training Twice as Much?
On top of the FineWeb-Edu version, another 3.26B tokens (6.5B in total):
- Validation loss 3.693 → 3.661, less than 1% improvement;
- FineWeb validation set 4.16 → 4.13;
- Fine-tuning score 16.62 — slightly better, but marginal.
Twice the time, twice the electricity bill, no qualitative jump. Diminishing returns kick in very early at the 163M scale.
10. Where It Differs From the Original
Laying my model against the OpenAI weights point by point, the sources of the gap are basically clear:
- Data volume and epochs: the original had ~400B tokens (dozens of epochs); mine had 3.2–6.5B (1–2 epochs). Deeply "grinding" the data is beyond a single GPU;
- Architecture details: the original kept the QKV bias and weight tying; I dropped both per modern practice. Those "retro" choices may have had special value in the training dynamics of the day;
- Training tricks: gradient clipping and more elaborate learning-rate schedules (e.g. cosine annealing) were not fully used here;
- Batch size: the original's global batch was 512; physically I could only reach 6 on one GPU. Bigger batches give a steadier gradient direction;
- Precision: the original was likely FP32 throughout; I used mixed precision for speed, conceding a step on numerical stability.
11. What This Mirror Shows
The answer to the original question is "yes": a single 3090, two days, a few dozen dollars of electricity, and anyone can train a GPT-2-scale model with basic language ability from zero. Five years ago that was unimaginable.
But the boundaries are equally clear:
- Chinchilla-optimal is an efficient starting point, not a finish line;
- Getting close to an industrial-grade model takes exponentially more data, compute, and engineering polish;
- An individual developer can "learn the alchemy", but can't easily "brew industrial-grade elixir".
That final model will confidently get the author of Pride and Prejudice wrong. It is a mirror: on one side it shows the ceiling of capability that the open-source ecosystem and consumer hardware have already pushed high for individuals; on the other, it shows the chasm between "usable" and "excellent" that still demands massive investment to cross.
For an indie developer, this conclusion is actually not a loss — you don't need to train your own GPT-5; you just need to know where the chasm is, which steps you can stand on someone else's shoulders, and put your own energy into the product and the gameplay. Treat the "brewing models" story as noise; if you actually ship a product, you still need an existing strong model.
(Views in this article are the author's own.)