Skylion007/openwebtext
Viewer • Updated • 8.01M • 54.3k • 531
Both models are Decoder-only Transformers (GPT-style) implemented in pure PyTorch.
| Hyperparameters | Main Model (Target) | Draft Model (Small) | Draft Model (Medium) |
|---|---|---|---|
| Parameters | ~150M | ~30M | ~70M |
| Layers | 12 | 2 | 6 |
| Heads | 12 | 4 | 8 |
| Embedding Dim | 768 | 256 | 512 |
| Context Length | 1024 | 1024 | 1024 |
| Vocab Size | 50304 | 50304 | 50304 |
| Droupout | 0.1 | 0.1 | 0.1 |
| Dataset | OpenWebText (Sample) | OpenWebText (Sample) | OpenWebText (Sample) |
Token + Position Embeddings
(vocab_size, n_embd)Decoder Blocks Each block consists of:
Final RMSNorm + Language Modeling Head
n_embd → vocab_size to produce logitsThe main model was trained in 4 distinct phases to hande hardware constraints and optimize convergence
| Phases | Focus | Max Iterations | Warmup Steps | Eval Iterations | Eval Interval | Accumulation Steps | Learning Rate | Weight Decay | Batch Size |
|---|---|---|---|---|---|---|---|---|---|
| Phase 1 | Initial Warmup | 10_000 | 200 | 20 | 500 | 32 | 3e-4 | 0.1 | 8 |
| Phase 2 | Main Pre-Training | 30_000 | 0 | 20 | 1_000 | 32 | 3e-4 | 0.15 | 8 |
| Phase 3 | Large-Batch Scaling | 50_000 | 1_000 | 10 | 2_000 | 16 | 6e-5 | 0.1 | 16 |
| Phase 4 | Convergence | 40_000 | 500 | 10 | 2_000 | 16 | 2e-5 | 0.08 | 16 |
The draft model (small) was trained only once
max_iters = 40_000
warmup_steps = 1_000
eval_iter = 20
eval_interval = 1_000
accumulation_steps = 16
base_lr = 3e-4
weight_decay=0.1
batch_size = 16
The draft model (medium) was trained only once
max_iters = 40_000
warmup_steps = 2_000
eval_iters = 20
eval_interval = 2_000
accumulation_steps = 16
base_lr = 3e-4
weight_decay = 0.1