We would like to thank Axiomic Labs for allowing us to use their TrainWork framework to train this model.

ForgePlex-M2-9M

ForgePlex-M2-9M is a ~9.95M-parameter decoder-only language model from ForgeWorks. Trained on 30B tokens. Our second attempt at creating a <10m parameter model. We're proud of this product, while M1 was a promising start, M2 shows what we can do.

Q&A

What was the motivation behind M2?

"That is an excellent question. To be completely honest. George Mallory was asked why he wanted to climb Everest. He said, “Because it’s there.” I just see M2 as a mountain to climb"

Metric Value
Unique parameters 9,949,698
Intelligence Index 9.14
HellaSwag 28.02%
ARC (combined) 30.00%
PIQA 57.18%
ArithMark-3 34.60%
License Apache-2.0

Training

Trained on an all new dataset

Data Percentage
Finephrase 50%
DCLM - Baseline 30%
Proprietary Axiomic Labs dataset releasing soon 10%
Finemath 10%

Architecture

GQA + NeoX-style RoPE + RMSNorm + SwiGLU, with Qwen3.5-style attention output gates and Axiomic Labs TX4 style refresh gates on inject layers [5, 10] (kernel 9). XSA is off. Weights keep training key layout (no Llama remapping).

Component Details
Position encoding RoPE (theta=5,000, NeoX even/odd)
Normalization RMSNorm (eps=1e-6)
Feed-forward SwiGLU (intermediate 707)
Attention GQA — 8Q / 2KV, head_dim=32 + attn output gate
Refresh Layers 5, 10, kernel 9
Bias None
Embedding Weight tying
Depth × width 11 layers × 256 hidden
Context 1024 tokens
Vocab 4,096 custom BPE

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_path = r"C:\slm\ForgePlex-M2-9M"
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_path,
    trust_remote_code=True,
    torch_dtype=torch.float32,
    device_map="auto",
)

prompt = "Once upon a time"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.inference_mode():
    out = model.generate(**inputs, max_new_tokens=80, do_sample=False)
print(tokenizer.decode(out[0], skip_special_tokens=True))

Or run python usage.py from this folder.

Downloads last month
821
Safetensors
Model size
9.95M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including ForgeWorks/ForgePlex-M2-9M