bob-int
The weights and every file in this repository are licensed Apache-2.0. The bob name and marks are not part of that grant.
Model: bob 2-ck_13000 (bob-talk)
A small integer language model. The same, every time.
Try it: https://bob-talk.org ; the model in this repository is the model answering that page, verifiable by digest (below).
Bob exists because I asked the question: "If fixed point buys exactness, can exact multithreading/parallelism then provide speed?" I'm still working on that question. This project was inspired by Norm Winningstead, and my own dumb questions. Bob has demonstrated the ability to resume his own training, cross-ISA and byte-exact reproducible. I do not know if Bob's training model can do this on every possible combination- and make no claims to his ability outside of what I have already demonstrated.
What is in this repository
ck_13000.weights.bin the final weights, training step 13,000
sha256 53ae7609f22e580ac1d67ad1995dfc2d87c0ebcd5f60f3ab7136345f502c2890
ck_13000.opt.bin the optimizer state at the same step, so the
trajectory can be resumed, not only replayed
registration/exponents.json the registered fixed-point exponents and bit
widths for every activation and gradient site
registration/width-sites.json the accumulator width registrations
registration/init/ the initial weights and ordering- the
trajectory is replayable from step 0
SHA256SUMS digests for every file above
The property these files carry
Every number in training and inference is fixed-point integer arithmetic under a registered contract: each site has a registered exponent and width, and arithmetic that would leave its registration refuses- it stops with a named site and an exit code rather than saturating or wrapping silently.
Because the arithmetic is integer, results do not depend on floating-point summation order. The digest of the reply tokens for a given (prompt, seed, effort) is the same on every conforming build. Four pre-published examples with their digests are on https://bob-talk.org and embedded in its page source.
Two portability results are in the project record, each stated no wider than its measurement:
bob 2's own training (the weights in this repository) was saved mid-run and resumed on different rented machines across two GPU generations of one vendor, with checkpoint digests matching at the joints. All of bob 2's GPU lineages are one vendor; no cross-vendor claim is made for this trajectory.
A separate 80-step synthetic trial (2026-08-17, not these weights)
ran the same integer training on an Apple M4 via Metal (arm64, macOS)
and a Pascal-era NVIDIA GTX 1070 via Vulkan (x86_64, Linux)-> produced
byte-identical weights AND optimizer state- sha256
cddab458c1e2e8b9… / 436114216b977be2…- hashed independently on
each machine. The two devices also made byte-identical refusal
decisions: the enforcement machinery itself is deterministic across
vendors. The trial is synthetic and unregistered; it demonstrates the
arithmetic contract's portability, not this model's history.
I know of no shipped system combining an end-to-end integer training loop with integer optimizer state, overflow semantics that refuse with a named site and typed exit, and demonstrated mid-run save-and-resume with matching digests- resting on a documented 2026-08-17 survey of eleven named systems: unoccupied as of the search date, not impossible, not permanent. Please direct me toward prior art if available.
Verify this is the model behind bob-talk.org
sha256sum ck_13000.weights.bin
# 53ae7609f22e580ac1d67ad1995dfc2d87c0ebcd5f60f3ab7136345f502c2890
That digest is compiled into the public server, which refuses to start
if its embedded weights do not hash to it. Ask the live page
(POST /ask, {"prompt": "Alice was", "seed": "7"}) and compare the
reply digest against the published table on the page.
Format
Weights are raw little-endian integer tensors in the registered
layout; registration/exponents.json gives each site's fixed-point
interpretation. This is not a transformers/safetensors checkpoint and
will not load with from_pretrained. Geometry: d_model 512, 8 layers,
8 heads, head_dim 64, d_mlp 2048, context 8192, vocab 8192; parameter
count 33,608,704 with the learned position table (proved at both ends
of the pipeline; receipts in the project record).
Inference code: [TBD - LIVE GH REPO NOT RELEASED; https://bob-talk.org/ for live inference]
What is NOT claimed
bob 2 is deterministic in the sense stated above. He is not thereby accurate, clever, or safe, and no such claim is made. He is small (33.6M parameters), his prose is what a 33.6M-parameter model's prose is, and the property I claim to be interesting is the arithmetic, not the model's vocabulary.
Provenance
Trained August of 2026 by f32. Training step 13,000; two legs; the enforcement regime is part of the record (one leg ended in a registered disagreement event on a named site; the refusal machinery was demonstrated live, then the registration corrected and the run completed).