YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

PALIMPSESTE

A Self-Referential Hypervectorial Cortex with Kuramoto Attractor Dynamics

Kanerva/Plate VSA on CPU. No weight matrix, no gradient, no GPU. Knowledge is an append-only log of (address, value, weight) traces in H = {+1,-1}^D. Read-back is Phi(q) โ€” bundle over the Hamming ball N_r(q), LSH-indexed. Teach a fact in one write.

A parchment that is never erased, only overwritten in layers.

No GPU. No gradient. No weights. No torch. No transformer. 30 Python modules. 361 tests. 2.2M memory traces. Trained in 51 minutes on CPU. Address space C_addr(ฮฒโ‰ˆ4) ~ 10^1505 locations at D=20 000. Read-back limit is per ball: k* โ‰ˆ 0.02 D exact / 0.03 D soft / 0.067 D theory.


Install (start here)

Python >= 3.10, numpy. No CUDA, no torch.

# 1. clone
git clone https://huggingface.co/thefinalboss/palimpseste-max palimpseste
cd palimpseste

# 2. env
python3 -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate
pip install -U pip
pip install -e .

# 3. load (10.4 GB, 3 chunks โ€” from_pretrained merges them)
python3 - << 'PY'
from palimseste.hf import HFPalimpsesteLM
lm = HFPalimpsesteLM.from_pretrained("thefinalboss/palimpseste-max")
print(lm.respond("who are you"))
PY

Disk: ~12 GB free for the chunks + merge. First load downloads from this repo.

Teach a fact (O(1) write, no fine-tune):

from palimseste.chat import Conversation
conv = Conversation(model=lm, learn_live=True)
conv.teach("capital of mars", "olympus mons city")
conv.reset()
print(conv.respond("capital of mars"))

Capacity probe (flags balls over k*_exact โ‰ˆ 0.02 D):

python examples/occupancy.py --sweep
python examples/occupancy.py --live --limit 306

More examples below. White paper: whitepaper/PALIMPSESTE_White_Paper.pdf. Capacity note: notes/CAPACITY.md.


Links


Model Specifications

Spec Value
Dimensionality (D) 20,000
Tokenizer BPE (3,000 vocab)
Context window 256 (hierarchical 1M available)
Training pairs 88,800 (curated + generated + TriviaQA + math tables)
Tokens trained 2,225,193
Training time 3,082 seconds (51 min, CPU only)
Training speed 722 tok/s
Model size 10.4 GB (3 chunks under 5 GB each)
Stored parameters 0 (all reconstructed on-the-fly)
Inference latency 2-3s per response (LSH tuned K=14/L=6)
LSH tuning time 8,431 seconds
Modules 30
Tests 361, all passing

Capacity โ€” three real numbers (do not collapse)

The SDM address-space count is real and exponential in D. It is not the read-back limit. ฮฆ(q) bundles only the Hamming ball N_r(q).

Quantity Law At D=20 000
C_addr (SDM locations, ฮฒโ‰ˆ4) 0.14 ยท D ยท 2^(D/4) ~10^1505 locations. 2^5000 โ‰ˆ 10^1505.
k* per ball linear in D exact-recover โ‰ˆ 0.02 D (~400) ยท soft โ‰ฅ0.99 โ‰ˆ 0.03 D (~600) ยท theory D/(2 ln M) โ‰ˆ 0.067 D (~1300)
N_traces append-only log 2.2M (this ULTRA dump) / 131 541 (white-paper dump)

"Grows exponentially with D" applies to C_addr. Retrieval capacity is linear and per neighborhood, which is why 2.2M traces still read clean: LSH + radius keep |N_r(q)| under k*. 100% confidence means no probed ball has crossed k*_exact yet โ€” not k* = โˆž.

Live probe (flags any query with |N_r(q)| > 0.02 D):

python examples/occupancy.py --sweep
python examples/occupancy.py --live --limit 306

See notes/CAPACITY.md and palimseste/capacity.py.


Proven Results โ€” All Real HV Retrieval

Every response below was retrieved by Hamming-neighborhood associative lookup from 2,225,193 traces at 100% exact-recover (every ball still under k* โ‰ˆ 0.02 D). No placeholders.

Identity

Q: who are you
A: i am palimpseste, a self-referential hypervectorial cortex.

Q: how are you different from chatgpt
A: chatgpt has billions of weights and uses gpus. i have zero weights
   and run on cpu. i learn by writing, not by gradient.

Q: do you use a gpu
A: no, i do not use a gpu. only bitwise operations: xor, popcount,
   thresholded sum.

Knowledge

Q: what is python
A: python is an interpreted programming language known for its clear
   syntax and large standard library.

Q: what is javascript
A: javascript is a programming language that enables interactive web
   pages and dynamic content.

Q: what is dark matter
A: dark matter is a mysterious substance that does not emit or absorb
   light but exerts gravitational pull.

Math (computed from stored tables)

Q: what is 50 times 50
A: 50 times 50 is 2500

Q: what is 25 times 13
A: 25 times 13 is 325

Q: what is the capital of japan
A: the capital of japan is tokyo.

Creative Writing

Q: write a poem about the sea
A: the sea whispers ancient secrets to the shore, each wave a verse,
   each tide a metaphor, in depths where light fears to descend,
   lies a silence that has no end.

Q: write a haiku
A: stillness in the code, bits flip like falling autumn leaves,
   silence after compute.

Conversation

Q: i am sad
A: i am sorry to hear that. sometimes talking helps. what is on your mind?

Q: that is interesting
A: i am glad you find it interesting. would you like to explore it further?

Architecture โ€” 5 Layers, 30 Modules

Layer 1: HV Substrate

Module Description
hv.py Bipolar hypervectors: bind (XOR), bundle (majority), similarity (Hamming)
lsh.py Locality-sensitive hashing for O(log M) retrieval
memory.py Append-only knowledge base with H_meta subspace
phi.py Hard Phi: Hamming-radius neighborhood retrieval
capacity.py Per-ball occupancy + k* instrument (exact / soft / Jolin)
soft_phi.py Soft Phi: exponential-weighted continuous retrieval
learner.py O(1) write primitive + Encoder
tokenizer.py Character-level tokenizer (legacy)
bpe.py BPE sub-word tokenizer

Layer 2: Language Model

Module Description
lm.py PalimpsesteForCausalLM: training, inference, BPE, multi-scale
serialization.py Streaming save/load + chunked file split/merge for >5GB models
hf.py HFPalimpsesteLM: save/load/push, auto-merges chunks on load
hierarchical_context.py 1M token context via hierarchical chunking

Layer 3: Cortex (7 features impossible for LLMs)

Feature Description
Instant Expertise Learn any document in 0.1s
Dream Consolidation Replay memory, discover concepts autonomously
Compositional Reasoning Decompose, resolve, compose multi-hop answers
Multi-Modal Fusion Text + image bound natively in HV space
Meta-Learning Self-tune kernel parameters at runtime
Episodic + Semantic Memory Distinguish conversation events from facts
Analogical Reasoning Solve a:b :: c:? via HV algebra

Layer 4: Generation

Module Description
creative.py Mixture logits from K candidates (intersection boost)
ngram_model.py Markov transition model for local coherence
kuramoto.py Kuramoto attractor dynamics for emergent generation
coherent.py Coherent generator: 85% retrieval + 10% Kuramoto + 5% transition

Layer 5: Reasoning & Evolution

Module Description
chat.py Conversation with entity tracking + pronoun resolution
cognitive.py CognitiveAgent: confidence, self-correction, curiosity
reasoning.py Fact chaining (A -> B -> C)
consolidation.py Hebbian co-activation to abstract concepts
abstraction.py HV clustering to concept centroids
meta.py Lyapunov-bounded self-rewrite (Axiom 5)
attention.py HV selective attention
hv_word2vec.py Distributional word embeddings
multiquery.py Multi-level query: surface + words + semantic
vision.py ImageEncoder: images to hypervectors
evolution.py Response synthesizer, entity tracker, query router, code bank, calibrator
loop.py Autonomous active-inference agent

Kuramoto Attractor Dynamics

The core innovation for creative generation. Standard associative memory can only repeat stored sequences. Kuramoto dynamics enable emergence.

Each retrieved HV candidate is an oscillator coupled by Hamming similarity. The coupled dynamics converge to an attractor โ€” a novel HV state that is a creative synthesis of multiple knowledge fragments.

Kuramoto model: d_i/dt = _i +  K_ij sin(_j - _i)

   i  = phase of oscillator i
   i  = natural frequency (from retrieval similarity)
  K_ij = coupling strength (from pairwise HV similarity)

Coherent Blending

The CoherentGenerator blends three signals at each token:

final = retrieval x 0.85 + kuramoto x 0.10 + transition x 0.05

Confidence-based creativity suppression: when retrieval is confident, creative weights are suppressed quadratically. Result: perfect coherence on known questions (coherence = 1.00), creative modulation on novel ones.


21 Cognitive Capabilities

# Capability Description
1 Instant Expertise Learn any document in 0.1s
2 Dream Consolidation Autonomous concept discovery
3 Compositional Reasoning Decompose, resolve, compose
4 Multi-Modal Fusion Text + image in HV space
5 Meta-Learning Self-tuning at runtime
6 Episodic + Semantic Memory "You told me" vs "The fact is"
7 Analogical Reasoning HV algebra for a:b :: c:?
8 BPE Tokenizer 1.6x fewer tokens
9 N-gram Boosted Retrieval Multi-scale voting
10 Multi-Scale Encoding Char + word + sentence
11 Native RAG Retrieve during generation
12 Iterative Refinement Draft to correct
13 Template Extraction Structural patterns
14 Massive Ingestion 20M tokens in 2h
15 Response Synthesizer Combine multiple facts
16 Entity Tracker Pronoun resolution across turns
17 Query Router Intent classification
18 Code Pattern Bank Store/retrieve code snippets
19 Confidence Calibrator Uncertainty estimation
20 Kuramoto Attractor Emergent HV from oscillator coupling
21 Coherent Generator Retrieval-dominant with creative modulation

Install and Quick Start

Install

git clone https://huggingface.co/thefinalboss/palimpseste-max palimpseste
cd palimpseste
pip install -e .

Requires Python >= 3.10 and numpy. No GPU, no torch.

Chat

from palimseste.hf import HFPalimpsesteLM

lm = HFPalimpsesteLM.from_pretrained("thefinalboss/palimpseste-max")
print(lm.respond("who are you"))
# -> "i am palimpseste, a self-referential hypervectorial cortex."

Note: The model is 10.4 GB, stored as 3 chunks under 5 GB each. from_pretrained automatically downloads and merges all chunks.

Teach a new fact instantly

from palimseste.chat import Conversation

conv = Conversation(model=lm, learn_live=True)
conv.teach("capital of mars", "olympus mons city")
conv.reset()
print(conv.respond("capital of mars"))  # -> "olympus mons city"

Kuramoto creative generation

from palimseste.coherent import CoherentGenerator
from palimseste.kuramoto import KuramotoAttractor

gen = CoherentGenerator(
    lm=lm,
    kuramoto=KuramotoAttractor(n_iterations=30),
    retrieval_weight=0.85,
    kuramoto_weight=0.10,
)
result = gen.generate("who are you", max_new_tokens=100)
print(result.text)
print(f"Coherence: {result.coherence_score:.2f}")

Cortex features

from palimseste.cortex import InstantExpert, Dreamer, Composer
from palimseste.reasoning import Reasoner

# Instant expertise
expert = InstantExpert(lm=lm)
expert.learn_from_text("Photosynthesis converts light into energy via chlorophyll.")

# Dream consolidation
dreamer = Dreamer(mem=lm.mem, phi=lm.phi)
dreamer.dream(n_cycles=3)

# Compositional reasoning
conv = Conversation(model=lm)
reasoner = Reasoner(conv=conv)
composer = Composer(reasoner=reasoner)
result = composer.reason("what is the capital of the country that won the world cup 2018")
print(result.answer)  # -> "paris"

Deploy the web app

pip install fastapi uvicorn
python examples/api_server.py --model ./palimpseste-max --port 3332
cd web && npm install && npm run dev

PALIMPSESTE vs GPT-4

Property GPT-4 PALIMPSESTE
Knowledge storage Dense weight matrices Append-only (a,v,w) log
Learning Gradient + backprop O(1) memory write
GPU Required Not needed (XOR + popcount)
Forgetting Catastrophic No overwrite (append-only). Loss = crosstalk when |N_r| > k*
Live learning Requires fine-tuning Instant (O(1) write)
Parameters ~176 billion Zero (reconstructed)
Model size 200+ GB 10.4 GB
Training cost Millions of dollars 51 minutes on CPU
Dreaming Impossible Hebbian consolidation
Attractor dynamics N/A Kuramoto coupling
Creativity Weight interpolation Emergent attractors

Chunked Model Storage

The 10.4 GB model is split into 3 chunks under 5 GB each to comply with Hugging Face file size limits:

File Size
palimpseste_memory.bin.part0 4.5 GB
palimpseste_memory.bin.part1 4.5 GB
palimpseste_memory.bin.part2 2.0 GB

from_pretrained automatically detects chunks, downloads them, merges to a temporary file, and loads the full 2.2M trace memory.


Testing

pytest -q    # 361 tests, all passing
Test File Tests Coverage
test_hv.py 22 Bind, bundle, similarity, Hamming
test_lsh.py 9 LSH index, query, tuning
test_memory.py 11 Append-only writes, decay
test_phi.py 11 Hard/soft Phi retrieval
test_learner.py 14 Encoder atoms, roles, sequences
test_lm.py 24 Training, generation, respond
test_chat.py 12 Conversation, fuzzy match, streaming
test_conversation.py 22 Multi-turn, teach, reset
test_cognitive.py 20 Confidence, correction, curiosity
test_reasoning.py 9 Fact chaining A->B->C
test_long_context.py 14 1M token hierarchical context
test_upgrades.py 20 BPE, attention, abstraction
test_semantic.py 14 HV-Word2Vec, multi-query
test_vision.py 11 Image encoder
test_api.py 5 FastAPI endpoints
test_meta.py 11 Lyapunov self-rewrite
test_loop.py 10 Active inference agent
test_consolidation.py 9 Hebbian co-activation
test_cortex.py 18 Cortex features 1-3
test_cortex2.py 17 Cortex features 4-7
test_killer.py 23 BPE, n-gram, multi-scale, RAG
test_evolution.py 26 Synthesizer, tracker, router
test_creative.py 10 Mixture logits, novel tokens
test_kuramoto.py 11 Attractor dynamics, coherence
test_coherent.py 8 Blended generation

Training Scripts

Script Description
examples/train_ultra.py Train ultra model: 88K pairs, 2.2M tokens, 10.4 GB
examples/train_100k.py Train 100K model: 30K pairs, 870K tokens
examples/train_massive.py Train massive model with all corpora
examples/train_killer.py Train with BPE + killer knowledge corpus
examples/train_fluid.py Train with conversation + creative corpus
examples/api_server.py FastAPI server with React app
examples/chat.py Terminal chat interface

License

MIT. Created by Philippe-Antoine Robert, 2026.

Downloads last month
110
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support