·
AI & ML interests
Working on small sota models
Recent Activity
repliedto their post about 8 hours ago A Small Model is All You Need. Meet `palmer-006` (90M)
After 3 years of experiments, we are finally releasing our flagship tiny model: **palmer-006**.
If you are building for edge hardware, SBCs (Raspberry Pi, etc.), or low-power devices, this is for you. Inspired by Andrej Karpathy's idea of a self-contained "cognitive core," we wanted to see how much power we could pack into a sub-100M parameter footprint.
🧠 **How we "Palmerized" it:**
We believe in starting our experiments with the absolute strongest baseline possible.
1. Light fine-tuning on highly curated data
2. Model merging
3. Another light fine-tuning round
4. Adjusted Mamba for maximum token speed ⚡️
⚠️ *Note: This is a foundational language model. It has not been instruction-tuned yet!*
Also, since this needs instruction tuning next to become a chat assistant—**what dataset would you recommend we use for the instruct tune?**
---
🔗 **Quick Links & Info:**
* **License:** Open for research, education, hobby, and modification! (For commercial use/hosted APIs, shoot an email to nosoyhackercodigo@gmail.com. *PS: Donators can claim a free commercial license!*)
* **Attribution:** Built using AI tech from the Technology Innovation Institute (TII).
Can't wait to see what you build at the edge. Let me know your prompt completions below! 👇
https://huggingface.co/appvoid/palmer-006 repliedto their post about 9 hours ago A Small Model is All You Need. Meet `palmer-006` (90M)
After 3 years of experiments, we are finally releasing our flagship tiny model: **palmer-006**.
If you are building for edge hardware, SBCs (Raspberry Pi, etc.), or low-power devices, this is for you. Inspired by Andrej Karpathy's idea of a self-contained "cognitive core," we wanted to see how much power we could pack into a sub-100M parameter footprint.
🧠 **How we "Palmerized" it:**
We believe in starting our experiments with the absolute strongest baseline possible.
1. Light fine-tuning on highly curated data
2. Model merging
3. Another light fine-tuning round
4. Adjusted Mamba for maximum token speed ⚡️
⚠️ *Note: This is a foundational language model. It has not been instruction-tuned yet!*
Also, since this needs instruction tuning next to become a chat assistant—**what dataset would you recommend we use for the instruct tune?**
---
🔗 **Quick Links & Info:**
* **License:** Open for research, education, hobby, and modification! (For commercial use/hosted APIs, shoot an email to nosoyhackercodigo@gmail.com. *PS: Donators can claim a free commercial license!*)
* **Attribution:** Built using AI tech from the Technology Innovation Institute (TII).
Can't wait to see what you build at the edge. Let me know your prompt completions below! 👇
https://huggingface.co/appvoid/palmer-006 View all activity Organizations
Viewer
• Updated • 5k • 45
Viewer
• Updated • 595 • 11
Viewer
• Updated • 1.6M • 42
• 1
Viewer
• Updated • 3.91M • 11
appvoid/no-prompt-oasst-mini
Viewer
• Updated • 100 • 24
Viewer
• Updated • 1.06k • 12
Viewer
• Updated • 1.06k • 35
Viewer
• Updated • 19.9k • 39
Viewer
• Updated • 389 • 10
Viewer
• Updated • 1M • 47
appvoid/noisy-textbook-5k
Viewer
• Updated • 5k • 18
appvoid/noisy-textbook-15k
Viewer
• Updated • 15k • 5
appvoid/noisy-textbook-25k
Viewer
• Updated • 25k • 6
appvoid/noisy-textbook-50k
Viewer
• Updated • 50k • 4
appvoid/noisy-textbook-70k
Viewer
• Updated • 70k • 21
Viewer
• Updated • 75k • 16
Viewer
• Updated • 12.9k • 10
appvoid/simple-prompt-oasst
Viewer
• Updated • 12.9k • 9
Viewer
• Updated • 15k • 29
• 1
appvoid/no-prompt-openhermes
Viewer
• Updated • 242k • 50
Viewer
• Updated • 50k • 21