AI/ML and tech enthusiast, training & fine-tuning models π¦₯π
OSS supporter & Love helping people and always looking to learn, build, and contributeπ
<3
while preparing the last class of the Training Agents live series during the summer, i spent some time reading the post-training sections of many frontier model reports, to learn how they use RL environments to improve their models, and wrote a blog about it
if you use any kind of coding harness, or you saw the Blender scenes that went viral recently, this might be interesting to you
There are quite a few good smaller parameter models that are capable for Agentic tasks:
The ones from the chart, I have tried a few already in my Jetson Orin Nano, βGemma4 E2B IT - cannot fit my RAM usage if use with TTS and embedder βQwen3.5 4B - just barely fit my RAM usage, need to add think/no_think βSpark X2.5 4B - need to build the forked llama.cpp; no vision β‘οΈNanbeige 4.2 3B - need to build the forked llama.cpp; slower than Ministral3-3B by 25%; no vision but good for coding; maybe run this is separate server for doing coding tasks β‘οΈAgents A1 4B - This one is quite interesting. Another Qwen3.5 4B base. I just learnt this right now. This model may surpassed the Ministral3-3B that I'm currently running. β‘οΈNeoHorse 1 4B - wait for GGUF version comes out; Qwen 3.5 4B base with vision striped β‘οΈNeedle2 45M - need to use separately from llama.cpp server; currently testing to see if it can be used as spawning sub-agents to do parallel tasks
7 replies
Β·
reactedtodanielhanchen'spost with π₯π9 days ago