·
AI & ML interests
Contact: arxivgpt@gmail.com
Recent Activity
reacted to ginigen-ai's post with 👍 about 21 hours ago OpenRouter Leaderboard — every model, every provider, one comparable table. Price, precision, uptime, measured latency and language quality on the same axes.
Building it turned up three things.
We graded 330 models on Korean and two axes collapsed.
Honorifics — only 8.5% earn an A
Knowledge of Korean institutions — 9.4%
Every other axis sits above 31%
Fluency hides it. A model can write clean, natural Korean and still attach an honorific to a coffee cup. Fluent and wrong at the same time is worse than obviously broken, because nobody catches it in review.
A 2023 model beats the 2026 flagships. gpt-3.5-turbo-16k scores a perfect 3.00. Korean cannot be inferred from release date, parameter count or English benchmarks — it has to be measured, per model.
Quality, value and speed are three different models. Across five axes, the same model almost never takes two columns.
425 models, latency measured on 329 on a paid API, Korean graded on 330. Three languages, three currencies, daily refresh, open API, no key.
📝 https://huggingface.co/blog/ginigen-ai/openrouter-leaderboard 🔎 https://huggingface.co/spaces/ginigen-ai/open-router-leaderboard View all activity Organizations
view article Same bytes, closer to the original: two lines of AutoRound we had wrong
FINAL-Bench
• • 12
view article Did the Civilization Emerge, or Was It Recited?
FINAL-Bench
• • 9
view article Writing Down the Line Between Luck and Skill
FINAL-Bench
• • 14
view article We changed one line and the benchmark score moved 0.21 AUROC
FINAL-Bench
• • 15
published an article about 1 month ago view article Who Tells You Whether the Molecule Your AI Just Designed Is Any Good?
FINAL-Bench
• • 11
published an article about 1 month ago view article AX-Ray, Finding Causal-Leakage Defects in Two General-Purpose Public Models
FINAL-Bench
• • 13
published an article about 2 months ago view article The Fast Gemma Challenge: our verified-SOTA recipe, in full
published an article about 2 months ago view article POCKET: a 35-billion-parameter model that runs on your iPhone — and on your PC with no GPU
FINAL-Bench
• • 11
published an article about 2 months ago view article Aether-7B-5Attn: A 100% Open-Source Sovereign Foundation Model — and a Controlled Experiment in Heterogeneous Attention
FINAL-Bench
• • 21
view article VKUE: No GPU? Runs Anyway — a 34.7B Reasoner on a Laptop and on Bare CPU
FINAL-Bench
• • 17
view article Quantum Cryptanalysis on Real Hardware: Pushing Symmetric-Structure Key Recovery Beyond the Published Frontier
view article Chitos: From Detection to Proof — An Autonomous Security AI That Actually Exploits
FINAL-Bench
• • 19
view article FINAL-Bench Quantum: An Open, Neutral Benchmark for Quantum-Computing Methods
FINAL-Bench
• • 17
view article Training-Free Reasoning at 88.89% on GPQA Diamond: How Darwin Family Hit Frontier Scores Without a Single Gradient Step
FINAL-Bench
• • 18
view article Darwin-TTS: We Gave a TTS Model 3% of an LLM's Brain — It Started Showing Emotion
FINAL-Bench
• • 13
view article "Darwin-27B-Opus: Surpassing the Foundation Model Without Training"
FINAL-Bench
• • 16
view article Darwin V6: Diagnostic-Guided Evolutionary Model Merging
view article "The Child That Surpassed Both Parents Through MRI-Guided Evolutionary Merge"
FINAL-Bench
• • 15
view article Introducing WM Bench: A Benchmark for Cognitive Intelligence in World Models
FINAL-Bench
• • 13