-
FlowRL: Matching Reward Distributions for LLM Reasoning
Paper • 2509.15207 • Published • 119 -
Kwaipilot/KAT-Dev-72B-Exp
Text Generation • 73B • Updated • 235 • • 157 -
Agentic Entropy-Balanced Policy Optimization
Paper • 2510.14545 • Published • 109 -
Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
Paper • 2511.13288 • Published • 19
Malkesh Dalia
malkesh2911
AI & ML interests
None yet
Recent Activity
upvoted an article 17 days ago
State of Open Models: Summer 2026 Observations upvoted a paper about 2 months ago
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs liked a model 2 months ago
zai-org/GLM-5.2Organizations
None yet