Ian Cole
codingiancole
AI & ML interests
Reinforcement learning, reward modeling, RLHF, policy optimization, offline RL
Recent Activity
upvoted a paper about 10 hours ago
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training upvoted a paper about 10 hours ago
Post-Training Language Models for Gold-Medal Performance in Coding Competitions upvoted a paper about 10 hours ago
Agentic Visual Generation: From Generative Models to Agentic ControlOrganizations
None yet