Léo Martin
leo-mart
AI & ML interests
Reinforcement learning, reward modeling, RLHF, policy optimization, offline RL
Recent Activity
liked a model about 19 hours ago
zhiqings/LLaVA-RLHF-13b-v1.5-336 upvoted a paper about 19 hours ago
VideoGen-Agent: Reinforcing Video Generation Agents upvoted a paper about 19 hours ago
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention MechanismsOrganizations
None yet