Léo Martin
leo-mart
·
AI & ML interests
Reinforcement learning, reward modeling, RLHF, policy optimization, offline RL
Recent Activity
upvoted a paper about 5 hours ago
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks liked a model 1 day ago
zhiqings/LLaVA-RLHF-13b-v1.5-336 upvoted a paper 1 day ago
VideoGen-Agent: Reinforcing Video Generation AgentsOrganizations
None yet