Lucas Oliveira
lucasolive
AI & ML interests
Reinforcement learning, reward modeling, RLHF, policy optimization, offline RL
Recent Activity
liked a dataset about 17 hours ago
andersonbcdefg/red_teaming_reward_modeling_pairwise_no_as_an_ai upvoted a paper about 17 hours ago
LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches upvoted a paper about 17 hours ago
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language ModelsOrganizations
None yet