Tasks:
- * Analyze rollout trajectories and training metrics
- * Design agent training tasks and verifiable rewards
- * Improve long-horizon reinforcement learning algorithms
- * Reproduce and validate agent reinforcement learning research
- * Research LLM reinforcement learning post-training
- * Run ablation experiments and diagnose training failures
Perks/Benefits:
Skills/Tech stack required:
[Advantage Estimation] [Credit Assignment] [Experiment analysis] [Gradient methods] [GRPO] [LLM post training] [Off Policy] [Off Policy Learning] [OPD] [Policy Gradient] [Policy Gradient Methods] [Policy learning] [Post-training] [PPO] [Python] [PyTorch] [Reinforcement Learning] [Reward Design] [RLHF] [RLVR] [SFT] [Transformer Architecture]
Educational requirements:
N/A
Role(s):
[Intern] [LLM Agent Research Intern] [Reinforcement Learning Research Intern] [Research Intern]