About Nebius
Nebius is leading a new era in cloud infrastructure for the global AI economy, building a full-stack AI cloud platform. Listed on Nasdaq (NBIS) and headquartered in Amsterdam.
The Role
This role is for Nebius AI R&D, a team focused on applied research in AI, including reinforcement learning for agent training, scaling task data collection for SWE agents, and decontaminated evaluation frameworks. We are looking for senior- and staff-level ML engineers to work on guided search and reinforcement learning for agentic systems, reinforcement learning for reasoning models, and efficient model distillation.
Example Responsibilities
- Conducting experiments to train large language models on traces of interactions with various environments.
- Exploring methods of guided generation and search in the trajectory space.
- Mining relevant data at web scale and using it efficiently in model post-training.
- Conducting reinforcement learning experiments in verifiable domains, and exploring non-verifiable reward signals.
We Expect You to Have
- A profound understanding of theoretical foundations of machine learning and reinforcement learning.
- Deep expertise in modern deep learning for language processing and generation.
- Substantial experience training large models on multiple computational nodes.
- Strong software engineering skills (mostly Python) and deep experience with JAX.
- Strong communication and leadership abilities.
Nice to Have
Experience with deep RL for LLMs (reward modeling, DPO, PPO), familiarity with RoPE, ZeRO/FSDP, Flash Attention, and quantization.