Ema · Vancouver, Canada
Join Ema, a company that builds AI employees to carry out complex workflows across enterprise applications. As a Machine Learning Software Engineer, you will work on improving the performance of AI agents through data analysis, training, and evaluation. You will design and implement systems for multi-step agents, curate training environments, and evaluate data quality and performance. This role requires a master's or PhD in a relevant field, strong statistical judgment, and experience in production engineering.
- A master’s or PhD in a relevant field, or equivalent work or research experience. Papers, substantial open-source contributions, trained models and well-documented experiments can demonstrate that depth
- Statistical judgment. You can size an experiment, choose meaningful baselines and held-out tests, and account for variation across tasks, seeds and repeated runs. You can distinguish a real improvement from judge bias, data leakage or a benchmark shortcut
- Evidence of zero-to-one ownership. A system, model or research project you took from an ambiguous problem to a working result. We want to understand your contribution, the tradeoffs you made, and what changed when the work met real users or realistic tasks
- These are project-specific strengths; post-training experience is optional, and no candidate needs the entire list. Bring a repository, paper, model or technical write-up that lets us examine how you think and what you built
- Depth you can defend. Substantial work in at least one of agent/tool-use systems, post-training, reward modeling or RL environments, retrieval and memory, or evaluation design. Be ready to explain the mechanism, the alternatives you rejected and the failure modes you found. One area you can teach us beats five you’ve touched
- Production engineering judgment. You can debug across the model and system boundary, isolate a failure, and turn the result into maintainable production code
- Honest measurement. You would rather retire your own approach after a clean negative result than ship an improvement that disappears under a stronger evaluation
- Interactive agent environments/harnesses for software engineering, web or tool use; large-scale trace analysis, data curation or synthetic generation
- Practical security work on prompt injection, data governance or permission boundaries for agents that can act and improve themselves
- Open-model post-training with TRL, veRL, OpenRLHF, or similar or a custom loop, especially debugging reward hacking or unstable optimization
- Designing systems to support complex, long-horizon agent work across a multitude of modalities and platforms
- Serving with vLLM or SGLang, distillation, quantization, or multi-node GPU training
Inicia sesión para generar una carta de presentación para esta vacante.
Iniciar sesiónInicia sesión para ver cómo encaja este empleo con tu perfil.
Iniciar sesión