In this role, you'll research and evaluate misalignment in frontier AI systems to understand loss-of-control risks such as research sabotage and reward-seeking.
Build and run alignment evaluations that test for risks not captured by current benchmarks.
Conduct pre-deployment testing of AI systems and report findings to frontier AI companies and governments.
Design and develop software and tooling to improve alignment evaluation efficiency and usability.
Contribute to publications and technical reports advancing the field's understanding of misalignment risks.
Tags
Carta de presentación
Inicia sesión para generar una carta de presentación para esta vacante.