Clera · Zürich
Join an engineering team building and scaling cloud infrastructure for AI research workloads. You will own key systems across compute, distributed workflows, data storage, and experiment orchestration, helping improve platform reliability, performance, and efficiency.
Build and operate cloud compute infrastructure that supports research workloads at scale.
Design and optimize distributed workflow scheduling for performance and cost efficiency.
Develop recovery safeguards for worker interruptions, application crashes, and partial results while preventing duplicate work.
Improve monitoring, logging, and diagnostic tools to help identify and resolve system failures.
Partner with researchers and engineers to develop platform capabilities and shape infrastructure architecture.
Typically 5 to 10 or more years of relevant experience in infrastructure or software engineering.
Strong Python and Linux skills, including scripting, process management, and debugging.
Experience operating services on a major cloud platform, including compute, storage, networking, and access management.
Understanding of distributed systems concepts such as queues, timeouts, retries, and idempotency.
Experience investigating production issues and balancing reliability, complexity, and cost.
Self-directed, collaborative approach, with sound judgment and strong follow-through.
Experience with KVM/QEMU, VM images or snapshots, desktop application automation, Terraform, monitoring tools, batch scheduling, experiment orchestration, or sandboxing untrusted code is beneficial.
Equity is offered, and visa sponsorship is available.
On-site in Zürich, Switzerland.
Inicia sesión para generar una carta de presentación para esta vacante.
Iniciar sesiónInicia sesión para ver cómo encaja este empleo con tu perfil.
Iniciar sesión