Lambda · United States
Join Lambda's Cloud Services Engineering organization as a Senior Software Engineer. You will build and operate the distributed systems that power Lambda's GPU cloud, owning the full engineering lifecycle from design to continuous improvement. This role is ideal for engineers who enjoy cloud infrastructure, distributed systems, and solving complex problems. Benefits include health, dental, vision, 401k match, flexible PTO, and paid parental, medical, and caregiver leave.
- Depth in at least one general-purpose language, we work primarily in Go and Python, and candidates interview in the language they know best. We look for someone who can reason about concurrency, error handling, and testing in that language, not someone who has used it
- A track record of owning complex work through delivery and operation, including testing, staged rollout, monitoring, incident response, and root-cause improvement
- 7 or more years of professional software engineering experience, or equivalent evidence of impact building production systems
- Practical understanding of system design, data models, APIs, failure modes, performance, and the tradeoffs required to run reliable software in production
- Proven track record of aligning cross functional partners and gaining consensus around decisions and tradeoffs
- Experience designing, building, and operating backend services, distributed systems, infrastructure, or platform capabilities at meaningful scale
- 2+ years of experience building cloud services or platform infrastructure, or operating large-scale production systems on AWS, GCP, Azure, or a comparable cloud platform
- Depth in one or more cloud infrastructure or platform domains, such as compute, storage, networking, identity and access, developer platforms, container orchestration, usage metering and billing, databases, or fleet management
- Experience with Kubernetes, container orchestration, schedulers, controllers, or cloud control-plane systems
- Experience with infrastructure automation, durable workflow systems, event-driven architectures, or infrastructure as code
- Experience designing highly available, multi-region, or rapidly scaling distributed systems
- Familiarity with GPU infrastructure, HPC environments, or large-scale AI/ML training and inference workloads