Tasks:
- * Apply AI to alerting, root cause analysis, and capacity forecasting
- * Build, upgrade, and stabilize Kubernetes clusters
- * Develop and optimize operations automation tools
- * Handle operational changes, inspections, and incidents
- * Improve CI/CD pipelines and monitoring, logging, and alerting systems
- * Plan and deploy server, storage, network, and GPU infrastructure
Perks/Benefits:
Skills/Tech stack required:
[Alerting] [Capacity forecasting] [Cause analysis] [CI/CD] [Containerization] [Go] [GPU infrastructure] [Infrastructure automation] [Kubernetes] [Logging] [Monitoring] [Python] [Root cause] [Root Cause Analysis] [Shell Scripting]
Educational requirements:
[Bachelor's Degree] [Master's Degree]
Role(s):
[DevOps Engineer] [Engineer] [Reliability Engineer] [Site Reliability Engineer]