Tasks:
- * Automate incident diagnosis and resolution
- * Build intelligent operations platform
- * Design operations workflows and approval processes
- * Ensure hybrid-cloud availability and stability
- * Expose operations capabilities through AI Agent skills and MCP services
- * Implement automated rollback mechanisms
- * Improve reliability and resource operations
- * Investigate large-scale GPU cluster issues
- * Monitor and maintain services and resources
- * Operate business services and infrastructure
- * Optimize cost, quality, efficiency, and security
- * Troubleshoot AI training system stability
Perks/Benefits:
Skills/Tech stack required:
[AI Agents] [AI Training] [AI training infrastructure] [Cloud infrastructure] [Cloud Native] [Cloud-native operations] [Cluster management] [Golang] [GPU Cluster] [GPU Cluster Management] [Hadoop] [Hybrid Cloud] [Hybrid cloud infrastructure] [Istio] [Kubeflow] [Kubernetes] [Langchain] [Linux] [LLM APIs] [MCP] [Middleware operations] [Operations automation] [Prometheus] [Python] [RDMA Networking] [TCP/IP] [Training Infrastructure] [VictoriaMetrics] [Workflow Orchestration]
Educational requirements:
N/A
Role(s):
[DevOps Engineer] [Engineer] [Infrastructure] [Infrastructure Engineer] [Reliability Engineer] [Site Reliability Engineer]