Tasks:
- * Adapt GPU hardware including domestic accelerators
- * Build and operate large-scale Kubernetes GPU clusters
- * Implement GPU fault recovery and end-to-end observability
- * Improve GPU utilization through scheduling optimization and training-inference workload co-location
- * Optimize mixed-precision computing, model quantization, and GPU memory
- * Optimize topology-aware and gang scheduling
- * Pool and virtualize GPU resources
Perks/Benefits:
Skills/Tech stack required:
[C++] [Cluster management] [CUDA] [Fault Recovery] [Gang Scheduling] [Go] [GPU Cluster] [GPU Cluster Management] [GPU fault recovery] [GPU memory] [GPU Memory Optimization] [GPU observability] [GPU resource pooling] [GPU scheduling] [GPU virtualization] [Kubernetes] [Megatron] [Memory Optimization] [Mixed Precision] [Mixed Precision Computing] [Model Quantization] [Precision computing] [Python] [PyTorch] [Resource pooling] [SGLang] [Topology Aware Scheduling] [Volcano]
Educational requirements:
[Bachelor's Degree]
Role(s):
[AI Infrastructure Engineer] [Computing Engineer] [Engineer] [GPU Computing Engineer] [Infrastructure Engineer]