Tasks:
- * Adapt domestic and emerging GPU chips
- * Build and operate large-scale Kubernetes GPU clusters
- * Develop topology-aware and gang scheduling
- * Implement GPU failure recovery and end-to-end observability
- * Optimize GPU scheduling and utilization
- * Optimize large-model training and inference
- * Pool and virtualize GPU resources
- * Support mixed training and inference workloads
Perks/Benefits:
Skills/Tech stack required:
[C++] [Cluster management] [CUDA] [Gang Scheduling] [Go] [GPU Cluster] [GPU Cluster Management] [GPU memory] [GPU Memory Optimization] [GPU resource pooling] [GPU scheduling] [GPU virtualization] [Inference Optimization] [Kubernetes] [Large model] [Large model training] [Large-model training optimization] [Megatron] [Memory Optimization] [Mixed Precision] [Model Quantization] [Model training optimization] [Python] [PyTorch] [Resource pooling] [SGLang] [Topology Aware Scheduling] [Training Optimization] [Volcano]
Educational requirements:
[Bachelor's Degree]
Role(s):
[AI Infrastructure Engineer] [Computing Engineer] [Engineer] [GPU Computing Engineer] [Infrastructure Engineer]