Tasks:
- * Analyze data loading and communication bottlenecks
- * Build distributed training platforms
- * Build online and offline inference frameworks
- * Develop performance monitoring tools
- * Implement inference optimization techniques
- * Improve hardware utilization
- * Maintain distributed training frameworks
- * Optimize inference speed, memory, and energy use
- * Optimize training performance and resource utilization
- * Support model deployment
Perks/Benefits:
Skills/Tech stack required:
[C++] [CUDA] [CUDNN] [DeepSpeed] [Distributed Training] [Docker] [GPU Computing] [Horovod] [JAX] [Kubernetes] [Mixed Precision] [Mixed-precision training] [Model Inference] [Model Pruning] [Model Quantization] [MPI] [NCCL] [NPU Computing] [Performance optimization] [Python] [PyTorch] [Ray] [TensorFlow]
Educational requirements:
[Bachelor's Degree]
Role(s):
[AI Infrastructure Engineer] [Engineer] [Inference Engineer] [Infrastructure Engineer] [Machine Learning Infrastructure Engineer]