Tasks:
- * Analyze data loading and communication bottlenecks
- * Build distributed training platforms
- * Build performance monitoring tools
- * Develop model-specific training and inference strategies
- * Develop online and offline inference frameworks
- * Implement model quantization and speculative decoding
- * Implement prefill-decode disaggregation and context caching
- * Maintain distributed training frameworks
- * Optimize hardware utilization
- * Optimize inference speed, memory, and energy use
- * Optimize training performance and resource utilization
- * Support model deployment
Perks/Benefits:
Skills/Tech stack required:
[C++] [CUDA] [CUDNN] [DeepSpeed] [Distributed Training] [Docker] [GPU] [Horovod] [JAX] [Kubernetes] [Language Models] [Large Language Models] [Mixed Precision] [Mixed-precision training] [Model Pruning] [Model Quantization] [NCCL] [NPU] [Performance Tuning] [Python] [PyTorch] [Ray] [TensorFlow] [Terraform] [Transformer Models]
Educational requirements:
[Bachelor's Degree]
Role(s):
[Distributed Training Engineer] [Engineer] [Inference Infrastructure Engineer] [Infrastructure Engineer] [Machine Learning Infrastructure Engineer] [Training Engineer]