Tasks:
- * Accelerate data loading and preprocessing
- * Build performance monitoring systems
- * Design and optimize large-model training frameworks
- * Profile and tune training performance
- * Resolve distributed training communication and memory bottlenecks
Perks/Benefits:
Skills/Tech stack required:
[C++] [CI/CD] [Communication Computation Overlap] [Data parallelism] [Data Preprocessing] [DeepSpeed] [Distributed Training] [Expert parallelism] [High Performance] [High-Performance Computing] [Infiniband] [Megatron-LM] [Memory Optimization] [NCCL Profiler] [Nsight] [Performance Computing] [Performance Profiling] [Pipeline parallelism] [Python] [PyTorch] [RoCEv2] [Tensor Parallelism] [Training performance] [Training performance profiling]
Educational requirements:
N/A
Role(s):
[Distributed Training Engineer] [Engineer] [Framework Engineer] [Large Language Model Training Framework Engineer] [Learning Systems Engineer] [Machine Learning Systems Engineer] [Systems Engineer] [Training Engineer]