Tasks:
- * Build training execution and state management
- * Develop pretraining framework modules
- * Develop reusable infrastructure for pretraining and reinforcement learning
- * Implement checkpointing and fault recovery
- * Improve training monitoring and diagnostics
- * Optimize distributed parallelism and communication overlap
- * Optimize end-to-end training performance
- * Profile data, compute, memory, and communication bottlenecks
- * Resolve training failures and stability issues
- * Scale training from hundreds to thousands of GPUs
Perks/Benefits:
Skills/Tech stack required:
[AllGather] [AllReduce] [CUDA] [Data pipeline] [Data pipeline optimization] [DeepSpeed] [Diffusion Models] [Distributed Checkpointing] [Distributed Training] [FSDP] [GPU memory] [GPU memory management] [GPU Performance] [GPU performance profiling] [Low Precision] [Low-precision training] [Megatron-LM] [Memory Management] [Multimodal Models] [NCCL] [Nsight] [Parallel training] [Performance optimization] [Performance Profiling] [Pipeline Optimization] [Python] [PyTorch] [PyTorch Profiler] [ReduceScatter] [Transformer] [Triton] [Video Generation]
Educational requirements:
N/A
Role(s):
[AI Infrastructure Engineer] [Distributed Training Engineer] [Engineer] [Infrastructure Engineer] [Training Engineer]