Tasks:
- * Deploy and accelerate models on edge devices
- * Design and optimize large language model inference services
- * Develop CUDA operators and optimize computation graphs
- * Implement model quantization, pruning, distillation, and operator fusion
- * Optimize KV cache and dynamic batching for high-throughput inference
Perks/Benefits:
Skills/Tech stack required:
[Cache optimization] [C/C++] [CUDA] [CUDA kernel] [CUDA kernel development] [Dynamic batching] [Graph compilation] [Kernel development] [Knowledge Distillation] [KV cache] [KV cache optimization] [Language Model Inference] [Large Language Model] [Large language model inference] [Linux] [Model Inference] [Model Pruning] [Model Quantization] [Operator fusion] [Python] [TensorRT] [Transformer Architecture] [Triton] [VLLM]
Educational requirements:
[Bachelor's Degree] [Master's Degree]
Role(s):
[AI Inference Deployment Engineer] [Deployment Engineer] [Engineer] [Inference Deployment Engineer] [Learning Engineer] [Machine Learning Engineer]