Tasks:
- * Adapt and optimize inference operators
- * Analyze inference latency, FPS, utilization, memory, and bandwidth
- * Convert, compile, deploy, and debug models for edge inference
- * Deploy and optimize AI models on heterogeneous CPU/GPU/NPU/DSP platforms
- * Deploy robotics perception and VLM/LLM models on edge devices
- * Develop inference engines, runtimes, and heterogeneous resource scheduling software
- * Optimize memory management and data transfer
- * Quantize models and validate accuracy
Perks/Benefits:
- + Exposure to Qualcomm and NVIDIA platforms
- + Exposure to robotics and VLM/LLM applications
- + Hands-on work on core inference software
- + Potential conversion to a full-time role
- + Structured training and evaluation
Skills/Tech stack required:
[AI Deployment] [Attention] [C++] [CMake] [CNN] [CUDA] [Edge AI] [Edge AI deployment] [GDB] [Git] [Heterogeneous computing] [Inference Optimization] [Linux] [Model Quantization] [ONNX] [ONNX Runtime] [Operator optimization] [Perf] [Python] [PyTorch] [TensorRT] [Transformer]
Educational requirements:
[Bachelor's Degree]
Role(s):
[AI Inference Engineer] [Deployment Engineer] [Engineer] [Inference Engineer] [Model Deployment Engineer]