Tasks:
- * Adapt models to hardware
- * Co-optimize models, systems, and hardware
- * Deploy large language models on edge devices
- * Develop inference frameworks and optimize operators
- * Research and implement edge model optimization techniques
Perks/Benefits:
- + Potential conversion to a full-time position
Skills/Tech stack required:
[C#] [C++] [Continuous batching] [CUDA] [Heterogeneous systems] [Hexagon HVX] [Linux] [Llama.cpp] [LLM architectures] [MNN] [NPU Development] [ONNX Runtime] [OpenCL] [Python] [Qualcomm NPU] [Qualcomm NPU development] [Qualcomm QNN] [Quantization] [Speculative decoding] [TensorRT] [VLLM] [VLM architectures]
Educational requirements:
N/A
Role(s):
[Engineer Intern] [Intern] [Machine Learning Engineer] [Machine Learning Engineer Intern]