Tasks:
- * Build unified inference runtime and serving framework
- * Design and implement embodied AI model inference architecture
- * Develop model conversion, benchmarking, validation, and deployment toolchains
- * Optimize model compression, quantization, accuracy, memory, and power
- * Schedule heterogeneous CPU, GPU, and NPU resources
Perks/Benefits:
Skills/Tech stack required:
[Accuracy Alignment] [C++] [Concurrency programming] [Concurrent execution] [CPU] [Embedded NPU] [Embedded NPU deployment] [GPU] [Heterogeneous computing] [Inference benchmarking] [Inference runtimes] [Linux systems] [Linux Systems Programming] [Memory Management] [Mixed Precision] [Mixed precision inference] [Model Conversion] [Model Deployment] [Model Inference] [Model inference runtimes] [Model Quantization] [Model Scheduling] [NPU] [NPU deployment] [Performance Profiling] [Precision Inference] [PTQ] [Python] [QAT] [Systems programming] [Transformer] [VLA] [VLM]
Educational requirements:
[Bachelor's Degree]
Role(s):
[AI Inference Architect] [Architect]