Active listing
Senior Quantization Engineer
About the role
Edge AI Model Optimization team at NXP focuses on advancing on‑device intelligence by optimizing CNN, LLM and VLM models for the Ara2 NPU family. The Senior Quantization Engineer researches state‑of‑the‑art quantization and compression techniques, prototypes and integrates them into production C++/Python pipelines, and guides cross‑functional teams on performance trade‑offs. Hyderabad, on‑site, hybrid schedule possible with relocation support.
What you’ll do
- Survey latest model compression research
- Develop and adapt quantization methods for NXP hardware
- Implement production‑grade optimization code in C++/Python
- Document trade‑offs and create deployment recipes
- Mentor engineers on numerical methods
- Collaborate across AI research, hardware, and software teams
- Contribute to IP via patents and publications
What you’ll bring
- MSc or Ph.D in Computer Science, EE or Mathematics
- 3+ years AI/ML experience with CNNs and Generative AI
- Proficient in PyTorch, ONNX, Python and C++
- Experience with embedded constraints (latency, power, memory)
Nice to have
- Published research or patents in model optimization
- Familiarity with NPUs and hardware profiling
Skills
Benefits
- Health insurance
- Retirement plan
- Paid time off