NPU Software Engineer SDK
You will design, build, and operate SDK release pipelines, port and optimize AI models for NPU hardware, integrate open-source inference and serving frameworks, and build verification and debugging utilities. You will author developer documentation and operate its publishing pipeline.
Responsibilities
- Design, build, and operate CI/CD pipelines for SDK releases.
- Port and optimize deep-learning models for NPU hardware.
- Analyze performance bottlenecks and validate numerical parity against GPU reference implementations.
- Integrate the SDK with inference and serving frameworks.
- Design model-feeding frameworks and debugging utilities.
- Author and maintain API references, tutorials, model-support matrices, and release notes.
- Operate the documentation build and publishing pipeline.
Requirements
- Master’s degree or higher in Computer Science, Electrical Engineering, or a related field.
- Knowledge of CI/CD tools including GitHub Actions, Airflow, and Buildkite.
- Knowledge of LLM, multimodal, vision, and speech model architectures.
- Knowledge of PyTorch internals, model customization, and graph transformations.
- Python and modern C++ development skills.
- Experience integrating Hugging Face Transformers, Diffusers, and vLLM.
- Experience deploying and troubleshooting AI inference workloads in Python and Kubernetes environments.
- Experience optimizing and deploying models with CUDA, TensorRT, TensorRT-LLM, MLIR, Triton, or TVM.
- Docker and Kubernetes ML infrastructure experience.
- Knowledge of JAX or TensorFlow.