Point your AI agent at freehire and let it find you a job.

Get the CLI →

Rebellions

NewBe an early applicant

NPU Software Engineer Runtime

Posted
Discussion

You will build runtime modules that connect compilers, drivers, and ML frameworks for model deployment. You will support PyTorch execution, develop profiling capabilities, extend vLLM, optimize multi-NPU distributed inference, benchmark performance, and help deploy scalable inference services.

Responsibilities

  • Design and implement runtime modules that interface with compilers and drivers
  • Maintain native PyTorch execution support, torch.compile integration, and compiler toolchains
  • Develop a user-facing performance profiler for the SDK
  • Extend vLLM to improve NPU inference performance
  • Design and optimize distributed multi-NPU inference and collective communication
  • Benchmark, profile, and optimize runtime-system performance
  • Deploy and scale inference services with ML and infrastructure engineers

Requirements

  • Over 5 years of software-engineering experience with ML frameworks, inference runtimes, or AI accelerator toolchains
  • Bachelor’s degree or higher in Computer Science, Electrical Engineering, or a related field
  • Strong proficiency in C++ and Python
  • Understanding of deep learning, LLM architectures, generative AI, and inference optimization
  • Experience with LLM serving frameworks such as vLLM or TensorRT-LLM
  • Understanding of tensor parallelism, KV-cache optimization, and memory-efficient execution
  • Familiarity with compilers, runtimes, drivers, firmware, and hardware acceleration
  • Debugging and performance-profiling skills for high-throughput inference
  • Written and verbal communication skills

Skills

Apply

See also

Software Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available