Lead AI Engineer (Inference Serving & Performance)
You will own the inference serving stack and define its technical direction and measurement roadmap. You will improve latency throughput memory efficiency and serving cost while maintaining model-quality standards. You will identify GPU bottlenecks, optimise kernels and memory usage, support distributed MoE serving, implement quantization techniques, build reproducible benchmarks, improve instrumentation, establish engineering standards, and guide engineers while remaining hands-on.
Responsibilities
- Own and improve the inference serving stack including prefill and decode disaggregation continuous batching KV-cache management cache-aware routing and speculative decoding.
- Improve throughput latency memory efficiency and serving cost while maintaining model-quality standards.
- Identify and eliminate GPU performance bottlenecks using CUDA Triton profiling kernel optimisation memory optimisation and precision strategies.
- Drive efficient serving of large Mixture-of-Experts models through expert parallelism all-to-all communication expert-load balancing and distributed execution.
- Implement and evaluate quantization and other model optimisation techniques against full-precision references.
- Own rigorous reproducible benchmarking across serving performance quality and cost.
- Build and improve instrumentation to understand system behaviour identify bottlenecks compare configurations and validate improvements.
- Make technical decisions around routing scheduling caching and workload distribution across diverse hardware environments.
- Define the AI serving and measurement roadmap and prioritise performance and infrastructure improvements.
- Establish technical standards for inference performance benchmarking reproducibility testing and production readiness.
- Guide and mentor engineers across serving metrics and testing while remaining involved in implementation and technical problem solving.
Requirements
- Strong hands-on experience building optimising or operating LLM inference-serving systems in production.
- Deep understanding of inference stacks such as vLLM SGLang TensorRT-LLM or similar technologies.
- Strong understanding of paged attention continuous batching disaggregated serving KV-cache management and modern inference architectures.
- Strong GPU performance engineering experience using CUDA or Triton.
- Strong understanding of GPU performance characteristics including memory bandwidth compute utilisation model bandwidth utilisation FLOPs utilisation and FP8 and FP4 precisions.
- Experience serving large Mixture-of-Experts models including expert parallelism all-to-all communication and expert-load balancing.
- Hands-on experience with model quantization and measuring quality against full-precision references.
- Strong benchmarking experience including load generation latency percentiles throughput-at-SLO quality measurement and reproducibility.
- Strong understanding of distributed systems particularly routing and scheduling workloads across varied hardware.
- Ability to analyse complex performance bottlenecks across models serving infrastructure networking and hardware.
- Demonstrated technical leadership experience leading an inference serving ML infrastructure or similarly complex engineering initiative.
- Ability to set technical direction while remaining hands-on with implementation and performance optimisation.
- Experience with speculative decoding and its behaviour under production load is a strong plus.
- Open-source contributions to inference or serving frameworks are a strong plus.
- Experience serving reasoning models long-context models or agentic multi-turn workloads is a plus.
- Experience with energy-aware serving including throughput-per-watt or power-constrained environments is a plus.
- Experience with confidential computing or data-residency-aware serving is a plus.
- Production experience with structured sparsity activation sparsity or KV compression is a plus.
- Advanced English level (C1).
Benefits
- 24 days annual leave plus public holidays
- Health insurance