Senior Software Engineer (Performance)

Summary

Senior engineer optimizing large-scale LLM inference systems for speed, scalability, and efficiency across the full stack, from kernels to distributed execution.

Job Description:

We are looking for a Senior Inference Engineer with a strong foundation in software engineering, distributed systems, and performance optimization to build and optimize inference engines for large-scale LLM serving systems. You will work across both research and production environments, ensuring our LLM serving systems are fast, scalable, and efficient. The role spans the entire inference stack — from kernel and runtime to scheduling, memory management, and distributed execution


Key Responsibilities:

  • Profile, benchmark, and analyze bottlenecks for LLM inference workloads across multiple layers: kernel, memory, networking, and scheduler
  • Optimize inference engines (vLLM, SGLang, TensorRT-LLM) for throughput, latency, memory efficiency, GPU utilization, and cost
  • Implement and fine-tune inference optimization techniques including batching, KV-cache management, quantization, speculative decoding, parallelism strategies, and disaggregated serving
  • Build instrumentation and profiling tools to identify bottlenecks
  • Ensure the reliability of the inference pipeline through A/B launches, rollback, model versioning, and fault tolerance
  • Collaborate with the Platform Engineering team to improve serving architecture based on performance findings
  • Document and share knowledge, contributing to internal best practices and AI open-source projects whenever possible

Requirements

1 - Mandatory:

  • At least 5 years of experience as a Software Engineer, Performance Engineer, or equivalent.
  • Strong foundation in Software Engineering, Software Architecture, and Distributed Systems.
  • Proficiency in at least one of the following languages: Python, Go, or C++.
  • Experience developing or optimizing distributed systems, high-throughput backends, or large-scale serving systems.
  • Experience with benchmarking, profiling, and performance tuning in production environments.
  • Ability to analyze CPU, Memory, Network, or Storage bottlenecks.
  • Strong systems thinking, Root Cause Analysis capabilities, and the ability to solve complex performance problems.
  • Strong ownership mindset and the ability to work independently.

2 - Nice to Have:

  • Experience with Linux internals, kernel tuning, or custom Linux kernel.
  • Understanding of GPU Architecture or CUDA Programming.
  • Experience with AI/ML Serving Systems or LLM Inference.- Have worked with one of the inference engines such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server.
  • Understanding of batching, KV Cache, quantization, speculative decoding, tensor/pipeline parallelism, or disaggregated serving.
  • Experience with the NVIDIA inference stack (TensorRT, Triton, CUTLASS, NCCL, cuBLAS, cuDNN).
  • Experience with observability stacks such as Prometheus, Grafana, or OpenTelemetry.
  • Open-source contributions or research related to AI Infrastructure, ML Systems, or Performance Optimization.