Software Engineer - Training/Inference (C++)
You will design and optimize large-scale model-serving systems from distributed infrastructure through low-level GPU optimizations. You will improve inference latency, throughput, reliability, deployment, and testing infrastructure.
Responsibilities
- Architect and implement scalable distributed infrastructure for model serving.
- Optimize inference latency and throughput under production workloads.
- Build reliable, high-concurrency serving systems.
- Benchmark, fine-tune, and accelerate inference engines.
- Develop tools to trace, replay, and resolve full-stack issues.
- Create CI/CD infrastructure for endpoint deployment, image publishing, and inference-engine updates.
- Accelerate research on test-time compute, RL rollout, and model-hardware co-design.
Requirements
- Deep low-level systems programming in C/C++ or Rust.
- Experience with large-scale high-concurrency production serving.
- Experience with GPU inference engines such as vLLM, SGLang, Triton, or TensorRT-LLM.
- Strong background in batching, caching, load balancing, and parallelism.
- Experience with GPU kernels and code generation.
- Knowledge of quantization, speculative decoding, distillation, and low-precision numerics.
- Experience testing, benchmarking, and ensuring reliability of inference services.
- Experience designing and implementing CI/CD infrastructure for inference.
Benefits
- Equity
- Medical coverage
- Vision coverage
- Dental coverage
- 401(k) retirement plan
- Short-term disability insurance
- Long-term disability insurance
- Life insurance
- Discounts and perks