ML Infrastructure Engineer
You will build and scale GPU compute infrastructure, training frameworks, experimentation tools, and data pipelines. You will integrate large-scale training and inference systems, work with ML teams to productionize models, ensure platform reliability and efficiency, solve full-stack problems, and mentor junior engineers.
Responsibilities
- Design, build, and scale GPU compute infrastructure, training frameworks, and experimentation tools
- Develop data pipelines and integrate large-scale data, training, and inference systems
- Collaborate with ML teams to productionize models
- Ensure scalability, reliability, and efficiency of machine learning systems
- Solve full-stack problems independently
- Mentor junior engineers
Requirements
- Bachelor's, Master's, postgraduate degree or PhD in computer science, machine learning, another quantitative discipline, or equivalent work experience
- 2+ years of industry experience with high-traffic or large-scale production environments, distributed systems, GPU infrastructure, or deep learning applications
- 2+ years of experience with ML platforms, training infrastructure, or collaboration with modeling engineers and data scientists
- Strong proficiency with Python
- Experience with C++ or Rust
Benefits
- Equity
- Medical coverage
- Vision coverage
- Dental coverage
- 401(k) retirement plan
- Short-term disability insurance
- Long-term disability insurance
- Life insurance
- Discounts and perks
