ML Engineer (Senior/Staff)
Summary
Senior/Staff ML Engineer owning the inference platform architecture—GPU scheduling, autoscaling, and model serving with vLLM/SGLang on Kubernetes—plus evaluation systems for LLM and agent changes at a climate-focused AI startup.
About AZX
Our mission is to accelerate positive impact in critical industries through AI transformation. We specialize in physics-informed ML and enterprise AI solutions that directly address climate and sustainability challenges.
We’re growing quickly and already work with category-leaders in real estate (CBRE), energy (LevelTen Energy), logistics (Flexe) and utilities.
We’re a public benefit corporation, founded in 2024, and have been profitable from inception.
We work on challenges in clean energy, decarbonization, climate risk, energy systems, and global economics. We’re building our company for long-term success and aim to create the ultimate place to work for those passionate about AI and making a positive impact.
About the Role
We're looking for an ML Engineer to own the technical backbone of how AZX serves and evaluates models at scale. This is a high-leverage IC role spanning our inference platform — GPU scheduling, autoscaling, and serving infrastructure for vLLM/SGLang across cloud and customer-managed clusters — and the evaluation systems that tell us whether model, prompt, and agent changes actually make things better.
You'll create technical direction for how AZX serves models reliably. This role suits someone who wants architectural ownership over hard ML infrastructure problems, paired with the judgment to build the guardrails that let the rest of the team move fast safely.
Responsibilities:
Own architecture for inference serving and GPU scheduling — Kubernetes operators, autoscaling, and dynamic capacity across vLLM/SGLang deployments on cloud and customer-managed infrastructure.
Design and calibrate eval systems for model, prompt, and agent changes, including golden datasets, LLM-as-judge pipelines, and regression gates wired into CI.
Advise on cost-aware model routing and cascading decisions, balancing latency, cost, and quality across providers and model tiers.
Apply physics-informed ML and enterprise AI expertise to the hardest client and platform problems, drawing on the team's research depth.
Set technical standards for ML infrastructure and evaluation practice across the org, and mentor engineers working in this space.
Partner closely with the inference platform, gateway, and evals-focused engineers to keep architecture coherent as the platform grows.
Core Qualifications:
3+ years of experience with ML infrastructure and inference serving — vLLM, SGLang, TensorRT-LLM, or comparable systems — at production scale.
Strong background in evaluation and reliability engineering for ML/LLM systems, or the seniority to build this practice from scratch.
Solid Kubernetes experience, ideally including GPU-specific scheduling constraints (node pools, autoscaling under GPU bottlenecks).
A track record of technical leadership at a staff or senior level — setting direction, not just executing tickets.
Research fluency is a plus (PhD, publications, or equivalent depth) given the technical bar of our existing ML team, though this is an infrastructure-and-systems role first.
Bonus Qualifications:
Advanced ML/AI frameworks and techniques (e.g., PyTorch Lightning, JAX, HuggingFace, ONNX optimizations)
Lower-level or performance-focused languages for ML acceleration (e.g., C++, Rust, CUDA)
Large-scale data and distributed training paradigms (e.g., Spark, Ray, Horovod, Dask)
Advanced data infrastructure (e.g., vector/graph databases, feature stores, data lakes)
Why AZX!
Be part of a fast-growing, profitable, mission-driven company with industry-leading clients tackling the massive opportunity of AI transformation in critical industries.
Competitive early-stage startup compensation (based on capabilities, experience, and location)
Bonus eligibility
Health insurance with meaningful coverage for dependents
Flexible paid time off
Equity
Fully remote culture with a cluster of teammates in Seattle
Additional Information:
Must be willing to travel to Seattle area for final interview and travel 2x/year for company summits
Applicants must be currently authorized to work in the United States on a full-time basis.
We are unable to sponsor or take over sponsorship of employment visas at this time.
Next Steps:
If this job sounds like a great fit but don’t check ALL of these qualification boxes, we’d still love to hear from you!
As published by ashby
Name, Email, Resume
- Primary Phone Number optional
- Are you currently located in North America? yes / no
- Do you currently, or will you in future, require visa sponsorship to work in the US? yes / no
- Do you have at least 7+ years of professional software engineering experience yes / no
- Have you led the technical design and implementation of a significant feature or system from conception to deployment? yes / no
- Do you have experience directly mentoring junior engineers on technical design or coding best practices? yes / no
- Are you comfortable working on projects where the initial requirements or solutions are highly ambiguous, requiring you to help define them? yes / no
- Describe a time you significantly improved the reliability, performance, or maintainability of an existing system. (1-2 sentences) written answer
- In your last role, how often did you interact directly with non-technical stakeholders (e.g., product managers, clients) to gather requirements or explain technical concepts? choose one
- In a high-throughput FastAPI application using PostgreSQL, you need to implement a feature that involves writing data to the database, but also performing a relatively slow, external API call (e.g., to a third-party service for enrichment or notification) that doesn't need to block the primary database write operation. Describe your approach to integrating this external API call. Specifically, discuss the technical patterns you would use within the FastAPI application and any considerations for ensuring data consistency or handling failures in the external call. written answer
- At AZX, we're building systems that matter in critical industries like energy, real estate, and infrastructure, often navigating ambiguous problems. Beyond the technical challenge, what aspects of solving these kinds of real-world, high-impact problems truly excite or motivate you? Feel free to share a brief anecdote or personal reflection. written answer
- Did anyone refer you to AZX? If so, who? optional