Software Engineer (Backend-Focused)
Summary
Build and maintain backend services for an LLM gateway (routing, rate limiting, key management, observability) and support Kubernetes-based platform services, using Go, Rust, or async Python with Envoy, gRPC, and OpenTelemetry.
Software Engineer (Backend-Focused)
About AZX
Our mission is to accelerate positive impact in critical industries through AI transformation.
We’re growing quickly and already work with category-leaders in real estate (CBRE), energy (LevelTen Energy), logistics (Flexe) and utilities (can’t say which large utility yet).
We’re a public benefit corporation, founded in 2024 and have been profitable from the beginning (bootstrapped with consulting).
We work on challenges in clean energy, decarbonization, climate risk, energy systems and global economics. We’re building our company for long term success and aim to build the ultimate place to work if you’re passionate about AI and positive impact.
About the Role
We're looking for a Staff or Senior ML Engineer to own the technical backbone of how AZX serves and evaluates models at scale. This is a high-leverage IC role spanning our inference platform — GPU scheduling, autoscaling, and serving infrastructure for vLLM/SGLang across cloud and customer-managed clusters — and the evaluation systems that tell us whether model, prompt, and agent changes that make things better.
You'll create technical direction for how AZX serves models reliably. This role suits someone who wants architectural ownership over hard ML infrastructure problems, paired with the judgment to build the guardrails that let the rest of the team move fast safely.
What you will do
You will work on software projects in client engagements, and over time, internal platform capabilities.
You will:
Build and maintain backend services for our LLM gateway — routing, rate limiting, key management, and observability in front of the inference fleet.
Contribute to sandboxing and isolation infrastructure that keeps agent-generated code safe to execute, working alongside our security-focused engineers.
Support Kubernetes-based platform services, including operators and autoscaling logic adjacent to our inference platform.
Write high-performance backend code in Go, Rust, or async Python (FastAPI/Starlette), working with infrastructure like Envoy and gRPC.
Instrument services with OpenTelemetry so behavior, latency, and cost stay observable as the platform scales.
Collaborate across the gateway, sandbox, and inference platform teams, flexing across areas as priorities shift.
Core Qualifications - Technical and foundational
4+ years of experience in backend engineering fundamentals: distributed systems, API design, and production experience in Go, Rust, or async Python.
Familiarity with LLM-specific backend concerns (rate limiting, caching, token accounting) is a plus, though not required on day one.
Exposure to Kubernetes and containerization; interest in sandboxing or security is a plus.
Comfort working across a range of platform concerns rather than one narrow specialty — this role is intentionally broader than our specialist infra profiles.
Eagerness to grow into deeper specialization in gateway, sandbox, or inference infrastructure over time.
Values and Culture Qualifications
High emotional intelligence and a learning mindset
Strong collaboration skills
Enjoy others' success and a fun, positive environment.
Comfortable making decisions in the face of ambiguity and course correcting as needed.
Bonus Qualifications (not required but a huge plus)
Experience in both startup and enterprise environments
Past work in energy, real estate, utilities, climate or related fields
Bonus if you have experience and passion in one or more of
Additional web frameworks (e.g. Svelte, Vue, Angular)
Lower-level languages e.g. C++, Rust
Networking paradigms e.g. GraphQL, Websockets
ML capabilities e.g. Sk-learn, xgboost, Pytorch/Tensorflow/JAX, Onnx…
Additional database types such as graph or vector databases
DevOps e.g. CI/CD pipelines, Docker, Kubernetes, Terraform, Pulumi and/or Bicep
Generative AI e.g. prompt engineering, RAG, fine-tuning, tooling ecosystem
Compensation & benefits
Competitive early-stage startup compensation (based on capabilities, experience, and location)
Bonus eligibility
Health insurance with meaningful coverage for dependents
Flexible paid time off
Equity
Fully remote culture with a cluster of teammates in Seattle
Training and learning opportunities
Be part of a fast-growing, profitable, mission-driven company with industry leading clients tackling the massive opportunity of AI transformation in critical industries
Logistics
Remote but only USA/Canada
Must be willing to travel to Seattle area for final interview and travel 2x/year for company summits
Applicants must be currently authorized to work in the United States on a full-time basis.
We are unable to sponsor or take over sponsorship of employment visas at this time.
Next steps
If this job sounds great, we’d love to hear from you. If you feel aligned to the company but don’t check all these boxes, we’d still love to hear from you!
As published by ashby
Name, Email, Resume
- Primary Phone Number optional
- Are you currently located in North America? yes / no
- Do you currently, or will you in future, require visa sponsorship to work in the US? yes / no
- Do you have at least 7+ years of professional software engineering experience yes / no
- Have you led the technical design and implementation of a significant feature or system from conception to deployment? yes / no
- Do you have experience directly mentoring junior engineers on technical design or coding best practices? yes / no
- Are you comfortable working on projects where the initial requirements or solutions are highly ambiguous, requiring you to help define them? yes / no
- Describe a time you significantly improved the reliability, performance, or maintainability of an existing system. (1-2 sentences) written answer
- In your last role, how often did you interact directly with non-technical stakeholders (e.g., product managers, clients) to gather requirements or explain technical concepts? choose one
- In a high-throughput FastAPI application using PostgreSQL, you need to implement a feature that involves writing data to the database, but also performing a relatively slow, external API call (e.g., to a third-party service for enrichment or notification) that doesn't need to block the primary database write operation. Describe your approach to integrating this external API call. Specifically, discuss the technical patterns you would use within the FastAPI application and any considerations for ensuring data consistency or handling failures in the external call. written answer
- At AZX, we're building systems that matter in critical industries like energy, real estate, and infrastructure, often navigating ambiguous problems. Beyond the technical challenge, what aspects of solving these kinds of real-world, high-impact problems truly excite or motivate you? Feel free to share a brief anecdote or personal reflection. written answer
- Did anyone refer you to AZX? If so, who? optional