Staff MaaS Backend Engineer
Summary
Architect and build a globally distributed Model-as-a-Service platform using Go, Kubernetes, PostgreSQL, Redis, and Kafka—handling API compatibility, token metering, multi-tenant security, and production reliability for a crypto infrastructure company.
You will co-own the architecture and implementation of a globally distributed Model-as-a-Service platform. You will develop Go services and compatible APIs, optimize model serving performance, manage Kubernetes-based deployments, enforce multi-tenant security, build accurate token metering and billing, execute safe migrations, and own reliability, observability, SLOs, and on-call operations.
Responsibilities
- Co-own the end-to-end MaaS system design and document technical decisions.
- Set Go and API standards and mentor engineers.
- Own OpenAI and Anthropic API compatibility including streaming tool calling and structured output.
- Develop intelligent routing with circuit breaking fallback and versioned APIs.
- Optimize token throughput KV caching and time-to-first-token latency.
- Integrate model deployment tooling LoRA multiplexing and autoscaling with Kubernetes.
- Scale regional inference pools and active-active control planes.
- Define and defend service level objectives and run load tests.
- Implement fail-closed authorization multi-tenant isolation rate limits quotas and abuse controls.
- Manage API keys and OAuth credentials.
- Build exactly-once token metering usage ledgers and billing reconciliation.
- Execute incremental migrations using shadow traffic and dual writes.
- Deliver request tracing cost telemetry and strict log hygiene.
Requirements
- 8+ years of backend engineering experience.
- 3+ years owning a high-traffic multi-tenant API platform for paying customers.
- Experience with multi-region active-active systems caching backpressure and performance engineering.
- Experience with zero-downtime migrations for stateful metering or ledger systems.
- Deep proficiency with Go-based services.
- Production Kubernetes experience including Envoy and GPU-aware scheduling.
- Experience with PostgreSQL Redis and Kafka.
- Experience with OpenTelemetry and high-cardinality analytics stores.
- Understanding of LLM serving including server-sent events KV caching and throughput optimization.
- Experience building exactly-once metering and billing systems.
- Experience designing fail-closed authorization and cross-tenant isolation.
- Experience with on-call operations and incident reviews.