Backend Software Engineer
Summary
Build and scale the core AI inference and orchestration layer that powers user interactions, ensuring low-latency, high-reliability systems handling LLM and multimodal workflows.
Our client is seeking a high-caliber Backend Engineer to own the critical inference and orchestration layer that powers every single AI interaction in the platform. You’ll sit right between state-of-the-art models and millions of user interactions, where latency, correctness, and absolute reliability determine product success.
What You’ll Be Doing
Inference & Orchestration: Design, build, and scale the core orchestration layers, service boundaries, and inference pipelines that serve production AI features across mobile and desktop apps.
Tame Non-Deterministic AI: Build high-reliability, long-running workflows that handle multi-step reasoning, external tool interaction, and persistent context despite non-deterministic model behavior.
Performance Optimization: Drive low-latency and high-throughput across inference, caching, batching, and streaming strategies.
Production Excellence: Own end-to-end operational quality—from monitoring, logging, and alerting to rapid incident response for high-traffic API endpoints.
What We’re Looking For
Solid Backend Fundamentals: Proven track record building and operating high-throughput, low-latency production services.
AI/LLM Familiarity: Practical experience with AI inference patterns (LLMs, embeddings, multimodal systems) and integrating providers like OpenAI, Anthropic, or open-source models.
Distributed Systems Mastery: Comfortable debugging distributed backend architectures under heavy load.
-
Product-Minded Execution: A strong bias toward shipping fast, learning from production metrics, and taking full ownership of your systems.
Tech Stack
Languages: Python, Node.js
AI/ML: PyTorch, OpenAI / Anthropic APIs, Open-Source LLMs
Data: SQL & NoSQL databases
Infra & Ops: Kubernetes, Docker, modern cloud architecture
