AI Backend Engineer - Audena AI
Remote · Full-time · Engineering
Audena AI is building a multi-tenant voice-AI platform that enables businesses to run AI agents over their own phone numbers. Our platform connects live telephony with LLMs to power real-time conversations, intelligent agents, retrieval, integrations, billing, and customer operations—with a target of under 800ms voice-to-voice latency.
We're looking for an AI Backend Engineer to build and own the Python services behind this platform.
What You'll Do
Build and maintain high-performance Python/FastAPI backend services.
Develop real-time voice pipelines using WebSockets, streaming, telephony providers, and live LLM APIs.
Build AI agent infrastructure including tool calling, structured outputs, prompt execution, and fallback handling.
Develop retrieval and RAG systems using PostgreSQL, pgvector, BM25, and LangGraph.
Build and maintain multi-tenant infrastructure covering authentication, tenant isolation, credentials, rate limits, billing, and integrations.
Develop reliable async workers and background systems for campaigns, webhooks, synchronization, and event processing.
Design APIs, database schemas, migrations, and integrations with external platforms.
Implement observability, testing, evaluation, and performance monitoring for production AI systems.
Take ownership of production issues involving latency, concurrency, reliability, and data isolation.
What We're Looking For
Required
2–4 years of production backend development experience with Python.
Strong understanding of async Python and event-driven systems.
Experience with FastAPI, Pydantic, PostgreSQL, and REST APIs.
Experience with WebSockets, streaming, or other long-lived connections.
Production experience integrating LLM APIs, including streaming, tool/function calling, structured outputs, retries, and fallbacks.
Strong database fundamentals—queries, indexing, execution plans, and migrations.
Experience writing unit and integration tests.
Comfortable working with mypy, Ruff, and strongly typed codebases.
Ability to reason about system performance, concurrency, failure modes, and security.
Nice to Have
Experience with voice AI, telephony, VoIP, WebRTC, VAD, or real-time audio.
Experience with Twilio, Plivo, Vonage, Gemini Live, or OpenAI Realtime.
Experience with RAG, pgvector, BM25, hybrid search, or LangGraph.
Experience building multi-tenant SaaS platforms.
Knowledge of Redis, distributed locks, idempotency, and background workers.
Experience with Docker, AWS, Terraform/OpenTofu, and GitHub Actions.
Familiarity with LLM evaluation and prompt-injection/untrusted-input protection.
Tech Stack
Python 3.12 · FastAPI · Pydantic v2 · PostgreSQL 16 · pgvector · asyncpg · Redis · WebSockets · Twilio · Plivo · Vonage · Gemini Live · OpenAI Realtime · LangGraph · Docker · AWS · Terraform/OpenTofu · GitHub Actions · pytest · mypy · Ruff
This Is Not a Prompt-Engineering Role
This role is about building production AI infrastructure.
The core challenges are latency, concurrency, real-time communication, failure handling, model reliability, distributed systems, and multi-tenant data isolation.
You'll be expected to understand not only how to integrate an LLM, but also what the model sees, what it can do, how its output is validated, and what happens when it is wrong or unavailable.
Why Audena AI?
You'll work on a real production voice-AI platform across:
Real-time voice AI
LLMs and AI agents
RAG and retrieval
Telephony
Distributed async systems
Multi-tenant SaaS
Billing and integrations
Production observability and evaluation
This is a small-team, high-ownership engineering role where your work directly impacts live AI conversations and production reliability.
