Backend Engineer - Voice AI Platform
Summary
Build real-time voice AI pipelines and backend services for Hams.ai’s voice agents, handling speech-to-text, LLM orchestration, telephony, and low-latency WebRTC sessions using Python, FastAPI, PostgreSQL, and Kubernetes.
As a Backend Engineer you’ll build the core systems behind our voice AI agents. You’ll architect the real‑time pipelines that process live speech, orchestrate LLM‑powered conversations, manage concurrent voice sessions at scale and ensure every call is fast, reliable and intelligent. This means working hands‑on with real‑time audio processing, speech‑to‑text and text‑to‑speech pipelines, LLM orchestration, telephony (SIP, Asterisk) and distributed backend infrastructure. You’ll work daily with Python, FastAPI, Pipecat, LiveKit, WebRTC, Celery, PostgreSQL, Redis and Kubernetes.
What You’ll Do
- Build and optimize real‑time voice AI pipelines (speech recognition, STT, LLM processing and speech synthesis, TTS) running in sub‑second loops.
- Design and maintain the backend services that manage voice‑agent sessions, call routing and telephony integration (SIP, Asterisk).
- Architect multi‑agent orchestration systems, conversation flows, agent hand‑offs and context passing between voice AI agents.
- Build and scale the infrastructure for handling thousands of concurrent voice calls with low latency.
- Develop and optimize WebSocket and WebRTC‑based real‑time communication layers.
- Implement and manage distributed task processing for batch calling campaigns, call analytics and post‑call AI analysis.
- Architect, optimize and maintain PostgreSQL databases, Redis caching and message queues.
- Implement observability across voice pipelines: latency tracking, call quality metrics, distributed tracing (OpenTelemetry), Sentry.
- Handle production debugging of real‑time voice systems, diagnosing audio quality issues, latency spikes and session failures.
- Work closely with AI, product and DevOps teams to ship voice AI features end‑to‑end.
Who You Are
- Strong proficiency in Python with deep understanding of async/await and real‑time concurrency patterns.
- Experience building production backend systems with FastAPI or similar async frameworks.
- Deep understanding of PostgreSQL, relational database design and ORM patterns.
- Experience with real‑time systems (WebSockets, streaming, audio pipelines or low‑latency communication).
- Hands‑on experience with distributed task processing (Celery, Redis).
- Comfortable owning production systems end‑to‑end from design to deployment to incident response.
- Thrives in fast‑paced, high‑ownership environments where voice AI is the core product.
Nice to Have
- Experience with voice AI, conversational AI or speech processing (STT, TTS, VAD).
- Experience with telephony systems (SIP, Asterisk, WebRTC).
- Familiarity with AI LLM orchestration (OpenAI, LangChain, LangGraph, Pipecat, LiveKit).
- Experience with real‑time audio frameworks or voice bot platforms.
- Cloud infrastructure experience (AWS, GCP, Kubernetes).
- Vector databases and RAG patterns for knowledge‑powered voice agents.