freehire launches on Product Hunt on 26 August.

Follow →

Backend Engineer, AI

Summary

Builds and optimizes the backend inference and orchestration layer for an AI assistant, ensuring low-latency, reliable interactions between LLMs and end users.

More than five billion users rely daily on foundational applications—such as email clients, note-taking tools, and task managers—that remain fundamentally non-AI-native. Our mission is to bridge this gap by building a proactive, intelligent assistant designed for everyday users. We strive to infuse intelligence seamlessly into daily conversations, routine errands, personal organization, and complex workflows with minimal prompting overhead.

Our product architecture is engineered to achieve high reliability across long-running workflows, maintain persistent contextual awareness, and execute real-world tasks effectively. The system must navigate multi-step reasoning pathways, interface reliably with external software tools, and maintain fault tolerance despite non-deterministic model behaviours. Our core objective is to delight users by reducing task execution time by over 90% in their daily routines.

Role Overview:

You will own the critical inference and orchestration layer that powers every artificial intelligence interaction within our product ecosystem. Positioned directly between foundational models and end users, your engineering contributions would span latency optimization, system correctness, architectural reliability, and operational cost. Your impact would directly dictate the real-world user experience.

You will design, build, and operate production-grade systems capable of transforming raw model capabilities into high-performance, stable, and observable APIs consumed seamlessly across mobile and desktop client applications.

Core Responsibilities:

  • System Engineering & Operations: Architect and operate robust backend infrastructure that serves AI-powered features reliably within high-scale production environments.
  • Inference & Orchestration: Design advanced inference pipelines, orchestration layers, and clean service boundaries around large language models and multimodal systems.
  • Production Reliability: Take full ownership of production operational concerns, including comprehensive monitoring, distributed logging, proactive alerting, and incident response.
  • Performance Optimization: Continuously optimize latency, throughput, and resource utilization across model inference endpoints, intelligent caching layers, request batching, and real-time streaming protocols.

Technical Stack:

  • Languages: Python, Node.js
  • Machine Learning & Frameworks: PyTorch, OpenAI API, Anthropic API, Open-Source LLMs
  • Data Stores: SQL and NoSQL database systems
  • Infrastructure & Orchestration: Kubernetes, Docker

What you’ll bring:

Candidates should possess strong backend engineering expertise and experience working with AI inference systems. Key qualifications include:

  • Demonstrated excellence in backend software engineering within rigorous, production-grade environments.
  • Proven track record of designing, deploying, and operating low-latency, high-throughput distributed services.
  • Direct familiarity with modern AI inference patterns, including large language models (LLMs), dense vector embeddings, and multimodal processing pipelines.
  • Comfort and proficiency in diagnosing and debugging complex distributed systems operating under heavy production load.
  • A strong bias toward rapid shipping, iterative deployment, and rigorous learning derived from real-world production feedback.

See also