Senior Engineer AI
You will own and evolve production AI applications, working across backend, frontend and inference services. You will design and deploy RAG systems, build agentic workflows, write and refine prompts, optimise inference costs, and run evaluations. You will deploy and operate services on Kubernetes and integrate AI features with third-party platforms.
Responsibilities
- Design, build and maintain production AI applications end-to-end (backend, frontend, inference services)
- Architect RAG systems using vector databases, embedding models and chunking strategies optimized for accuracy and latency
- Build agentic workflows with tool/function calling, multi-step reasoning and structured output parsing
- Write and iterate system prompts, few-shot examples and prompt chains to maximise output quality
- Implement function calling, tool-use patterns and structured JSON/XML output handling across provider models
- Drive cost optimization through model selection, caching, token budgeting and request batching
- Build and maintain evaluation frameworks to measure accuracy, relevance, hallucination rates and regressions using observability tools
- Work with message queues (RabbitMQ), caching layers (Redis) and relational databases (PostgreSQL)
- Deploy and manage AI services on Kubernetes with CI/CD pipelines on AWS and GCP
- Integrate AI capabilities with third-party platforms such as Telegram bots and chat widgets
- Contribute to architectural decisions regarding model selection, hosting approaches and build-vs-buy trade-offs
Requirements
- 5+ years shipping production software systems
- 2 years building AI/LLM-powered applications end-to-end with real users and volume
- Strong experience with RAG architectures including vector databases, embedding models and chunking/indexing strategies
- Deep understanding of LLM capabilities and limitations including prompt engineering, function/tool calling and context window management
- Experience with LLM provider APIs and abstraction layers such as OpenAI, Anthropic, LiteLLM or OpenRouter
- Proficiency in Python (Flask/FastAPI) and/or Node.js/TypeScript (Next.js, Vercel AI SDK); Golang experience is a plus
- Hands-on experience building evals, tracking quality metrics and debugging nondeterministic outputs in production
- Familiarity with cost optimisation techniques including model routing, caching and token usage monitoring
- Solid fundamentals in data structures, algorithms and system design
- Experience with containerised deployments (Docker, Kubernetes) and cloud platforms (AWS/GCP)