Software Engineer
Summary
Build and maintain the infrastructure for AI-powered audit agents, focusing on evaluation, observability, and scalable data pipelines in Python and GCP.
Project details
The project focuses on building an AI-powered platform that enables production-grade LLM agents for
complex audit workflows. It combines intelligent agents, evaluation pipelines, and observability tooling to
continuously improve quality, reliability, performance, and cost efficiency in production. We're looking for an
experienced engineer with strong data engineering and AI systems expertise to build the infrastructure that
evaluates, monitors, and optimizes AI agents at scale. You'll work across backend services, data pipelines,
analytical databases, and agent architecture, designing systems for evaluation, experimentation, logging,
tracing, and debugging. Many of the challenges involve building reliable infrastructure around
non-deterministic AI systems with no established blueprint to follow. You should be comfortable working in
complex backend environments, making sound architectural decisions, and delivering scalable solutions
that help production AI systems continuously improve.
Key responsibilities
You'll help design, build, and operate the infrastructure that supports our production AI agents, with a strong
emphasis on data platforms, evaluation, observability, and continuous optimization. Your responsibilities will
include:
- Developing online and offline evaluation frameworks for LLM agents using benchmark datasets, ground-truth labels, human feedback, and experimental results.
- Implementing automated validation pipelines that verify changes to prompts, models, context, and agent behavior before deployment.
- Investigating large-scale execution logs and agent traces to uncover quality issues, performance bottlenecks, reliability problems, and opportunities to reduce inference costs.
- Designing and maintaining data solutions built on analytical databases and column-oriented storage technologies such as BigQuery, ClickHouse, or comparable platforms.
- Building systems for long-term storage, replay, and analysis of production agent activity to support debugging and continuous improvement.
- Creating observability capabilities, including monitoring, trace inspection, logging, dashboards, and debugging tools for AI services.
- Contributing directly to the backend platform and agent ecosystem by enhancing existing AI agents and developing new capabilities when required.
Requirements
- Solid experience developing backend applications in Python or similar server-side technologies.
- Confident writing advanced SQL queries and working with large-scale datasets.
- Experience in deploying and maintaining cloud-based applications, preferably on Google Cloud Platform.
- Experience designing and implementing data pipelines, ETL/ELT processes, streaming systems, or production feedback loops.
- Familiarity with analytical databases, data warehouses, columnar storage solutions, and processing high-volume event or trace data.
- Strong understanding of distributed systems, observability, monitoring, logging, debugging, and the trade-offs involved in building reliable software.
- Ability to quickly navigate complex codebases and gain a deep understanding of existing architectures.
- Senior engineering judgment by making sound architectural decisions, evaluating technical trade-offs, and building solutions that are easy to maintain and extend.
- Ability to work in ambiguous environments, approach problems from first principles, and enjoy building infrastructure that powers production AI systems.
Nice-to-haves
- Hands-on experience developing infrastructure for LLM-powered applications or AI agents, including prompt optimization, model selection, context management, or inference optimization.
- Experience analyzing production traces from distributed applications.
- Building internal developer platforms or tooling for engineering, operations, or business teams.
- Experience with workflow orchestration platforms such as Temporal, Airflow, or similar technologies.
- Familiarity with highly regulated industries such as finance, compliance, risk, or audit.
- Previous experience working in an early-stage startup or another fast-paced engineering environment.