freehire launches on Product Hunt on 26 August.

Follow →

Lead Backend & Data Engineer

Open 25d

Summary

Lead a team building the backend and data infrastructure for a zero-hallucination AI RAG platform, using Python, Apache Beam, FastAPI, and Google Cloud Spanner.

This is us

At Avenga, we believe that human creativity empowers technology that matters. Operating globally, our 6000+ specialists provide a full spectrum of services, including business and tech advisory, enterprise solutions, CX, UX and Ul design, managed services, product development, and software development.

This is the job

We are looking for a Lead Backend & Data Engineer to join an innovative AI initiative focused on building a highly deterministic, zero-hallucination Dual-Engine Retrieval-Augmented Generation (RAG) platform. You will be building the foundational data infrastructure and API layer that powers a next-generation zero-hallucination RAG platform. Your work will enable downstream AI agents to query a deterministic "Truth Engine" — fusing vector search with strict physical topologies to eliminate hallucinations at the data layer.

The core focus of this role is service engineering - building, deploying, and scaling microservices and API layers. While familiarity with graph technologies is a plus, the team already has strong collective knowledge graph expertise, so your primary responsibility will be designing and implementing robust, scalable backend services.

This is you

  • 7+ years in backend engineering, data engineering, or distributed systems development

  • Extensive experience building and deploying microservices and API layers

  • Strong proficiency in Python and modern backend development (FastAPI, Pydantic)

  • Hands-on experience with containerization and orchestration (Docker, Kubernetes)

  • Experience designing and deploying cloud-native applications on GCP (Cloud Run, Cloud Functions)

  • Deep expertise in Apache Beam and distributed data processing concepts, including windowing, watermarks, and idempotent writes

  • Strong experience with Google Cloud Spanner or another distributed SQL database

  • Advanced proficiency in asynchronous Python development using FastAPI and Pydantic

  • Hands-on experience with workflow orchestration platforms such as Temporal or Cadence

  • Strong understanding of deterministic workflow execution and separating non-deterministic operations into Activities

  • Experience implementing retries, compensation logic, and state management in distributed systems

  • Solid understanding of distributed systems architecture and cloud-native application design

  • Strong problem-solving skills with the ability to drive technical decisions

  • Upper-Intermediate or higher level of English

Nice-to-have skills:

  • Working knowledge of graph query languages such as ISO GQL, Cypher, or Gremlin

  • Experience with LangChain / LangGraph agent payloads

  • Understanding of RAG architectures and retrieval systems

  • Experience with Dataflow (Google Cloud's managed Beam runner)

  • Degree in Computer Science, Engineering, or a related field

This is your role

  • Design, build, and deploy scalable microservices and API layers using FastAPI and Python

  • Design, build, and optimize Apache Beam data ingestion pipelines running on multiple execution engines such as Dataflow and Flink

  • Develop scalable backend services and optimize Google Cloud Spanner performance through advanced gRPC connection pooling

  • Define robust API contracts with Pydantic to support seamless integration with downstream AI orchestration frameworks such as LangChain and LangGraph

  • Design backend architectures that efficiently separate synchronous request handling from long-running asynchronous processing

  • Develop and maintain Workflow Definitions, Activity Definitions, and Temporal clients using the Temporal Python SDK

  • Deploy and manage a horizontally scalable fleet of Temporal Workers responsible for complex database operations and AI routing logic

  • Implement resilient distributed workflows with automated retries, compensation mechanisms, and reliable state management

  • Containerize services using Docker and orchestrate deployments on Kubernetes and GCP Cloud Run

  • Structure and optimize multi-hop graph queries using ISO GQL to trace physical and logical data provenance (secondary)

  • Collaborate closely with AI, Platform, and Data Engineering teams to deliver scalable, production-ready solutions

  • Contribute to architectural decisions, code reviews, and engineering best practices while ensuring high standards of performance, reliability, and maintainability

What awaits you at Avenga?

At Avenga, everyone matters. We provide equal opportunities in recruitment, career development, and leadership, regardless of race, ethnicity, gender identity, sexual orientation, disability, age, religion, or any other characteristic. We are committed to fostering a work environment where our diverse community of employees, candidates, and business partners actively shapes our growth. By bringing together people from different backgrounds and experiences, we build a workplace where everyone feels free to be themselves while honoring the boundaries of others.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available