Point your AI agent at freehire and let it find you a job.

Get the CLI →

Rapsodo

NewBe an early applicant

Senior Data Engineer

Posted Updated
Discussion

Summary

Senior Data & AI Engineer who owns Rapsodo's production cloud data platform in Kuala Lumpur — the warehouse, multi-layer star schema, pipelines, and company-wide BI — and also designs and ships agentic AI workflows (LLM, RAG, tool-calling) on top of it. Core stack: SQL, Python, a major cloud (GCP/AWS/Azure), IaC, Kubernetes, dbt/Dataform-style transformations, and LLM APIs.

About Rapsodo

Rapsodo is a Sports Technology company with offices in the USA, Singapore, Turkey, Malaysia & Japan. We build data-driven sports analytics products used by athletes worldwide — from Major League Baseball players to golf enthusiasts — and by the coaches who train them. We are looking for team players who will help us deliver state-of-the-art solutions as part of Team Rapsodo.

Overview

As a senior Data & AI Engineer at Rapsodo, you will own the data platform that powers analytics, reporting, and decision-making across the company — and you will drive development of the agentic workflows built on top of it. This is a data engineer's role first: the custodianship of a production-grade warehouse spanning multiple domains, a multi-layer star schema, and a BI platform serving users company-wide. On top of that foundation, you will lead the design and delivery of agentic workflows that turn trusted data into informed, reliable action.

What You’ll Do

Own and Operate the Cloud Data Platform

  • Maintain cloud analytics infrastructure as code, treating the IaC repository as the source of truth for every cloud resource.
  • Operate a modern data stack spanning workflow orchestration, container management, ingestion, BI, and serverless compute; lead upgrades and patches with minimal disruption.
  • Operate the company-wide BI platform, including SSO integration and supporting dashboard workflows at scale.
  • Optimize warehouse cost and performance through partitioning, clustering, and query tuning; set up alerting across pipelines, connectors, and scheduled queries.
  • Enforce least-privilege access, rotate credentials on schedule, maintain backups, and keep warehouse documentation current.
  • Act as the final word on data quality and trust — every agent, dashboard, or automated workflow built on the warehouse is only as reliable as the custodianship behind it.

Build and Evolve Data Pipelines & Integrations

  • Develop transformation repositories on a multi-layer (raw → staging → curated) star schema, with incremental hash-based loading that keeps refreshes performant on very large datasets.
  • Author and maintain orchestration DAGs covering ingestion, transformation, retries, scheduling, and alerting; diagnose incidents with strong root-cause discipline.
  • Build and maintain ingestion for real-time and batch sources — database CDC, ERP, identity provider syncs, and SaaS connectors across e-commerce, payments, support, marketing, and marketplaces.
  • Lead new ingestion projects end-to-end (e.g. product event analytics, device telemetry) and drive the design of a unified semantic layer across product, CRM, billing, marketing, and support data — the layer that both humans and agents will query against.

Drive Agentic Workflow Development

  • Own the roadmap for AI agents built on top of the data warehouse — deciding which workflows justify an agent, what "grounded and trustworthy" means for each, and how they're deployed and monitored.
  • Design and build production agentic systems that read and act on warehouse data, complementing business workflows and adding value.
  • Architect end-to-end AI app stacks — serverless backends, LLM integrations, RAG, SQL generation, tool-calling/agentic orchestration, memory and context management — with access controls, evaluation harnesses, and accuracy safeguards appropriate to customer-facing use.
  • Establish engineering standards for agentic development at Rapsodo: observability (logging, tracing, quality monitoring) for AI systems, cost/latency-aware model and fallback strategy, safe and repeatable CI/CD for AI services, and guardrails against hallucination or unsafe autonomous action.
  • Push the organization's use of the warehouse forward — identify where an agent, a reporting dashboard or another solution, is the right answer to a business problem, and make the case for it.

Partner with Stakeholders and the Team

  • Support stakeholder reporting across Product, Sales, Support, Engineering, and leadership; own quarterly business processes that depend on the warehouse (e.g. commission calculations) in partnership with Sales and Finance.
  • Validate ingested data against business sources and lead investigations into discrepancies until they are resolved — the same rigor applies whether the consumer is a dashboard or an agent.
  • Manage data team and support other departments.
  • Document architectures, runbooks, agent evaluation results, and incident learnings; contribute to engineering standards across code review, testing, and release processes.

Requirements

Requirements

  • 6+ years in data engineering, analytics engineering, or platform engineering, with meaningful ownership of production systems and warehouse architecture end-to-end. Deep, hands-on data engineering experience is non-negotiable — this role leads with data custodianship, not AI engineering alone.
  • Demonstrated experience designing and shipping production AI agent systems — tool-calling architectures, RAG pipelines, context/memory orchestration, and evaluation frameworks — experience beyond prototyping with an LLM API is a plus.
  • Bachelor's degree in Computer Science, Engineering, or a related discipline; master's a plus.
  • Strong hands-on experience with a major cloud data platform (e.g. GCP, AWS, or Azure) — managed compute, serverless functions, container workloads — and Infrastructure-as-Code tooling.
  • Deep proficiency in SQL and Python; experience with a SQL transformation framework (Dataform, dbt, or equivalent) and container orchestration (Kubernetes or equivalent).
  • Experience building ingestion pipelines — batch and real-time (e.g. CDC) — and integrating SaaS sources with proper credential hygiene; familiarity with BI platforms and supporting non-technical users at scale.
  • Experience optimizing warehouse cost and performance on high-volume event data (partitioning, clustering, query tuning).
  • Practical experience with LLM APIs (OpenAI, Claude, or similar), vector databases/embedding pipelines, and evaluating LLM quality, hallucinations, and failure modes in production.
  • A track record of making the call on when an AI agent is — and isn't — the right solution to a business problem, and owning that decision through to a reliable, monitored production system.
  • Strong security mindset; comfort partnering with business stakeholders on cross-functional deliverables; excellent communication, documentation, and ownership in fast-paced environments.
  • Interest in sports technology is a plus.

Benefits

Why Rapsodo?

At Rapsodo, you won't just build agents on someone else's data platform — you'll own the platform itself and decide how AI gets built on top of it. This role is built for someone who wants both: the discipline of being the trusted custodian of a company's data, and the license to drive where agentic AI takes the organization next. If you want to shape both the foundation and the frontier, apply now.

Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available