Senior Data Engineer, Behavior ML Planning & Prediction
Summary
Senior Data Engineer to build and own the end-to-end data architecture for autonomous driving ML — designing scalable pipelines from petabyte-scale multimodal fleet and simulation data ingestion to visualization, and serving canonical datasets that power motion planning across onboard and offboard systems. Work involves close collaboration with ML and platform teams, setting technical direction or
TEAM
WHO ARE WE LOOKING FOR?
RESPONSIBILITIES
- Own and set the technical direction for the end-to-end data architecture that unifies our data into a coherent, discoverable, and reliable platform — evaluating design and operational trade-offs across scalability, reliability, and cost with a long-term view rather than optimizing locally
- Design and build data pipelines from ingestion through transformation to serving and visualization — sourcing, modeling, and delivering the canonical datasets that turn raw fleet and simulation logs into trusted, reusable data, and keeping them consistent across teams
- Set shared technical direction across teams: partner with stakeholders org-wide to understand their data needs, weigh technical trade-offs rigorously and objectively, influence roadmaps, and drive consensus toward a single, trusted data foundation — representing key insights clearly for both technical and non-technical audiences
- Define and own data products, Service Level Agreements, and the self-serve dashboards and tooling that scale analytics across the organization, along with the monitoring, alerting, and operational practices that keep those promises
- De-risk major architectural bets before the organization commits to them, using rapid prototypes and focused technical investigations to turn open questions into evidence-based decisions
- Document architecture, data models, interfaces, and decisions clearly, so that designs, trade-offs, and the resulting datasets are easy for others across the organization to understand, adopt, and maintain
- Act as a technical leader beyond the team: mentor engineers, establish data engineering best practices and standards that other teams adopt, and raise the data capability of the wider Autonomy organization
MINIMUM QUALIFICATIONS
- 7+ years of experience building and operating production data pipelines and data platforms at scale
- Deep command of SQL and a modern programming language (e.g. Python), and hands-on expertise designing robust data models and multi-step ETL/ELT jobs
- Experience with a cloud data warehouse (e.g. BigQuery, Snowflake, Redshift) and with orchestration and transformation tooling (e.g. dbt, Airflow, or equivalents)
- Demonstrated ownership of the data architecture for large-scale systems — setting technical direction and reasoning explicitly about scalability, reliability, security, and cost trade-offs
- A track record of technical leadership across teams — setting engineering standards that others adopt, aligning peers who have competing priorities or differing technical choices, and driving org-wide decisions to closure without formal authority, while growing the capability of other engineers
- Comfort operating in ambiguity — taking a loosely-defined, cross-team problem and creating the clarity, structure, and momentum to solve it
- Excellent communication skills in English, with the ability to explain complex technical trade-offs clearly and persuasively
NICE TO HAVES
- Experience unifying or consolidating data across multiple pipelines, formats, or storage systems onto a common platform, or migrating from bespoke dataset formats to an open table format (e.g. Apache Iceberg) as the basis of a lakehouse architecture
- Experience establishing data products, contracts, and SLAs for widely-used datasets, along with data-quality frameworks and observability
- Experience with large-scale, multimodal data — including spatial and temporal/sequential data (e.g. sensor, log, simulation, time-series, trajectory, or scene/snapshot representations) — and modeling it for reliable downstream use
- Familiarity with autonomous driving or robotics domain concepts (e.g. vehicle motion — kinematics and dynamics, trajectories, coordinate frames; motion planning and prediction; perception; mapping and localization) and how they shape the data we work with
- Familiarity with distributed data processing (e.g. Spark, Ray), workflow orchestration (e.g. Flyte/Union, Airflow), and columnar/lakehouse formats (e.g. Parquet, Iceberg)
- Experience building self-serve analytics products, semantic layers, or BI/dashboarding tooling for cross-functional users
- Business-level proficiency in Japanese