Data Engineer, Amazon Traffic Engineering
Posted Updated
We are seeking an experienced Data Engineer to build and operate the Core Data Infrastructure that underpins our ML and Science initiatives for Bot Management. You will design and own production-grade data pipelines that ingest billions of events from diverse source systems and transform raw signals into ML-ready feature groups that Science and ML Platform teams depend on for training, evaluation, and inference.
This is a hands-on engineering role at the center of a fast-moving ML organization. Our Science teams build increasingly sophisticated models, each requiring different data formats, latencies, and serving patterns. You will build the pipelines and feature infrastructure that make this possible: ingesting from Trails and Non-Trails sources, transforming disparate datasets into versioned feature groups, and operating the Feature Store that serves features consistently across all model types. You will partner directly with Applied Scientists to translate model data requirements into reliable pipelines, and work across data engineering, software engineering, and science teams to deliver data that is accurate, timely, and well-governed.
Key job responsibilities
Data Pipelines & Ingestion — Build and own batch and near real-time pipelines spanning Trails (raw and aggregated) and Non-Trails sources (Clickstream, Customer Segmentations, OPS). Evolve pipelines from Cradle/POC to production-grade using AWS Glue. Implement data quality checks, drift detection, and governance frameworks.
Feature Engineering & Serving — Build versioned feature groups across multiple storage backends (S3 for tabular data, OpenSearch for embeddings). Develop production pipelines that transform raw signals into ML-ready features, and help operate the Feature Store that serves them consistently to Science teams.
Streaming & Real-Time Systems — Develop and operate Apache Flink applications and stream processing for near real-time feature computation. Build event-driven data flows leveraging Kinesis and Kafka to support low-latency bot detection signals.
Science Partnership — Partner with Applied Scientists and ML Platform engineers to define data contracts and SLAs, understand model data requirements, and ensure feature pipelines integrate cleanly with training and inference systems.
Operational Excellence — Own the reliability, monitoring, and cost efficiency of the pipelines you build. Participate in on-call, root-cause data issues, and drive improvements that reduce operational load.
About the team
Traffic Engineering's Bot Management organization protects Amazon's ecosystem by detecting and mitigating automated threats at scale. Our Core ML Data Infrastructure team is responsible for building and operating the foundational data infrastructure that powers bot detection, AI agent identification, and content exfiltration defense. We are building a unified, model-agnostic, production-grade ML platform that brings together training, evaluation, and inference pipelines into a cohesive system serving multiple model types across the organization.
This is a hands-on engineering role at the center of a fast-moving ML organization. Our Science teams build increasingly sophisticated models, each requiring different data formats, latencies, and serving patterns. You will build the pipelines and feature infrastructure that make this possible: ingesting from Trails and Non-Trails sources, transforming disparate datasets into versioned feature groups, and operating the Feature Store that serves features consistently across all model types. You will partner directly with Applied Scientists to translate model data requirements into reliable pipelines, and work across data engineering, software engineering, and science teams to deliver data that is accurate, timely, and well-governed.
Key job responsibilities
Data Pipelines & Ingestion — Build and own batch and near real-time pipelines spanning Trails (raw and aggregated) and Non-Trails sources (Clickstream, Customer Segmentations, OPS). Evolve pipelines from Cradle/POC to production-grade using AWS Glue. Implement data quality checks, drift detection, and governance frameworks.
Feature Engineering & Serving — Build versioned feature groups across multiple storage backends (S3 for tabular data, OpenSearch for embeddings). Develop production pipelines that transform raw signals into ML-ready features, and help operate the Feature Store that serves them consistently to Science teams.
Streaming & Real-Time Systems — Develop and operate Apache Flink applications and stream processing for near real-time feature computation. Build event-driven data flows leveraging Kinesis and Kafka to support low-latency bot detection signals.
Science Partnership — Partner with Applied Scientists and ML Platform engineers to define data contracts and SLAs, understand model data requirements, and ensure feature pipelines integrate cleanly with training and inference systems.
Operational Excellence — Own the reliability, monitoring, and cost efficiency of the pipelines you build. Participate in on-call, root-cause data issues, and drive improvements that reduce operational load.
About the team
Traffic Engineering's Bot Management organization protects Amazon's ecosystem by detecting and mitigating automated threats at scale. Our Core ML Data Infrastructure team is responsible for building and operating the foundational data infrastructure that powers bot detection, AI agent identification, and content exfiltration defense. We are building a unified, model-agnostic, production-grade ML platform that brings together training, evaluation, and inference pipelines into a cohesive system serving multiple model types across the organization.