Data Engineer II, Compliance Shared Services, RISC
Amazon Regulatory Intelligence Safety and Risk (RISC) team mission is to protect customers from products that are unsafe, illegal, illegally marketed, controversial or otherwise in violation of Amazon's policies while enabling our Selling Partners to offer their broadest selection of safe and compliant products. We achieve these objectives worldwide by: (1) taking a science-first approach to offer trustworthy listings to our customers, (2) inventing intuitive and precise tools to simplify our selling partners' compliance journey and (3) innovating to reduce our cost to serve. The RISC Data Engineering team is seeking an experienced Data Engineer to join our team. In this role, you will be responsible for designing, building, and maintaining large scale robust data pipelines and infrastructure to empower our data and analytics initiatives. You will collaborate closely with tech teams and business stakeholders to understand their requirements and support with scalable data solutions. Join our expert team to build scalable data solutions, improving Amazon business efficiency and simplifying our selling partners' compliance journey.
Key job responsibilities
1. Design, develop, and maintain automated ETL/ELT pipelines with monitoring using Python, Spark, SQL, and AWS services such as Redshift, S3, Glue, Lambda
2. Optimize data warehouse and data lake architectures using best practices for DDL, physical and logical table design, data partitioning, compression, and parallelization
3. Implement and support reporting and analytics infrastructure for internal business customers
4. Develop optimized data models and transformations to ensure high-quality, well-structured data for business and analytics applications
5. Develop and maintain enterprise-scale data security solutions including data encryption, database user access controls, logging, and permissions management for data warehouse and data lake implementations
6. Maintain data warehouse and data lake metadata, data catalog, and comprehensive user documentation
7. Collaborate with internal business customers and technical teams to gather, document, and implement requirements for data publishing and consumption via data warehouse, data lake, and analytics solutions
8. Build self-service tooling and automation (including GenAI-powered agents) that removes recurring manual effort from the data engineering workflow and reduces turnaround time for common requests such as dataset provisioning, table subscriptions, and compliance validation
9. Design and operate compliance-critical data validation pipelines (e.g., multi-stage matching/verification funnels, referential integrity checks, dead-letter-queue handling) that convert raw or model-generated data into accurate, auditable, and regulator-ready outputs
10. Monitor and optimize cluster/compute resource utilization (e.g., workload management tiers, job scheduling, cost-aware compute architectures) to improve reliability and reduce infrastructure spend
11. Stay current with emerging technologies, tools, and trends (including AI advancements), evaluating and incorporating them into the existing data ecosystem for continuous improvement
A day in the life
You'll build and maintain data pipelines that feed compliance reporting across RISC, manage the infrastructure behind them (clusters, cost, reliability), and help teammates apply the right security and data-handling guardrails before anything reaches production. You'll also work on solutions that let the broader analytics team scale, including GenAI tooling that automates repetitive requests so BIEs and BAs spend less time on pipeline upkeep and more on analysis. Some days you're deep in Spark or SQL, others you're reviewing an agent design or advising a stakeholder on the right data pattern to use.
About the team
Who Are We
We are a team of engineers building scalable data solutions to improve Amazon business efficiency and simplify our selling partners' compliance journey. Our mission is to give every analyst and compliance program at Amazon fast, reliable access to trustworthy data, so they can protect customers and help selling partners stay compliant without waiting on manual pipeline work. We own the data foundation behind the compliance programs, and the people who use what we build range from BIEs and program managers to compliance and legal teams who depend on that data being accurate and on time. As a data engineer here, you'll work directly with those teams to understand what they need and build the pipelines and tools that get them there.
Work/Life Balance
We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud.
Mentorship and Career Growth
We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.
Key job responsibilities
1. Design, develop, and maintain automated ETL/ELT pipelines with monitoring using Python, Spark, SQL, and AWS services such as Redshift, S3, Glue, Lambda
2. Optimize data warehouse and data lake architectures using best practices for DDL, physical and logical table design, data partitioning, compression, and parallelization
3. Implement and support reporting and analytics infrastructure for internal business customers
4. Develop optimized data models and transformations to ensure high-quality, well-structured data for business and analytics applications
5. Develop and maintain enterprise-scale data security solutions including data encryption, database user access controls, logging, and permissions management for data warehouse and data lake implementations
6. Maintain data warehouse and data lake metadata, data catalog, and comprehensive user documentation
7. Collaborate with internal business customers and technical teams to gather, document, and implement requirements for data publishing and consumption via data warehouse, data lake, and analytics solutions
8. Build self-service tooling and automation (including GenAI-powered agents) that removes recurring manual effort from the data engineering workflow and reduces turnaround time for common requests such as dataset provisioning, table subscriptions, and compliance validation
9. Design and operate compliance-critical data validation pipelines (e.g., multi-stage matching/verification funnels, referential integrity checks, dead-letter-queue handling) that convert raw or model-generated data into accurate, auditable, and regulator-ready outputs
10. Monitor and optimize cluster/compute resource utilization (e.g., workload management tiers, job scheduling, cost-aware compute architectures) to improve reliability and reduce infrastructure spend
11. Stay current with emerging technologies, tools, and trends (including AI advancements), evaluating and incorporating them into the existing data ecosystem for continuous improvement
A day in the life
You'll build and maintain data pipelines that feed compliance reporting across RISC, manage the infrastructure behind them (clusters, cost, reliability), and help teammates apply the right security and data-handling guardrails before anything reaches production. You'll also work on solutions that let the broader analytics team scale, including GenAI tooling that automates repetitive requests so BIEs and BAs spend less time on pipeline upkeep and more on analysis. Some days you're deep in Spark or SQL, others you're reviewing an agent design or advising a stakeholder on the right data pattern to use.
About the team
Who Are We
We are a team of engineers building scalable data solutions to improve Amazon business efficiency and simplify our selling partners' compliance journey. Our mission is to give every analyst and compliance program at Amazon fast, reliable access to trustworthy data, so they can protect customers and help selling partners stay compliant without waiting on manual pipeline work. We own the data foundation behind the compliance programs, and the people who use what we build range from BIEs and program managers to compliance and legal teams who depend on that data being accurate and on time. As a data engineer here, you'll work directly with those teams to understand what they need and build the pipelines and tools that get them there.
Work/Life Balance
We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud.
Mentorship and Career Growth
We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.