Software Dev Engineer II, Alexa for Shopping
As a Software Development Engineer, you will solve challenging data and distributed-systems problems at the center of Amazon’s next-generation shopping AI. You will build scalable data platforms and intelligence capabilities that transform billions of customer interactions into high-quality training data, learning signals, and insights that directly improve large language models, agentic systems, and customer experiences.
You will own important components spanning data ingestion, behavioral analysis, dataset generation, and experimentation and evaluation for large language models and AI agents. Working closely with applied scientists and experienced engineers, you will turn complex research and product needs into reliable production systems and help pioneer automated, agent-driven workflows for continuous model improvement.
Key job responsibilities
1, Design, build, and operate scalable systems that transform real customer interactions into high-quality datasets, behavioral signals, and actionable insights for model training, post-training, and evaluation.
2, Develop data intelligence capabilities to analyze customer behavior, identify meaningful patterns and model quality gaps, and discover signals that improve large language models and agentic systems.
3, Build reliable pipelines and self-service tools for data ingestion, filtering, sampling, aggregation, dataset generation, quality validation, and exploratory analysis.
4, Partner with applied scientists to define metrics, analyze experiments, evaluate training data effectiveness, and translate findings into improved data recipes and learning signals.
5, Develop automated and agent-driven workflows for data curation, anomaly detection, experimentation, evaluation, and continuous model improvement while maintaining high standards for privacy, security, scalability, and operational excellence.
Preferred qualifications
1, Experience building large-scale data platforms, analytical systems, or data products using technologies such as Spark, Flink, SQL, distributed storage, or workflow orchestration systems.
2, Experience analyzing large and complex datasets to identify customer behavior patterns, data quality issues, and opportunities to improve machine learning models.
3, Experience defining metrics and building data quality monitoring, anomaly detection, or self-service analytics capabilities.
4, Knowledge of experimental design, statistical analysis, sampling methodologies, and techniques for measuring the impact of data or model changes.
5, Experience partnering with applied scientists, data scientists, or machine learning engineers to translate ambiguous analytical requirements into scalable production systems.
6, Knowledge of machine learning and large language model workflows, including training data preparation, post-training, experimentation, evaluation, model behavior analysis, agentic systems, and feedback loops.