Senior Manager, Software Development, AWS SageMaker HyperPod Infra
Key job responsibilities
- Define and own a multi-year technical strategy for your area, aligning product roadmaps and engineering investments to deliver measurable customer value across multiple software teams.
- Lead, mentor, and grow a team of software development managers and senior engineers, building an inclusive culture where people develop their careers and deliver their best work.
- Partner with product management and cross-functional stakeholders to translate ambiguous business problems into clear engineering plans, making thoughtful trade-offs between opportunity, resources, and long-term sustainability.
- Establish mechanisms to audit customer experience, operational health, and software quality across your teams, driving continuous improvement in system design, performance, security, and reliability.
- Collaborate across organizational boundaries to influence technical direction, resolve prioritization conflicts, and ensure your architecture evolves to meet future business needs at scale.
A day in the life
You start your morning reviewing operational metrics and identifying areas that need attention across your teams. Mid-morning, you join a design review with senior engineers to evaluate a proposed architecture change, asking probing questions to ensure scalability and maintainability. After lunch, you meet with peer managers and product leaders to align on quarterly priorities and resolve a resource constraint. Later, you spend time in a one-on-one coaching session with one of your managers, discussing their team's growth plans and career development strategies.
About the team
The SageMaker HyperPod Infra team is building next-generation AI infrastructure for large-scale training, model customization, and inference. HyperPod is the only AWS service unifying training and inference across Slurm, Kubernetes, and Ray - delivering fault-tolerant infrastructure with sub-2-minute recovery, automated GPU health monitoring, and sophisticated capacity management across thousands of GPU nodes globally. Our team is focused on building AI infra software systems that directly impact how customers experience AI.
Our team values builders who thrive in ambiguity, think long-term, and want to define the future of AI infrastructure. We ship fast, iterate with purpose, and support some of the most demanding AI workloads in the world. We operate in a collaborative environment where engineering leaders work closely with product and business partners to solve complex, high-impact problems. We are committed to creating an inclusive culture where every team member has the opportunity to grow, contribute ideas, and make a real difference. If you're excited to do meaningful work at the frontier of generative AI at massive scale, we'd love to talk.