Software Engineer III - AI ML Platform Reliability
Summary
Build and maintain the AI/ML platform infrastructure, ensuring reliability and scalability using Python, Kubernetes, Terraform, and CI/CD pipelines.
Salary: £100,000 - 100,000 per year
Requirements:- Formal training or certification in software engineering concepts, with proficient applied experience.
- Hands-on experience delivering system design, application development, testing, and operational stability.
- Advanced proficiency in Python.
- Proficiency across the Software Development Life Cycle.
- Experience with infrastructure-as-code and cloud-native delivery practices, including Terraform, containers, Kubernetes, CI/CD pipelines, and automated deployment workflows.
- Experience designing and developing large-scale distributed systems and cloud-native architecture.
- Experience building large-scale infrastructure and cloud-native delivery practices in Google Cloud, AWS, or Azure, with Terraform.
- Extensive experience implementing advanced observability using tools such as OpenTelemetry, Dynatrace, Grafana, or cloud-native services.
- Strong systematic problem-solving and troubleshooting skills in complex systems.
- Hands-on experience using enterprise-authorized AI-assisted software development tools, with the ability to critically evaluate and validate AI-generated outputs.
- Understanding of responsible AI use in engineering workflows, including data sensitivity, secure handling of inputs and outputs, and resiliency and security expectations.
- Design and implement solutions that enhance the reliability and scalability of AI/ML platforms and applications to support fast-growing demands.
- Own non-functional requirements and develop tooling for observability, security, resilience, infrastructure management, and operational excellence.
- Build and maintain scalable infrastructure to support the deployment and operation of large-scale AI platforms and applications.
- Build strong cross-functional relationships to drive engagement across the organization and deliver solutions to user problems.
- Participate in on-call rotations and escalation workflows.
- Debug and solve production issues, taking full ownership of problems and developing solutions.
- Deliver creative software solutions through design, development, and technical troubleshooting.
- Develop secure, high-quality production code and review and debug code written by others.
- Identify opportunities to eliminate or automate remediation of recurring issues to improve operational stability.
- Use enterprise-authorized AI coding assist tools to improve code quality, delivery speed, and productivity while validating outputs through peer review, automated testing, and secure coding standards.
- Apply knowledge of SDLC tools, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
- Contribute to a team culture of diversity, opportunity, inclusion, and respect.
- AI
- AWS
- Azure
- CI/CD
- Cloud
- Dynatrace
- Grafana
- Support
- Kubernetes
- Machine Learning
- OpenTelemetry
- Python
- Security
- Terraform
- AI Agents
- Marketing
More:
We are partnering directly with JPMorganChase to hire for this Software Engineer III role within our AI/ML Data Platforms organization. You will join our Reliability Engineering team and help design and deliver trusted, market-leading technology products that power AI and machine learning at scale. We are focused on building resilient, secure, and scalable AI/ML platforms with operational excellence at the core. JPMorganChase is a global leader in financial services, serving corporations, governments, wealthy individuals, and institutional investors, and we value diversity, inclusion, and long-term client partnerships. Our Corporate Functions team supports key areas across the firm, helping set our businesses, clients, customers, and employees up for success.
last updated 33 week of 2026