Site Reliability Engineer III
There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems
As a Software Engineer III (Associate) at JPMorganChase within Employee Platform — Collaboration and Communication, you will be a seasoned member of an agile team designing and delivering trusted, market-leading technology products in a secure, stable, and scalable way. You will deliver critical technology solutions across multiple technical areas and business functions in support of firm objectives.
Job responsibilities
- Deliver end-to-end software solutions (design, development, troubleshooting) with strong problem decomposition and innovative thinking beyond routine approaches.
- Build and maintain secure, high-quality production code, including synchronous integrations and algorithm maintenance.
- Produce architecture/design artifacts for complex applications and ensure implementation adheres to design constraints.
- Analyze large, diverse datasets to create insights, visualizations, and reporting that drive continuous improvement.
- Proactively identify hidden issues/patterns and use them to improve coding hygiene, reliability, and system architecture.
- Champion site reliability culture: define, educate, and operationalize SLOs/SLIs/error budgets with product, engineering, and stakeholders; monitor and review regularly.
- Drive delivery excellence through self-learning, issue/blocker resolution, and robust change & release management practices.
- Apply knowledge of tools within the SDLC toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
- Lead cross-team operational responsibilities including hosting platform/infrastructure planning, data center migrations, partner engagements, audits/application hygiene, and contribute to engineering communities, knowledge sharing, and an inclusive, resilient team culture.
- Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 3+ years applied experience.
- Hands-on experience across system design, application development, testing, and production/operational stability.
- Strong coding skills in Linux, Bash, and Python
- Proven ability to develop, debug, and maintain enterprise-scale code, including database querying; solid understanding of SDLC and end-to-end delivery.
- Practical experience with Agile ways of working, including CI/CD, application resiliency, and security practices; exposure to modern technical domains such as cloud, AI/ML, and/or mobile, and their supporting technical processes.
- Excellent communication and stakeholder management, adapting messages for senior business and junior technologists.
- Pragmatic mindset to drive technical re-engineering/modernization while continuing business delivery.
- Deep expertise in SRE best practices: reliability, scalability, performance, security, enterprise architecture, and toil reduction; able to promote SRE culture within the team.
- Strong operational tooling knowledge: file transfers, SSL/certificates, load balancers, job scheduling, monitoring, backup/capacity, plus containers/orchestration (Docker/ECS/Kubernetes), CI/CD tools (Jenkins/GitLab/Terraform), and network troubleshooting; experience with enterprise-authorized AI-assisted dev tools and responsible AI use (validation, security, data sensitivity).
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
- Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
Preferred / good-to-have skills
- SIP / VoIP familiarity

