Site Reliability Engineer I
Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.
- Monitor application and infrastructure health using enterprise monitoring and observability tools, including ELF, to ensure availability, performance, and reliability of enterprise platforms
- Configure, tune, and maintain alerting mechanisms in ELF, aligned to service health indicators and SLOs, to enable timely incident detection and reduce noise and false positives
- Develop and maintain dashboards providing visibility into system performance, availability, reliability trends, and key operational metrics
- Analyze metrics, logs, and distributed traces across application and infrastructure layers to proactively identify issues and support effective root cause analysis (RCA)
- Own and execute blameless RCAs for production incidents, identify corrective and preventive actions, and track them to closure
- Implement minor code fixes, configuration updates, and reliability enhancements as part of incident remediation and preventive measures
- Collaborate with application development and platform teams to review defects, propose fixes, and improve overall service reliability
- Participate in Agile sprint planning ceremonies, backlog grooming, estimation, and delivery of SRE‑owned work items
- Drive reliability improvements through sprint‑based commitments, including automation, operational fixes, and platform enhancements
- Participate in Disaster Recovery (DR) planning, testing, and execution to ensure resilience of business‑critical services
- Perform regular system patching and maintenance activities in line with organizational security, compliance, and audit requirements
- Support ITIL‑based Incident, Problem, and Change Management processes, including planning, documentation, approvals, execution, and post‑implementation validation
- Monitor network performance and troubleshoot connectivity, latency, and access‑related issues impacting platform traffic
- Participate in certificate lifecycle management, including provisioning, renewal, validation, and troubleshooting of SSL/TLS certificates
- Maintain and manage service accounts (Service IDs), including access provisioning, credential rotation, and compliance with security policies
- Drive automation and operational toil reduction using scripting, CI/CD pipelines, and platform tooling to improve reliability and scalability
- Maintain accurate documentation of system configurations, runbooks, SOPs, platform operational guidelines, and troubleshooting procedures, and generate reports on system performance, incidents, and resolutions
- Participate and lead the Development change review and change validation processes
- Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues
- Uses AI-assisted coding and documentation tools to support development of automation scripts, runbooks, and infrastructure as code with guidance from senior engineers
Education Qualifications:
- Minimum of 5+ years of relevant experience in application development, maintenance, and production support, along with hands-on exposure to Java and distributed systems in enterprise environments.
- Bachelor’s degree in computer science, Information Technology, Engineering, or equivalent practical experience; advanced degree is a plus
- Strong knowledge of operating systems and application runtimes such as Java and .NET
- Knowledge of distributed systems and service‑based architectures from an operations and reliability perspective
- Strong knowledge of modern observability stacks and platforms, including Splunk, Elasticsearch, Prometheus, and Grafana
- Knowledge of observability practices including logging, monitoring, tracing, and performance analysis
- Knowledge of RDBMS and NoSQL databases including MySQL, PostgreSQL, Couchbase, HBase, and Cassandra
- Knowledge of scripting and automation using languages such as PowerShell and Python
- Knowledge of AI, analytics, or AIOps platforms from an operational perspective is a plus
Work Experience:
- Experience in Incident, Problem, and Change Management using ServiceNow or similar ITSM tools
- Experience supporting production systems in large‑scale enterprise environments with a focus on reliability and availability
- Experience in system administration, infrastructure operations, and network troubleshooting
- Experience with CI/CD pipeline implementation and support using tools such as Jenkins, GitHub Actions, XL Release (XLR), or similar
- Experience managing and troubleshooting technology infrastructure and services, including servers, networks, and cloud platforms
- Knowledge of cloud‑based Site Reliability Engineering (SRE) practices with hands‑on experience on public cloud platforms such as AWS, Azure, or Google Cloud Platform
- Knowledge of containerization and orchestration technologies such as Docker and Kubernetes, and microservices‑based architectures
- Experience using enterprise monitoring and alerting platforms such as ELF
- Exposure to AI‑assisted monitoring, automation, or AIOps tools is a plus
• Proficiency in connecting to and administering servers via SSH (Secure Shell) - Knowledge of core networking concepts including ports, protocols, firewalls, and secure remote access
Licenses & Certifications
- Certification in at least one programming language or runtime such as Java, .NET, or Python
- Certification in containerization and orchestration technologies (Docker, Kubernetes, OpenShift) is a plus
- Public cloud certification in AWS or GCP is a plus
- Certification or training related to AI platforms, analytics platforms, or AIOps is a plus
Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions.
Skills
- Agile
- AI
- Analytics
- Automation
- AWS
- Azure
- Bash
- Cassandra
- CI/CD
- Cloud
- Containerization
- Distributed Systems
- Docker
- .NET
- Elasticsearch
- Express.js
- Firewall
- GCP
- GitHub
- GitHub Actions
- Grafana
- HBase
- Infrastructure as Code
- ITIL
- Java
- Jenkins
- Kubernetes
- Microservices
- MySQL
- Networking
- NoSQL
- Observability
- OpenShift
- PostgreSQL
- PowerShell
- Prometheus
- Python
- RDBMS
- ServiceNow
- Splunk
- SSH
- SSL
- TLS
