Site Reliability Engineer - AWS (4 to 8 Years)
Summary
Site Reliability Engineer at PhonePe managing and scaling mission-critical AWS infrastructure for high uptime: daily RHEL Linux administration, EC2/IAM/VPC/S3/CloudWatch operations, network troubleshooting, security compliance, automation with Ansible, SaltStack, Python/Go/Bash, MySQL and Aerospike data stores, observability, and on-call incident response with RCA reporting.
We are seeking a highly motivated Site Reliability Engineer (SRE) with 4 to 8 years of experience to manage, scale, and ensure the high availability of our core infrastructure. This role is designed for experts specialized in AWS. You will leverage a profound background in Linux (specifically RHEL) to drive deep-level cloud architecture, automation, complex networking, and security compliance, supporting a high-volume, mission-critical environment that demands exceptional uptime and resilience.
Roles and Responsibilities
Perform daily administration, configuration, and troubleshooting of Linux systems (specifically RHEL). Manage system resources, file systems, and package installations to ensure optimal OS-level health and performance.
Provision, configure, and maintain AWS EC2 instances and cloud-native components (IAM, Load Balancers, S3, CloudWatch, etc), specializing in RHEL environments.
Proactively monitor for, identify, and remediate infrastructure vulnerabilities to maintain a secure environment. Implement security guardrails defined by organizational policies.
Configure and maintain AWS VPCs, Route Tables, and Security Groups. Perform network-level troubleshooting for routing issues using tcptraceroute, mtr, tcpdump, netstat etc.
Collaborate with stakeholders to execute infrastructure migration plans and component upgrades with strict adherence to established runbooks.
Write and maintain Ansible playbooks and SaltStack states for configuration management. Automate routine operational tasks using Python, Go, or Bash.
Manage MySQL and Aerospike data stores, including routine backups, maintenance, upgrades and scaling.
Build, maintain, and extend platform observability by configuring monitoring dashboards and actionable alerts.
Create, update, and maintain comprehensive runbooks, standard operating procedures (SOPs), and architecture documentation to ensure operational consistency and seamless knowledge sharing.
Participate in the on-call rotation to handle active incidents, mitigate downtime, and draft initial Root Cause Analysis (RCA) reports.
Minimum Requirements
Experience: 4 to 8 years in an SRE, DevOps, or Systems Engineering role.
Strong proficiency in hands-on Linux/RHEL administration and troubleshooting.
Deep, hands-on experience exclusively within the AWS ecosystem
Mandatory hands-on experience with containerization using Docker, Podman, or equivalent technologies.
Solid grasp of the TCP/IP stack, DNS, routing concepts
Understanding of industry best practices for maintaining a highly secure infrastructure
Working knowledge of SSL/TLS certificate and their lifecycle management.
Ability to communicate effectively in English, both written and verbally
Preferred Qualifications (A Plus)
Working experience with Nginx and HAProxy proxies
Understanding of MySQL/MariaDB/Percona or any other relational database concepts and administration
Knowledge of NoSQL databases like Aerospike
Familiarity with Kafka or similar distributed event streaming platforms
PhonePe Full Time Employee Benefits (Not applicable for Intern or Contract Roles)
Insurance Benefits - Medical Insurance, Critical Illness Insurance, Accidental Insurance, Life Insurance Wellness Program - Employee Assistance Program, Onsite Medical Center, Emergency Support System Parental Support - Maternity Benefit, Paternity Benefit Program, Adoption Assistance Program, Day-care Support Program Mobility Benefits - Relocation benefits, Transfer Support Policy, Travel Policy Retirement Benefits - Employee PF Contribution, Flexible PF Contribution, Gratuity, NPS, Leave Encashment Other Benefits - Higher Education Assistance, Car Lease, Salary Advance Policy
Our inclusive culture promotes individual expression, creativity, innovation, and achievement and in turn helps us better understand and serve our customers. We see ourselves as a place for intellectual curiosity, ideas and debates, where diverse perspectives lead to deeper understanding and better quality results. PhonePe is an equal opportunity employer and is committed to treating all its employees and job applicants equally; regardless of gender, sexual preference, religion, race, color or disability. If you have a disability or special need that requires assistance or reasonable accommodation, during the application and hiring process, including support for the interview or onboarding process, please fill out this form.
Read more about PhonePe on our blog.
Life at PhonePe
PhonePe in the news