Senior Platform Engineer (DevOps / Site Reliability Engineer) RTO 1x In A Month
Summary
Senior Platform Engineer designs and maintains AWS-based cloud infrastructure and Kubernetes clusters for 30+ teams, using IaC, CI/CD, and automation to ensure reliability and security.
Are you passionate about building reliable, scalable, and automated cloud platforms?
We're looking for a Senior Platform Engineer (DevOps / Site Reliability Engineer) to join our growing engineering team. In this role, you'll design and manage cloud infrastructure that supports 30+ engineering teams, helping deliver secure, highly available, and production-ready platforms.
If you enjoy Infrastructure as Code, AWS, Kubernetes, CI/CD, automation, and solving complex production challenges, we'd love to hear from you.
Work Schedule: Flexible Shift (Must be willing to participate in an On-Call Rotation)
What You'll Do
- Design, implement, and maintain highly available infrastructure supporting multiple engineering teams.
- Build and manage cloud infrastructure on AWS, ensuring scalability, reliability, performance, and security.
- Develop and maintain Infrastructure as Code (IaC) using GitOps principles and automation tools.
- Design, build, and optimize CI/CD pipelines to enable efficient software delivery.
- Manage and scale containerized workloads using Docker and Kubernetes.
- Monitor production environments, proactively identify issues, and ensure platform reliability.
- Troubleshoot infrastructure, application, and networking issues across cloud environments.
- Participate in an on-call rotation and lead incident response for critical production systems.
- Collaborate closely with software engineers to build resilient, production-ready platforms.
- Conduct code reviews, contribute to architecture discussions, and champion engineering best practices.
- Continuously improve platform automation, observability, and operational efficiency.
What We're Looking For
Required Qualifications
- Bachelor's Degree in Computer Science, Information Technology, Computer Engineering, or a related field.
- 3+ years of experience in Platform Engineering, DevOps Engineering, or Site Reliability Engineering (SRE).
- Strong hands-on experience managing AWS cloud infrastructure in production environments.
- Solid experience with Infrastructure as Code (IaC) and automation.
- Hands-on experience with:
- AWS
- Docker
- GitHub
- GitHub Actions / CI/CD
- Programming experience in at least one of the following:
- Python
- Ruby
- Go (Golang)
- Strong understanding of GitOps and modern DevOps practices.
- Passion for automation, continuous improvement, and platform reliability.
- Excellent troubleshooting and analytical skills.
- Strong written and verbal English communication skills.
Nice to Have
- Experience with Terraform, Chef, or other Infrastructure Automation tools.
- Experience with ELK Stack (Elasticsearch, Logstash, Kibana).
- Experience with RabbitMQ or other messaging platforms.
- Exposure to NoSQL databases.
- Experience working with globally distributed engineering teams.
- Knowledge of observability, monitoring, and incident management best practices.
Why Join Us?
- Be part of a modern engineering culture focused on automation, reliability, and innovation.
- Exposure to cloud-native technologies, Kubernetes, GitOps, and DevOps best practices.
- Collaborate with global engineering teams on business-critical platforms.