Senior DevOps Engineer (Hybrid Infrastructure & Cloud)
Job Details
Role
We are looking for a hands-on Senior DevOps Engineer responsible for managing and scaling a hybrid infrastructure environment across cloud and on-premise systems. This role requires strong ownership of production systems, incident response, and continuous improvement of system reliability, performance, and scalability.
Key Responsibilities
Infrastructure & Platform Management
- Manage and maintain hybrid infrastructure (cloud + on-premise)
- Operate Kubernetes clusters and NGINX ingress controllers
- Maintain and optimize Docker-based workloads and services
Cloud & Networking
- Manage AWS services including EC2, RDS, CloudFront, Route53, and ACM
- Configure and optimize NGINX (reverse proxy, SSL, buffering, routing)
- Troubleshoot and optimize CDN performance (e.g., Cloudfront, Cloudflare, object storage CDNs)
CI/CD & Automation
- Build and maintain CI/CD pipelines (GitHub Actions, Jenkins)
- Automate deployments, rollbacks, and environment provisioning
- Manage secrets and environment configurations securely
Observability & Monitoring
- Design and maintain monitoring stack (Prometheus, Grafana, Alertmanager, Loki)
- Implement alerting and logging strategies for production systems
- Investigate anomalies and performance bottlenecks
Storage & Data Systems
- Manage S3-compatible object storage systems
- Troubleshoot replication, checksum, and data integrity issues
- Optimize storage usage and data transfer performance
Incident Management & RCA
- Lead incident response and production troubleshooting
- Perform root cause analysis (RCA) and implement preventive measures
- Ensure system uptime, reliability, and SLA adherence
Scaling & Performance
- Plan and execute infrastructure scaling (auto-scaling, load balancing)
- Conduct capacity planning and resource optimization
- Identify and resolve performance bottlenecks across the stack
Location
Metro Manila, Misamis Oriental, Cebu
Job Type
Full Time
Work Setup
Work from Home
Schedule
8 Hours
Qualifications
Required Skills & Experience
- Strong experience with AWS (EC2, RDS, CloudFront, Route53, ACM)
- Solid experience with Kubernetes, Docker, and NGINX
- Hands-on experience with CI/CD tools (GitHub Actions, Jenkins)
- Deep understanding of monitoring and logging systems
- Prometheus, Grafana, Alertmanager, Loki
- Experience with CDN and edge networking
- Experience with S3-compatible object storage systems
- Strong troubleshooting and debugging skills in production environments
- Proven experience in incident management and RCA
- Strong understanding of scalability, high availability, and performance tuning
What We're Looking For
- Strong ownership mindset and accountability in production systems
- Ability to handle production issues under pressure
- Clear and structured troubleshooting approach
- Focus on reliability, scalability, and long-term improvements