DevOps - AWS/GCP
Summary
A DevOps/Cloud Engineer who manages and optimizes cloud infrastructure on GCP or AWS, builds and maintains CI/CD pipelines, and troubleshoots VMs, Kubernetes, Docker, and database infrastructure. Also handles monitoring, security/compliance, incident root cause analysis, and technical reporting while coordinating with backend, security, and infrastructure teams.
Minimum 3–5 years of experience as a DevOps Engineer, Cloud Engineer, or related role.
Strong understanding of system architecture and cloud infrastructure, particularly Google Cloud Platform (GCP) or AWS.
Experienced in designing and implementing CI/CD pipelines for cloud-based application deployment.
Hands-on experience with Virtual Machines (VM), Kubernetes, and Docker, including troubleshooting.
Understanding of cloud infrastructure security, compliance, access control, and vulnerability remediation.
Experience supporting and troubleshooting database infrastructure, including performance, backup, and availability.
Familiar with cloud monitoring, logging, and infrastructure troubleshooting.
Strong analytical and problem-solving skills with the ability to perform root cause analysis.
Good communication skills and able to prepare clear technical and infrastructure reports.
Able to work collaboratively with Backend, Security, Infrastructure, and other technical teams.
Key Responsibilities
Manage, maintain, and optimize cloud infrastructure on GCP or AWS.
Design, implement, maintain, and troubleshoot CI/CD pipelines for application deployment.
Monitor and troubleshoot VM, Kubernetes, Docker, and other cloud services to ensure system reliability.
Handle infrastructure incidents and perform root cause analysis to resolve technical issues.
Support cloud security, compliance, access control, and vulnerability remediation.
Support database infrastructure troubleshooting, including performance, backup, and availability.
Monitor system performance and identify opportunities for infrastructure optimization and improvement.
Prepare infrastructure reports, incident summaries, technical documentation, and improvement recommendations.
Coordinate with Backend, Security, Infrastructure, and other technical teams to ensure stable system operations.
Implement preventive and corrective actions to improve system availability, reliability, and security.