Cloud Engineer(ArgoCD , Prometheus , Flink , Kafka, Kubernetes, Terraform,CI/CD)
Summary
Designs, deploys, and manages cloud infrastructure on AWS, Kubernetes, and Terraform, while automating CI/CD pipelines and monitoring systems with Prometheus, Flink, and Kafka.
Key Responsibilities
- Design, deploy, and manage cloud infrastructure using AWS services while ensuring high availability and scalability.
- Build, administer, and optimize Kubernetes clusters for containerized applications in production environments.
- Develop and maintain Infrastructure as Code using Terraform for automated infrastructure provisioning and lifecycle management.
- Design, implement, and enhance CI/CD pipelines to streamline software delivery and deployment automation.
- Develop automation scripts using Python to improve operational efficiency and reduce manual effort.
- Monitor infrastructure, applications, and cloud services to ensure platform reliability and performance.
- Perform production incident management, root cause analysis, and implement preventive measures.
- Collaborate with development, security, and operations teams to improve platform reliability and deployment processes.
- Implement cloud security best practices, secrets management, and infrastructure automation.
- Optimize cloud resources, system performance, and operational costs across enterprise environments.
- Maintain technical documentation, operational runbooks, and deployment standards.
- Support on-call operations and contribute to continuous platform improvements.
Requirements
- 7+ years of experience in Site Reliability Engineering, Platform Engineering, or DevOps.
- Strong hands-on experience with AWS cloud services and cloud infrastructure management.
- Extensive experience managing Kubernetes environments and containerized workloads.
- Experience in Infrastructure as Code using Terraform.
- Strong experience building enterprise CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, ArgoCD, or similar tools.
- Hands on experience in Azure, OpenStack
- Proficiency in Python scripting for automation and infrastructure management.
- Experience with Docker, Linux administration, and cloud networking.
- Experience in monitoring, troubleshooting, and optimizing Flink applications in production environments.
- Hands-on experience with monitoring and observability tools such as Prometheus, Grafana, ELK/OpenSearch, and CloudWatch.
- Experience in DevSecOps practices, security automation, and configuration management.
- Strong troubleshooting, analytical, and production support skills.
- Experience working in Agile environments using Jira and Confluence.
- Experience in distributed data processing, stream analytics, and integration with messaging platforms such as Kafka.
- Preferred AWS Certifications or Kubernetes certifications.
- Experience with OpenShift, Rancher, or cloud-native technologies.
- Knowledge of AI/ML infrastructure or MLOps platforms.
- Experience with enterprise-scale production support and high-availability systems.