Senior Engineer DevOps
Senior Engineer DevOps:
Position Summary:
We are seeking a highly motivated and experienced Senior DevOps Engineer with experience in cloud infrastructure, platform engineering, automation, and DevOps practices. The ideal candidate will play a key role in designing, implementing, and managing scalable, secure, and reliable cloud platforms across Microsoft Azure and Google Cloud Platform (GCP) environments.
This role requires hands-on expertise in Kubernetes, CI/CD automation, Infrastructure-as-Code, observability, and cloud-native technologies to enable engineering teams to deliver high-quality solutions efficiently and securely.
Roles and Responsibilities:
Cloud Infrastructure & Platform Engineering:
- Design, deploy, and maintain scalable, highly available, and cost-effective cloud infrastructure on Azure and GCP.
- Manage and support containerized workloads using Azure Kubernetes Service (AKS) and Google Kubernetes Engine (GKE).
- Implement Kubernetes best practices including:
- Autoscaling
- Ingress Controllers
- Network Policies
- Workload Security
- Contribute to cloud architecture reviews and infrastructure optimization initiatives.
CI/CD & DevSecOps:
- Build and maintain automated CI/CD pipelines using GitHub Actions.
- Implement DevSecOps practices by incorporating security scanning, compliance checks, and automated quality gates into deployment pipelines.
- Support GitOps adoption and infrastructure automation initiatives.
- Collaborate with development teams to improve deployment frequency, reliability, and release quality.
Streaming & Messaging Platforms:
- Deploy, manage, and monitor Apache Kafka and Confluent Kafka clusters.
- Support event-driven architectures using Google Cloud Pub/Sub.
- Ensure high availability, performance optimization, and operational stability of messaging systems.
Observability & Reliability:
- Implement monitoring, logging, and alerting solutions using:
- Azure Monitor
- Prometheus
- Grafana
- ELK Stack
- GCP Operations Suite
- Monitor platform health and proactively identify performance bottlenecks.
- Support incident management, root cause analysis, and service reliability improvements.
- Assist in defining and tracking SLAs, SLOs, and SLIs.
Infrastructure Automation & IaC:
- Develop and manage Infrastructure-as-Code (IaC) solutions using:
- Terraform
- Bicep
- ARM Templates
- Automate infrastructure provisioning and operational workflows using:
- Python
- Bash
- PowerShell
- Maintain reusable automation frameworks and infrastructure modules.
Collaboration & Continuous Improvement
- Work closely with software engineering, security, data, and architecture teams.
- Participate in technical design reviews and cloud migration initiatives.
- Contribute to operational excellence through documentation, knowledge sharing, and process improvements.
- Mentor junior engineers and promote DevOps best practices across teams.
Required Qualifications:
Experience:
- Minimum 6+ years of experience in DevOps, Cloud Infrastructure, Platform Engineering, or Site Reliability Engineering (SRE).
- Experience supporting production environments in cloud-native and containerized ecosystems.
- Hands-on experience with modern infrastructure automation and deployment practices.
Technical Skills:
Strong expertise in:
- Microsoft Azure and Google Cloud Platform (GCP)
- Kubernetes platform administration (AKS, GKE)
- Docker and container orchestration
- GitHub Actions CI/CD pipelines
- Apache Kafka and Confluent Kafka
- Google Cloud Pub/Sub
- Terraform, Bicep, and ARM Templates
- Azure Monitor, Prometheus, Grafana, ELK Stack, and GCP Monitoring
- Python, Bash, and PowerShell scripting
- NoSQL databases and distributed systems
Professional Skills:
- Strong troubleshooting and problem-solving skills.
- Experience supporting mission-critical production environments.
- Ability to work effectively in cross-functional teams.
- Excellent communication and stakeholder management skills.
- Strong focus on automation, reliability, and continuous improvement.
Preferred Qualifications:
- Experience working in multi-cloud or hybrid-cloud environments.
- Understanding of disaster recovery, backup strategies, and cloud cost optimization.
- Kubernetes, Azure, GCP, or Terraform certifications are highly desirable.