Staff Engineer DevOps
Staff Engineer DevOps
Position Summary:
We are seeking an experienced and highly skilled Senior Staff Engineer DevOps to lead the design, implementation, and optimization of cloud-native platforms, infrastructure automation, and DevOps practices across a multi-cloud ecosystem. This role requires deep expertise in Microsoft Azure, Google Cloud Platform (GCP), container orchestration, infrastructure-as-code, observability, and real-time data streaming technologies.
The ideal candidate will serve as a technical leader, driving platform engineering excellence, reliability, security, scalability, and automation while mentoring engineering teams and influencing strategic technology decisions.
Key Roles and Responsibilities:
Cloud Infrastructure & Platform Architecture:
- Design, implement, and manage highly available, scalable, and cost-optimized cloud infrastructure on Microsoft Azure and Google Cloud Platform (GCP).
- Lead deployment, administration, and lifecycle management of containerized applications using Azure Kubernetes Service (AKS) and Google Kubernetes Engine (GKE).
- Architect and implement advanced Kubernetes capabilities including:
- Horizontal and Vertical Pod Autoscaling
- Ingress Controllers
- Service Mesh (Istio)
- Network and Security Policies
- Establish cloud architecture standards, governance, security controls, and best practices for enterprise-scale environments.
CI/CD & DevSecOps:
- Architect and maintain enterprise-grade CI/CD pipelines utilizing GitHub Actions for automated, secure, and reliable software delivery.
- Integrate security, compliance, and quality controls throughout the software development lifecycle following DevSecOps principles.
- Drive adoption of GitOps methodologies and Policy-as-Code for infrastructure governance and deployment consistency.
- Collaborate with development, security, and operations teams to improve release quality and deployment velocity.
Streaming & Messaging Systems:
- Design, deploy, and manage Apache Kafka and Confluent Kafka platforms supporting mission-critical real-time streaming workloads.
- Implement scalable event-driven architectures using Google Cloud Pub/Sub.
- Optimize streaming infrastructure for performance, fault tolerance, scalability, and operational efficiency.
- Ensure high availability and security for enterprise messaging ecosystems.
Observability & Site Reliability Engineering (SRE):
- Define and implement enterprise-wide observability frameworks using:
- Azure Monitor
- Prometheus
- Grafana
- ELK Stack
- GCP Operations Suite
- Establish and maintain Service Level Agreements (SLAs), Service Level Objectives (SLOs), and Service Level Indicators (SLIs).
- Lead production incident management, root cause analysis, post-incident reviews, and resilience improvement initiatives.
- Promote reliability engineering best practices through automation and proactive monitoring strategies.
Infrastructure Automation & Infrastructure as Code (IaC):
- Champion Infrastructure-as-Code (IaC) adoption using:
- Terraform
- Bicep
- ARM Templates
- Automate cloud provisioning, configuration management, and operational workflows.
- Develop tooling and automation frameworks using:
- Python
- Bash
- PowerShell
- Cloud Functions
- Enforce infrastructure governance, version control, and repeatability across environments.
Technical Leadership:
- Provide technical direction and architectural guidance for cloud and DevOps initiatives.
- Mentor engineers and promote engineering excellence across teams.
- Drive continuous improvement in platform reliability, security, scalability, and operational efficiency.
- Collaborate with architects, developers, security teams, and business stakeholders to align technology solutions with organizational goals.
Required Qualifications:
Experience:
- 9+ years of experience in DevOps, Site Reliability Engineering (SRE), Cloud Infrastructure, or Platform Engineering roles.
- Minimum 3+ years in a senior technical leadership capacity.
- Strong track record of delivering cloud-native platforms and enterprise-scale automation solutions.
Technical Expertise:
Deep hands-on experience in:
- Microsoft Azure and Google Cloud Platform architecture and services
- AKS and GKE administration and optimization
- Kubernetes platform engineering and container orchestration
- Apache Kafka, Confluent Kafka, and Cloud Pub/Sub
- CI/CD pipeline development using GitHub Actions
- Observability platforms including Azure Monitor, Prometheus, Grafana, ELK Stack, and GCP Monitoring
- Scripting and automation using Python, Bash, and PowerShell
- Infrastructure-as-Code using Terraform, Bicep, and ARM Templates
- NoSQL databases and distributed systems
Core Experience:
- Enterprise system architecture and design
- Production operations and incident management
- Cloud security and compliance implementation
- Cross-functional collaboration and stakeholder management
- Performance tuning, scalability, and optimization
Preferred Qualifications:
- Experience operating in multi-cloud and hybrid-cloud environments.
- Familiarity with chaos engineering principles and resilience testing.
- Knowledge of disaster recovery, business continuity, and cloud cost optimization.
- Cloud certifications such as:
- Microsoft Certified: Azure Solutions Architect Expert
- Azure DevOps Engineer Expert
- Google Professional Cloud Architect
- Certified Kubernetes Administrator (CKA)
- Terraform Associate
Core Competencies:
- Strategic and systems-thinking mindset
- Strong analytical and problem-solving abilities
- Inclusive, collaborative, and empathetic leadership style
- High integrity and ownership mentality
- Innovation-driven with a focus on technical excellence
- Strong mentoring and coaching capabilities
- Customer-centric approach to technology delivery
- Commitment to continuous learning and professional development
Excellent written and verbal communication