Senior Devops Engineer
Job Description
Role & Responsibilities
Infrastructure Design & Automation
Hybrid & On-Prem Architecture: Architect, build, and maintain robust hybrid and on-premises environments for intensive compute and storage workloads.
Container Orchestration: Deploy, scale, and manage containerized microservices using Kubernetes across both on-prem and cloud infrastructures.
Infrastructure as Code (IaC): Automate end-to-end infrastructure provisioning and configuration management using Terraform, Ansible, and Helm .
Data Platform Support: Configure, optimize, and scale distributed data streaming and storage platforms (including Kafka clusters, Apache Spark, and Data Lakes ).
Scalability, Security & Governance Performance Engineering: Design frameworks for elastic scaling, high availability, load balancing, and resource isolation for high-throughput data services and APIs.
Hybrid Security: Secure hybrid deployments by implementing firewalls, VPNs, and strict identity/access management (IAM).
Data Protection: Implement granular data access controls, audit trails, and end-to-end encryption across services and storage layers while managing secrets securely.
Observability & CI/CD Pipelines Full-Stack Observability: Set up and maintain distributed monitoring and logging stacks using Prometheus, Grafana, ELK, and OpenTelemetry for real-time system insights.
Continuous Delivery: Build, enhance, and maintain CI/CD pipelines to deploy microservices, configurations, and heavy data jobs with zero-downtime and automated rollback capabilities.
SRE Practices: Drive System Reliability Engineering (SRE) practices, including active incident management, root-cause post-mortems, and comprehensive runbook creation.
Preferred Candidate Profile
Experience: 810+ years of dedicated experience in DevOps and System Reliability Engineering (SRE) roles.
Education: Bachelors or Masters degree in Computer Science, Engineering, or a related technical field.
Core Kubernetes Expertise: Proven, deep experience handling on-premises Kubernetes deployments alongside cloud variants (AWS/Azure).
Data Infrastructure Familiarity: Direct experience supporting infrastructure for real-time data platforms, event-driven architectures, microservices orchestration, and ETL pipelines.
Automation & Scripting: Strong proficiency in Bash, Python, or Golang for systems automation.
Networking Foundations: Solid understanding of core networking, distributed storage systems, firewalls, and security topologies for hybrid cloud models.
Good to Have
- Hands-on experience with Data Lake/Warehouse platforms (HDFS, S3-compatible object storage).
- Exposure to MLOps or data science workflows.
- Certifications in Kubernetes, cloud platforms, or security are a plus.