Lead DevOps Engineer
Summary
Lead DevOps Engineer at Evam will own and set the technical direction for the Kubernetes platform across AWS, Azure, and bare metal supporting 50+ microservices, while managing and growing the platform team. Day to day mixes people leadership with CI/CD, GitOps, observability, and incident response using Kubernetes, Terraform, Jenkins, Prometheus/Grafana, and Kafka.
Recognized as a Forbes Türkiye Top 50 Startup, an Endeavor High-Impact Venture, and a Mar-Tech Awards winner, Evam is also proudly a Happy Place to Work.
We're building a platform that runs mission-critical production workloads at scale. If you enjoy owning the technical direction of our Kubernetes platform, building and growing a strong team, and turning incident learnings into systemic fixes, you'll fit here.
This is a platform-ownership and people-management role, not a ticket queue, so we're looking for a talented Lead DevOps Engineer to build, lead, and grow our platform team!
Responsibilities
- Set the technical direction and roadmap for Kubernetes platforms across AWS, Azure, and bare-metal environments, supporting 50+ microservices
- Own people management for the DevOps/platform team: hiring, onboarding, 1:1s, performance reviews, and career development
- Coach and grow the team's engineers, reviewing designs and raising the technical bar
- Set team goals and workload priorities, balancing platform roadmap with individual growth areas
- Own and improve CI/CD and GitOps-based delivery pipelines, enabling safe, zero-downtime releases
- Build and evolve observability (metrics, logs, traces) to ensure deep visibility and rapid incident detection across distributed systems
- Implement autoscaling, self-healing, and resilience patterns across services and infrastructure
- Collaborate with data and ML teams to operate and scale real-time ML and inference workloads on Kubernetes
- Enhance monitoring and incident detection using AI-assisted analysis and anomaly detection techniques
- Integrate security controls and scanning into pipelines and platform layers (DevSecOps)
- Lead incident response and drive postmortems into systemic reliability improvements, tracking follow-through across teams
- Own platform strategy conversations with engineering leadership, balancing reliability, cost, and delivery speed
- Build self-service platform tooling and documentation that enable developers to ship safely and fast
Our Stack
Orchestration: Kubernetes, Docker, Docker Swarm Cloud: AWS, Azure Cloud Services: EKS, ECR, ALB, WAF CI/CD: Jenkins, Github, Nexus IaC: Terraform, Ansible Observability: Prometheus, Grafana, Loki Messaging: Kafka Databases: PostgreSQL, MongoDB, Redis, Elasticsearch Networking: Istio, reverse proxy, TLS, load balancing Security: Trivy, Grype, Snyk, Fortify
Requirements
• BSc/MSc in Computer Science• 7+ years of hands-on experience in DevOps, SRE, Platform Engineering, or Infrastructure Engineering in production environments, with at least 2 years directly managing engineers (performance reviews, career development, hiring)
• Strong Kubernetes and container orchestration experience (cluster lifecycle, networking, storage, performance, troubleshooting) at a level where you can set standards and review others' designs
• Experience operating cloud environments (AWS and/or Azure), ideally multi-cloud and OnPrem environments
• Proficiency in Infrastructure as Code (Terraform, Ansible) and automated platform management
• Experience designing and operating CI/CD pipelines (Jenkins, GitHub Actions, or similar)
• Strong Linux and scripting skills with confidence in distributed systems troubleshooting
• Experience with observability stacks (metrics, logs, traces) and production monitoring practices
• Proven track record of making and owning architectural decisions for infrastructure/platform systems
• Experience building and growing a technical team from hiring and onboarding to performance management and career planning
• Strong communication and stakeholder management skills — able to represent the platform team to other engineering leads and to senior leadership
Nice to Have:
• Experience with event-driven architectures and Kafka at scale
• Observability tooling: Prometheus, Grafana, SigNoz, OpenTelemetry, Mimir, OneUptime
• Database operations (PostgreSQL, MongoDB, Redis, Elasticsearch)
• Experience in fintech, banking, or regulated industries
• GitOps tooling (ArgoCD, Flux) and DevSecOps practices
• Experience supporting AI/ML workloads on Kubernetes (Kubeflow, KServe, model serving, GPU scheduling)
• Familiarity with MLOps lifecycle (model deployment, monitoring, versioning)
• JVM-based containerized applications
• Relevant certifications (CKA, CKS, AWS, Azure)
Skills
- AI
- Anomaly Detection
- Ansible
- Argo CD
- AWS
- Azure
- CI/CD
- Cloud
- DevOps
- DevSecOps
- Distributed Systems
- Docker
- EKS
- Elasticsearch
- Event Driven Architecture
- Fintech
- Flux
- GitHub
- GitHub Actions
- GitOps
- Grafana
- Infrastructure as Code
- Istio
- Jenkins
- JVM
- Kafka
- Kubeflow
- Kubernetes
- Linux
- Loki
- Machine Learning
- Microservices
- MLOps
- Model Deployment
- MongoDB
- Networking
- Observability
- OpenTelemetry
- Performance Management
- PostgreSQL
- Prometheus
- Redis
- Stakeholder Management
- Terraform
- TLS
- WAF

