Senior DevOps Engineer
Summary
Senior DevOps engineer who owns multi-region cloud (AWS-first) infrastructure end to end: GitOps-based deployments, CI/CD pipelines, observability/SLOs, and secrets management, while championing AI agents and agentic automation to reduce operational toil in a SaaS environment.
- Own and evolve our cloud infrastructure across multi-region production environments, end-to-end.
- Lead our GitOps deployment model - designing and maintaining declarative, automated deployment workflows with zero manual gates.
- Build, maintain, and optimize CI/CD pipelines with a strong focus on developer experience, reliability, and speed.
- Initiate, implement, and champion an AI-first DevOps & SRE ecosystem - identifying opportunities, building AI agents and intelligent automation, and driving their adoption across engineering operations.
- Develop automation frameworks for provisioning, scaling, observability, and incident response, leveraging AI-powered tooling and agentic workflows to reduce toil.
- Operate and improve our observability platform: metrics, logs, alerting, dashboards, SLOs/SLIs, and on-call tooling.
- Champion zero-trust secrets management and credential-less authentication patterns across the stack.
- Partner with architects and engineering leadership on cloud cost optimization, availability, and performance.
- Build internal tooling and automation that multiplies engineering velocity across the organization.
- 5+ years of hands-on DevOps experience in a SaaS product environment - Must.
- Demonstrated initiative in applying AI to engineering operations - designing and building AI agents, agentic workflows, LLM-powered automation, or Model Context Protocol (MCP) integrations that reduced operational toil and improved production reliability, quality, or velocity - Must.
- Strong scripting and programming skills - Python and Bash for automation, tooling, and AI agent development; Go is a plus.
- Strong motivation to continuously learn and adopt emerging technologies, and to share that knowledge across the team.
- Deep, hands-on AWS expertise; multi-cloud (AWS, GCP, Azure) experience is a strong plus - Must.
- Strong understanding of containers and orchestration - Docker, Kubernetes, including workloads, networking, service mesh (Istio), Helm/Kustomize, and autoscaling (KEDA, HPA, VPA).
- Strong experience with:
- Infrastructure-as-Code - Terraform, Crossplane, and/or cloud-native declarative tooling.
- GitOps principles and tooling (ArgoCD or equivalent).
- CI/CD platforms - building reusable, scalable, security-hardened pipeline templates (GitHub Actions or equivalent).
- Secrets management - dynamic injection, IRSA/Workload Identity, avoiding long-lived credentials.
- Experience embedding security into CI/CD: vulnerability scanning, SBOM generation, and supply chain security (Trivy, Grype, Syft, JFrog Xray).
- Solid observability knowledge - OpenTelemetry, Prometheus, Grafana, Datadog, ELK/OpenSearch, distributed tracing.
- Hands-on experience with AI/ML workloads or LLMOps infrastructure - a significant advantage.
- Cost-awareness (FinOps) - treating cloud spend as a core engineering metric.
- Clear communication skills - able to align engineers, security teams, and leadership around infrastructure decisions.
- A strong sense of ownership - proactively identifying gaps and driving improvements.
Skills
- Agentic AI
- AI
- Argo CD
- Authentication
- Automation
- AWS
- Azure
- Bash
- CI/CD
- Cloud
- Cloud Native
- Datadog
- Developer Experience
- DevOps
- Docker
- ELK
- FinOps
- GCP
- GitHub
- GitHub Actions
- GitOps
- Grafana
- Helm
- Infrastructure as Code
- Istio
- Kubernetes
- LLM
- LLMOps
- Machine Learning
- MCP
- Networking
- Observability
- OpenSearch
- OpenTelemetry
- Prometheus
- Python
- SaaS
- Secrets Management
- Terraform
- Vulnerability Scanning
- Zero Trust