Senior DevOps & SRE

Current Tech Stack & Infrastructure
 Cloud Infrastructure: AWS (ElastiCache, EKS, RDS Aurora MySQL)
 Orchestration: Kubernetes on EKS with Karpenter for node management
 Infrastructure as Code: Terraform (multiple repositories, experiencing drift and
collaboration challenges)
 Monitoring & Alerting: Datadog for monitoring, alerting, incident management, and
runbooks
 CI/CD: GitHub Actions with Atlantis for infrastructure PRs, Bitbucket Pipelines for Helm
deployments, Rundeck for scripted operations
 GitOps: ArgoCD for Kubernetes workloads
 Data Infrastructure: Transitioning to segmented data stores with ClickHouse, MySQL
Aurora, and Pulsar for event streaming
 APM: Limited Datadog APM usage due to cost ($45 per host/month, ~10 hosts)
 Additional Monitoring: Started Prometheus clusters within EKS for more verbose
metrics alongside Datadog
Scale & Workload
 Handling hundreds of thousands of events per second via API (mix of synchronous and
asynchronous processing)
 EC2 instances: M5/M6 xlarge instances, scaling between 15-100 instances per
environment depending on load
 Currently transitioning workloads from EC2 to Kubernetes (K8s workload still smaller
than core application)
 Real-time workloads requiring minimal downtime (minutes not hours for maintenance
windows)
Key Operational Challenges
 Infrastructure as Code: Terraform has become unmanageable due to drift reconciliation
and multi-person collaboration issues
 Legacy Environments: Some environments set up entirely manually, never in Terraform,
with major operational challenges to migrate while keeping them online
 Terraform Migration: Proven migration patterns in test environments, but moving to
production challenging due to real-time workload requirements
 Alert Management: Receiving too many alerts, need prioritization and structured
approach to reduce noise and recategorize/adjust thresholds
 Alert Distribution: Historically all alerts went to one tech ops team instead of being
distributed to five different dev teams; working to shift left closer to developers
 Service Catalog: Still defining service catalog and ownership model
 Runbooks: Exist in Datadog but need more polish and structure
 Database Scaling: Operational challenges with database scaling being addressed
through data store segmentation
 Monitoring Costs: Datadog is expensive; moved from CloudWatch ~8 years ago due to
cost
 Manual Changes: Infrastructure changes often implemented manually first, then
imported to Terraform and rolled out to other environments

Top 5 Required Skill Sets
1. Strong Terraform and CI/CD process expertise

2. Experience breaking down alerts to identify critical vs. non-critical and handling frequent
alarms (Datadog-specific experience preferred)
3. SLA/SLI definition experience (team currently being asked to do this without prior
experience)
4. AWS infrastructure expertise, DevOps-heavy background
5. Monitoring and alerting experience (Datadog preferred, though Prometheus/Grafana
experience also valuable)

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available