freehire launches on Product Hunt on 26 August.

Follow →

Senior Site Reliability Engineer (LatAm)

Open 32d

Summary

Senior hands-on Site Reliability Engineer role focused on AWS, Amazon EKS, Istio, and OpenTofu infrastructure-as-code. The engineer improves reliability, scalability, observability, and CI/CD pipelines for a cloud-based platform, defines SLIs/SLOs, handles incident response, and evaluates AI-assisted engineering tools. Works directly with a distributed US-based engineering team, requiring LatAm lo

Senior Site Reliability Engineer - AWS / EKS / Istio / OpenTofu

We are looking for a senior, hands-on Site Reliability Engineer to improve the reliability, scalability, observability, and delivery processes of a cloud-based platform running on AWS and Kubernetes.
The engineer must have strong production experience with AWS, Amazon EKS, Istio, and infrastructure as code using OpenTofu. The role will also focus on uptime and reliability metrics, build and CI/CD improvements, operational automation, and the evaluation and integration of AI-assisted capabilities into engineering workflows.

Requirements

- Strong commercial experience as a Site Reliability Engineer, DevOps Engineer, Platform Engineer, or Cloud Infrastructure Engineer.
- Advanced hands-on experience with AWS.
- Strong production experience with Amazon EKS.
- Strong Kubernetes administration and troubleshooting skills.
- Production experience with Istio or another comparable service mesh.
- Hands-on experience with OpenTofu or Terraform.
- Strong infrastructure-as-code practices, including:
• reusable modules;
• code review;
• state management;
• environment separation;
• controlled deployment processes.
- Experience designing, building, and maintaining CI/CD pipelines.
- Experience improving build and deployment processes.
- Strong experience with monitoring, logging, alerting, and observability.
- Experience defining uptime metrics, availability targets, SLIs, and SLOs.
- Experience troubleshooting distributed cloud applications and production infrastructure.
- Experience with incident response, root-cause analysis, and operational runbooks.
- Strong understanding of cloud networking, load balancing, DNS, certificates, IAM, and secrets management.
- Strong scripting and automation skills.
- Strong written and spoken English.
- Ability to work independently and collaborate directly with a distributed US-based engineering team.

Preferred Skills
- Experience with Datadog.
- Experience integrating AI-assisted tools into CI/CD, testing, code review, build diagnostics, or developer workflows.
- Familiarity with event-driven and distributed application architectures.
- Experience supporting multi-tenant SaaS platforms.
- Experience helping Technical Support or Production Support teams create runbooks and escalation procedures.
- Experience working in healthcare or another regulated software environment.
- Expected Seniority
- This is a senior, hands-on engineering role rather than a coordination-only DevOps position.
The candidate must be able to:
- independently assess the current infrastructure and delivery processes;
- identify reliability and operational risks;
- propose practical improvements;
- implement and validate those improvements in production environments;
- collaborate directly with engineering leadership and development teams.

Responsibilities

- Design, maintain, and improve AWS-based infrastructure.
- Provision and manage infrastructure using OpenTofu.
- Operate and improve Kubernetes environments running on Amazon EKS.
- Configure, maintain, and troubleshoot Istio service mesh components.
- Improve platform reliability, availability, scalability, and operational resilience.
- Define and implement uptime and service reliability metrics.
- Establish and improve SLIs, SLOs, dashboards, alerts, and service health indicators.
- Improve monitoring, logging, alerting, and incident detection.
- Review existing CI/CD and build pipelines and identify opportunities for improvement.
- Improve build reliability, execution time, repeatability, and developer feedback loops.
- Evaluate and integrate AI-assisted capabilities into CI/CD, build analysis, testing, code review, or other engineering workflows.
- Automate repetitive infrastructure, deployment, and operational processes.
- Troubleshoot infrastructure, networking, Kubernetes, service mesh, and deployment issues.
- Participate in production incident investigation, mitigation, and root-cause analysis.
- Create and maintain infrastructure documentation, operational runbooks, and recovery procedures.
- Collaborate with application engineers, Support Engineers, and platform stakeholders.
- Help create troubleshooting and production support playbooks that can be used by the Support Engineering team.
- Review infrastructure and deployment processes for reliability, security, and operational risks.

Working conditions

- 12 vacation days per year;
- 5 sick days per year;
- 3 paid company holidays per year;
- Access to therapist and psychologist support for mental well-being;
- Compensation for courses and conferences;
- English courses.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available