freehire launches on Product Hunt on 26 August.

Follow →

Senior AWS Site Reliability Engineer – Cloud Infrastructure

Summary

Senior engineer responsible for designing, automating, and maintaining highly reliable AWS cloud infrastructure, ensuring uptime and operational excellence for a cloud-focused client.

Empower Cloud Resilience — Drive Innovation and Reliability at Scale!

Portugal-based opportunity with remote work arrangement, allowing up to 5 days per week of telecommuting.

As a Senior AWS Site Reliability Engineer – Cloud Infrastructure, you will be working for our client, a leading player in the cloud and infrastructure domain. You will be responsible for ensuring the stability, scalability, and operational excellence of their AWS-based platform, supporting cutting-edge cloud solutions that empower digital transformation across industries. This role offers a unique opportunity to influence architecture standards, boost platform reliability, and grow your expertise in a fast-paced environment.

Your main responsibilities:

  • Define, measure, and enforce SLOs, SLIs, and error budgets, aligning engineering priorities with reliability goals.

  • Lead incident lifecycle activities including detection, response, mitigation, and post-incident analysis, enhancing system resilience.

  • Design and maintain highly available, fault-tolerant AWS architectures across compute, networking, storage, and data layers.

  • Develop and manage infrastructure-as-code (Terraform preferred, others acceptable) and CI/CD pipelines to ensure repeatable, low-risk deployments.

  • Systematically eliminate operational toil through automation, creating self-healing systems to optimize operational efficiency.

  • Manage observability stacks — metrics, logging, tracing, alerting — ensuring proactive problem identification and resolution.

  • Collaborate with development teams on capacity planning, performance tuning, and security compliance improvements.

  • Participate in on-call rotations, tuning alerting systems to minimize noise and fatigue.

  • At Lead level: set platform-wide reliability standards, influence architecture decisions, mentor engineers, and lead incident response strategies.

You're ideal for this role if you have:

  • 7+ years of experience in SRE, DevOps, or infrastructure engineering, with proven reliability ownership.

  • Strong hands-on expertise with core AWS services — EC2, ECS/EKS, Lambda, VPC, IAM, S3, RDS, CloudWatch.

  • Deep understanding of SRE principles — SLOs, error budgets, toil reduction, blameless post-mortems.

  • Proficiency with infrastructure-as-code (Terraform preferred; CloudFormation or CDK acceptable).

  • Experience with container orchestration (Kubernetes/EKS) and CI/CD tooling.

  • Scripting and automation skills in Python, Go, or Bash.

  • Strong background in observability tools such as Prometheus, Grafana, Datadog, or ELK stack.

  • Proven incident response and on-call ownership experience in a production environment.

  • Knowledge of networking, Linux internals, and distributed systems failure modes.

It is a strong plus if you have:

  • AWS certifications (Solutions Architect, DevOps Engineer, SysOps).

  • Experience with multi-region/multi-account AWS setups and cost optimization.

  • Chaos engineering or resilience testing background.

  • Familiarity with service mesh, GitOps tools (ArgoCD/Flux), or policy-as-code (OPA).

  • Experience in regulated industries such as finance or healthcare.

Language Required for the role:

  • Fluent English (spoken and written).

Eligibility to work in Europe:

  • Only candidates with a legal right to work in Portugal or the wider European Union will be considered.

#MAKEYourCareerBETTER

Interested? Apply now and include your CV (preferably in English) along with a statement confirming your consent to the processing and storage of your personal data.

See also