Senior AWS Site Reliability Engineer – Cloud Infrastructure
Summary
Senior engineer responsible for designing, automating, and maintaining highly reliable AWS cloud infrastructure, ensuring uptime and operational excellence for a cloud-focused client.
Empower Cloud Resilience — Drive Innovation and Reliability at Scale!
Portugal-based opportunity with remote work arrangement, allowing up to 5 days per week of telecommuting.
As a Senior AWS Site Reliability Engineer – Cloud Infrastructure, you will be working for our client, a leading player in the cloud and infrastructure domain. You will be responsible for ensuring the stability, scalability, and operational excellence of their AWS-based platform, supporting cutting-edge cloud solutions that empower digital transformation across industries. This role offers a unique opportunity to influence architecture standards, boost platform reliability, and grow your expertise in a fast-paced environment.
Your main responsibilities:
Define, measure, and enforce SLOs, SLIs, and error budgets, aligning engineering priorities with reliability goals.
Lead incident lifecycle activities including detection, response, mitigation, and post-incident analysis, enhancing system resilience.
Design and maintain highly available, fault-tolerant AWS architectures across compute, networking, storage, and data layers.
Develop and manage infrastructure-as-code (Terraform preferred, others acceptable) and CI/CD pipelines to ensure repeatable, low-risk deployments.
Systematically eliminate operational toil through automation, creating self-healing systems to optimize operational efficiency.
Manage observability stacks — metrics, logging, tracing, alerting — ensuring proactive problem identification and resolution.
Collaborate with development teams on capacity planning, performance tuning, and security compliance improvements.
Participate in on-call rotations, tuning alerting systems to minimize noise and fatigue.
At Lead level: set platform-wide reliability standards, influence architecture decisions, mentor engineers, and lead incident response strategies.
You're ideal for this role if you have:
7+ years of experience in SRE, DevOps, or infrastructure engineering, with proven reliability ownership.
Strong hands-on expertise with core AWS services — EC2, ECS/EKS, Lambda, VPC, IAM, S3, RDS, CloudWatch.
Deep understanding of SRE principles — SLOs, error budgets, toil reduction, blameless post-mortems.
Proficiency with infrastructure-as-code (Terraform preferred; CloudFormation or CDK acceptable).
Experience with container orchestration (Kubernetes/EKS) and CI/CD tooling.
Scripting and automation skills in Python, Go, or Bash.
Strong background in observability tools such as Prometheus, Grafana, Datadog, or ELK stack.
Proven incident response and on-call ownership experience in a production environment.
Knowledge of networking, Linux internals, and distributed systems failure modes.
It is a strong plus if you have:
AWS certifications (Solutions Architect, DevOps Engineer, SysOps).
Experience with multi-region/multi-account AWS setups and cost optimization.
Chaos engineering or resilience testing background.
Familiarity with service mesh, GitOps tools (ArgoCD/Flux), or policy-as-code (OPA).
Experience in regulated industries such as finance or healthcare.
Language Required for the role:
Fluent English (spoken and written).
Eligibility to work in Europe:
Only candidates with a legal right to work in Portugal or the wider European Union will be considered.
#MAKEYourCareerBETTER
Interested? Apply now and include your CV (preferably in English) along with a statement confirming your consent to the processing and storage of your personal data.