Senior Site Reliability Engineer
Summary
Senior Site Reliability Engineer at Claranet in Lisbon (hybrid), splitting time 50/50 between operational excellence (incident response, troubleshooting, reliability) and engineering projects (automation, observability, platform improvements). Core stack is Azure and Kubernetes (AKS), with Terraform, CI/CD, and Datadog.
We're fast learners,
hard workers, natural collaborators... and we Make Modern Happen!
Our ambition is to
unlock the potential of our digital world so that organisations everywhere can
innovate and thrive securely.
We aim to achieve
this goal by bringing together the world’s most talented people and the most
powerful technologies, combining them to address our customers' challenges and to
build something stronger together.
If you share our
vision, join us!
We are looking for a Site
Reliability Engineer to join our team and help us build and operate
reliable, scalable and secure technology platforms.
This role combines
two complementary areas of work:
- 50% Operational Excellence: ensuring the smooth
operation of our platforms, responding to service requests and incidents,
troubleshooting issues and continuously improving reliability.
- 50% Engineering & Improvement
Projects: designing and implementing automation, observability, infrastructure and
platform improvements that make our services more resilient and easier to
operate.
This role is
responsible for ensuring the reliability, performance, security, and
scalability of cloud-based platforms, primarily in Azure and Kubernetes environments. The position combines operational support, infrastructure
engineering, automation, and Site Reliability Engineering (SRE) practices.
Your responsibilities include:
- Monitoring and
maintaining cloud and Kubernetes platforms to ensure high availability and
performance.
- Investigating and
resolving incidents, conducting root cause analysis, and driving
continuous service improvements.
- Designing, deploying,
and managing scalable infrastructure in Azure.
- Managing Kubernetes
environments, preferably with AKS (Azure Kubernetes Service).
- Developing and
maintaining Infrastructure as Code using Terraform.
- Building and improving
CI/CD pipelines and automating operational processes.
- Implementing
observability solutions, including monitoring, logging, tracing, and
alerting tools such as Datadog.
- Defining and tracking
reliability and performance metrics (SLIs, SLOs, and error budgets).
- Collaborating with
development and infrastructure teams to deliver reliable, secure, and
maintainable platform solutions.
- Promoting DevOps,
automation, knowledge sharing, and a culture of continuous improvement.
You must have:
- Degree in Computer Science, Engineering or a related field, or
equivalent practical experience.
- At least 4 years of experience in SRE, DevOps, Platform
Engineering, Cloud Engineering or a similar role.
- Hands-on experience with Microsoft Azure, particularly compute,
networking and storage services.
- Practical experience with Kubernetes; experience with AKS is an
advantage.
- Experience with Terraform or another Infrastructure as Code tool.
- Familiarity with CI/CD practices and version control systems.
- Experience with monitoring, logging and alerting platforms such as
Datadog, Azure Monitor, Prometheus, Grafana or equivalent.
- Good scripting skills in Bash, Python or PowerShell.
- Understanding of software development and deployment practices.
- Experience with .NET and/or Java, microservices or business
applications deployed on Kubernetes is a strong advantage.
- Ability to troubleshoot complex technical issues in a structured
and collaborative way.
- Good written and verbal communication skills in English.
We value:
- Experience with AWS or Google Cloud.
- Experience with .NET or Java application development.
- Knowledge of SLI/SLO frameworks, error budgets and incident
management practices.
- Experience with distributed systems, APIs and cloud-native
architectures.
- Familiarity with security, networking and identity concepts in
Azure and Kubernetes.
- Relevant certifications, such as:
- Microsoft
Certified: Azure Fundamentals
- Microsoft
Certified: Azure Solutions Architect Expert
- Certified
Kubernetes Administrator
- HashiCorp
Terraform Associate
- Datadog
Fundamentals
We
offer:
- Regular
professional development;
- Certification
paths resources;
- Regular teambuilding programs;
- Friendly workplace.
Workplace: Lisbon (Hybrid)
Claranet: Make Modern
Happen!