SRE / DevOps Engineer
Summary
Experienced SRE/DevOps engineer joining Uni3's team in Gothenburg to build, operate, and improve reliable cloud services for a customer — running Kubernetes/Docker workloads, CI/CD pipelines, and IaC (Terraform/Ansible) across AWS/Azure/GCP, with observability via Prometheus, Grafana, and ELK.
About the Role
We are looking for an experienced SRE / DevOps Engineer to join our engineering team and help our customer to build, operate, and continuously improve reliable, scalable, and secure cloud-based services in Gothenburg.
In this role, you will work closely with development, infrastructure, and operations teams to improve system reliability, automate software delivery, strengthen observability, and troubleshoot complex production issues.
You will work in a modern cloud-native environment involving Kubernetes, CI/CD, public cloud platforms, infrastructure automation, monitoring, logging, and distributed microservices.
The role is particularly suitable for someone who enjoys solving complex technical problems and taking end-to-end ownership of production services.
Key Responsibilities
- Operate and continuously improve highly available production services and cloud infrastructure.
- Lead or participate in incident management, troubleshooting, root cause analysis, and post-incident reviews.
- Investigate complex production issues using logs, metrics, traces, databases, and application-level diagnostics.
- Design, build, and maintain CI/CD pipelines for automated build, testing, deployment, and release processes.
- Deploy and operate containerized applications using Docker and Kubernetes.
- Improve Kubernetes workload configuration, scalability, resource utilization, and cost efficiency.
- Develop and maintain Infrastructure as Code (IaC) and environment automation using tools such as Terraform and Ansible.
- Build and improve monitoring, logging, tracing, alerting, and observability solutions.
- Work with engineering teams to improve application reliability, performance, resilience, and operational readiness.
- Support cloud infrastructure across AWS, Azure, and/or GCP.
- Implement and maintain backup, disaster recovery, capacity planning, and high-availability solutions.
- Participate in network, access control, IAM, RBAC, and infrastructure security improvements.
- Automate repetitive operational tasks using Python, Bash, PowerShell, or similar scripting languages.
- Proactively identify system risks, performance bottlenecks, capacity issues, and opportunities for automation.
- Document technical solutions, operational procedures, runbooks, and troubleshooting knowledge.
Required Qualifications
- Several years of professional experience in DevOps, Site Reliability Engineering, Cloud Engineering, Platform Engineering, or Production Operations.
- Strong hands-on experience with Kubernetes and Docker.
- Experience designing and maintaining CI/CD pipelines using tools such as:
- Jenkins
- GitLab CI
- GitHub Actions
- Azure DevOps
- ArgoCD
- Strong Linux administration and troubleshooting skills.
- Experience with at least one major cloud platform: AWS, Microsoft Azure, or Google Cloud Platform.
- Hands-on experience with monitoring and observability technologies such as:
- Prometheus
- Grafana
- ELK / Elasticsearch / Kibana
- OpenTelemetry or equivalent technologies
- Experience troubleshooting distributed systems and microservice-based applications.
- Experience with Infrastructure as Code, preferably Terraform.
- Good scripting or programming skills in Python, Bash, PowerShell, Java, or similar languages.
- Good understanding of networking, security, IAM/RBAC, databases, and cloud infrastructure.
- Experience with incident management and root cause analysis.
- Strong communication skills and the ability to collaborate across development, infrastructure, and business teams.
- Professional working proficiency in English.
Preferred Qualifications
Experience in one or several of the following areas is considered an advantage:
- Multi-cloud environments involving AWS, Azure, and GCP.
- Service mesh technologies such as Istio.
- GitOps and continuous delivery using ArgoCD.
- Message and event platforms such as Kafka, RabbitMQ, MQTT/HiveMQ, or cloud Pub/Sub services.
- Databases such as PostgreSQL, MySQL, MongoDB, Redis, Elasticsearch, or Neo4j.
- Cloud networking including VPC, Transit Gateway, VPN, SD-WAN, load balancing, and DNS.
- Backup and disaster recovery architecture.
- Capacity planning and cloud cost optimization.
- Java / Spring Boot based microservice environments.
- Large-scale distributed or business-critical 24/7 production systems.
- Automotive, connected vehicle, IoT, telecom, or other high-availability platforms.
- AIOps, AI-assisted troubleshooting, or automation using LLM-based technologies.
Personal Qualities
We believe you are:
- Proactive and comfortable taking ownership of technical problems.
- Structured and analytical when troubleshooting complex systems.
- Comfortable working across both development and operations.
- Curious about new technologies and continuously looking for ways to improve automation and reliability.
- Able to work independently while collaborating effectively with multiple teams and stakeholders.
- Calm and solution-oriented when dealing with production incidents.
What You Will Work With
Depending on the project, the technology environment may include:
Kubernetes · Docker · Linux · AWS · Azure · GCP · Terraform · Ansible · Jenkins · GitLab CI · Azure DevOps · ArgoCD · Git · Prometheus · Grafana · ELK · Istio · Jaeger · Kafka · RabbitMQ · PostgreSQL · MongoDB · Redis · Python · Bash · Java / Spring Boot
Location
Gothenburg
Employment
Full-time
Öppen för alla Vi fokuserar på din kompetens, inte dina övriga förutsättningar. Vi är öppna för att anpassa rollen eller arbetsplatsen efter dina behov.
Skills
- AI
- Ansible
- Argo CD
- Automation
- AWS
- Azure
- Azure DevOps
- Bash
- CI/CD
- Cloud
- Cloud Native
- DevOps
- Distributed Systems
- DNS
- Docker
- Elasticsearch
- ELK
- GCP
- Git
- GitHub
- GitHub Actions
- GitLab
- GitOps
- Grafana
- IAM
- Infrastructure as Code
- Istio
- Java
- Jenkins
- Kafka
- Kibana
- Kubernetes
- Linux
- LLM
- Microservices
- MongoDB
- MQTT
- MySQL
- Neo4j
- Networking
- Observability
- OpenTelemetry
- PostgreSQL
- PowerShell
- Prometheus
- Python
- RabbitMQ
- RBAC
- Redis
- SD-WAN
- Spring
- Terraform
- VPC
- VPN
- WAN
