Site Reliability Engineer - Cloud Operations
Summary
Site Reliability Engineer at Swissquote migrating and modernizing production apps on Kubernetes, managing a service mesh platform, improving observability, and providing L3 on-call support.
In this role, you will:
- Migrate and modernize production applications on Kubernetes,
- Integrate third-party software into our production platforms and make it fit our operational standards,
- Work alongside Software and IT Engineers to improve reliability, performance and operational readiness,
- Design and operate applications on our service mesh platform,
- Integrate safe deployment patterns such as canary releases and progressive rollouts,
- Define SLOs, SLIs and useful operational KPIs, then use them to drive improvements,
- Improve observability across metrics, logs and traces so problems are easier to spot and understand,
- Explore and integrate AI tools that can help with troubleshooting, incident analysis and remediation,
- Test how systems behave under load, during failures and when dependencies disappear,
- Automate repetitive operational work whenever it makes sense,
- Provide Level-3 support and participate in the on-call rotation.
- At least 3 years of experience in SRE, DevOps, Platform Engineering or a similar production-focused role,
- Solid hands-on experience running production workloads on Kubernetes, OpenShift, EKS or a similar Kubernetes platform,
- Good knowledge of Helm and how to package, configure and maintain applications with it,
- Experience working with service mesh technologies such as Istio or Linkerd, or strong Kubernetes networking experience,
- A good understanding of service-to-service networking, traffic routing, mTLS and TLS,
- Experience with GitOps and modern deployment strategies such as canary or progressive delivery,
- A practical understanding of SRE concepts such as SLIs, SLOs and error budgets,
- Experience with observability and tracing tooling such as Prometheus, Grafana, Elastic Stack or OpenTelemetry,
- Strong Linux and networking fundamentals, including TCP/IP, DNS and load balancing,
- Comfortable troubleshooting JVM-based applications in production and able to investigate issues related to heap usage, garbage collection or JVM configuration,
- Comfortable automating things with Python, Go, Bash or another programming language,
- Experience or strong interest in applying AI to observability, incident response or operational automation,
- Experience or a strong interest in operating applications that depend on GPU resources or other AI infrastructure,
- Familiarity with Infrastructure as Code tools such as Terraform, Ansible or Puppet.
Nice-to-Haves
- Experience with Argo CD, Argo Rollouts or Argo Workflows,
- Deeper experience with Istio, Linkerd or Envoy-based service mesh platforms,
- Experience designing or operating Kubernetes platforms at scale,
- Experience running Java or Spring Boot applications in production,
- Hands-on experience tuning JVM applications for performance or low-latency workloads,
- Experience integrating applications with self-hosted AI platforms such as vLLM,
- Experience troubleshooting AI infrastructure integrations, including model access, GPU availability and NVIDIA MIG configurations.
- Knowledge of Cilium, eBPF or other modern Kubernetes networking technologies,
- Experience with public cloud or large private cloud environments,
- CKAD, CKA, CKS or equivalent hands-on Kubernetes experience,
- A homelab, self-hosted services or side projects where you get to experiment, break things and build them again.
Who You Are
- You like understanding why systems behave the way they do, especially when something goes wrong,
- You automate repetitive work instead of accepting it as part of the job,
- You’re comfortable working across development, infrastructure and operations teams,
- You don’t mind getting deep into software you didn’t build yourself,
- You’re curious about AI and where it can genuinely improve day-to-day operations,
- You are fluent in English and have good conversational French,
- You enjoy keeping up with cloud-native technologies and trying new approaches when they solve a real problem.
Please note that Swissquote never requests sensitive personal information or payment of any kind during the recruitment process. Any such request is fraudulent.
SQ2