DevOps/SRE
Posted Updated
xFarm Technologies is seeking a hands-on DevOps / Site Reliability Engineer to join the CloudOps team. You will own the operational infrastructure, collaborating with the Cloud Architect and other engineers across EU and LATAM to keep the backend platform running smoothly.
The role emphasizes Kubernetes operations, CI/CD ownership, and building reliable, observable systems with a focus on reducing toil and enabling fast, safe deployments across Europe and Latin America.
Kubernetes Operations: deploy, scale, secure and troubleshoot containerized backend services on Kubernetes. CI/CD Pipelines: own and improve GitLab CI/CD pipelines end to end. Observability: govern Elastic Cloud stack and Prometheus/Grafana dashboards and alerts. Reliability & Incident Response: define SLIs/SLOs, respond to incidents, root-cause analysis. Infrastructure as Code: manage environments with Terraform/Helm/GitOps. Developer Enablement: support engineering teams with tooling and runbooks. Security & Hygiene: secrets management and access control across pipelines and clusters. Hands-on DevOps/SRE/Platform Engineer in production. Kubernetes in production — deployment, operations & troubleshooting. Experience designing and maintaining CI/CD pipelines, preferably GitLab CI/CD. Observability with Elastic Stack and Prometheus/Grafana stacks. Infrastructure as Code and cloud fundamentals (containers, networking, secrets, IAM). Strong scripting ability (Bash / Python). Willingness for Full European-hours availability (09:00–18:00 CEST). Solid English for working across EU and LATAM. Hybrid work model