Sr. DevOps Engineer 2
About the role:
While primarily remote, this role requires occasional visits to the office in Coimbra. We plan to open offices in Aveiro and Porto in the future. This approach gives team members the flexibility to work remotely while also coming together in the office for collaboration and teamwork.
We are looking for a Principal DevOps Engineer to serve as the technical lead and architect for infrastructure, automation, and deployments. You will define standards and reference architectures, lead complex initiatives across Windows/Linux platforms (with deep expertise in Windows clustering), and own end-to-end reliability—from Infrastructure as Code to release engineering and observability.
- Act as technical lead for DevOps/Platform/Release engineering: set direction, standards, and best practices
- Architect and govern end-to-end delivery: infrastructure provisioning, configuration management, CI/CD, release processes, and operations
- Design and support Windows-based high availability solutions, with deep ownership of Windows clustering (failover/HA patterns, maintenance, upgrades, troubleshooting)
- Lead Linux automation and platform standardization (configuration, patching, hardening, performance tuning)
- Own Infrastructure as Code strategy with Terraform (modules, environments, state, governance)
- Own automation strategy with Ansible (reusable roles, inventories, secure secrets handling, idempotency)
- Build and standardize deployments using Octopus Deploy, GitHub, and Ansible (templates, shared steps, release promotion, rollback)
- Design and mature CI/CD pipelines (artifact versioning, approvals, promotion strategy, policy-as-code where applicable)
- Establish observability standards using VictoriaMetrics/Prometheus (metrics strategy, alerting, SLO/SLA monitoring, dashboards)
- Provide production leadership: incident response, RCA/postmortems, reliability improvements, capacity planning
- Mentor engineers, review designs/code, and raise overall engineering quality across teams
- Produce and maintain architecture docs, runbooks, and platform roadmaps
- Bachelors degree in Computer Science or related field
- 7+ years (or equivalent) in DevOps / SRE / Infrastructure Engineering, including leadership in complex environments
- Expert-level experience designing and operating Windows Server HA and clustering (Failover Clustering and related components)
- Strong Linux administration and automation experience (systemd, networking, storage, performance)
- Advanced skills with Terraform and Ansible (architecture, reusable components, secure operations)
- Strong deployment/release engineering experience with Octopus Deploy and GitHub (release governance, environment promotion, rollback)
- Monitoring/observability expertise with VictoriaMetrics and/or Prometheus (alerting strategy, metrics design, operational readiness)
- Production experience running Redis, RabbitMQ, Nginx (HA, tuning, troubleshooting)
- Strong understanding of networking and security fundamentals (TLS, DNS, load balancing, firewalling, least privilege)
- Proven ability to lead cross-team initiatives, make architectural decisions, and communicate clearly
- Kubernetes and container ecosystems (Docker, Helm)
- CI/CD platforms beyond GitHub (GitLab CI, Jenkins)
- Logging platforms (ELK/EFK, Loki)
- DR/BCP design, backup automation, zero-downtime upgrade strategies
- PowerShell and advanced scripting, configuration governance, secrets tooling (Vault/SOPS)
- Experience with Virtualization platforms such as VMWare or HyperV
- Experience in building, configuring, and tuning highly available MS SQL Server environments
- Experience in managing VOIP components and protocols (SIP , FreeSwitch, OpenSIP, session border controllers)
- Experience with load balancing components ( F5 LTM, F5 GTM)
- Experience with administering AWS or Azure tenants