Point your AI agent at freehire and let it find you a job.

Get the CLI →

F5 Networks

NewBe an early applicant

Principal Site Reliability Engineer, Infrastructure & Platform

Posted Updated
Discussion

Responsibilities:

Infrastructure Automation & Configuration Management

  • Author, maintain, and refactor Ansible playbooks and roles across a large-scale multi-datacenter inventory, covering the full lifecycle from bare-metal provisioning to application deployment
  • Develop and improve CI/CD pipelines (GitLab CI) for infrastructure automation, including linting, testing, and staged rollout across regions
  • Manage secrets lifecycle using HashiCorp Vault, including AppRole authentication, secret rotation, and PKI integration
  • Maintain CMDB/IPAM accuracy in NetBox as a source of truth for all infrastructure assets

Compute & Virtualization

  • Deploy and manage Proxmox VE hypervisor clusters on bare-metal HPE hardware, including cluster formation, OVS networking, ZFS storage, and VM replication
  • Provision and lifecycle-manage virtual machines using cloud-init, QCOW2 images, and Proxmox API automation
  • Manage physical server provisioning end-to-end via HPE iLO (firmware updates, SPP deployment, OS installation via virtual media)

Container & Kubernetes Platforms

  • Manage self-hosted Kubernetes clusters on-premises, including control plane operations, node provisioning, workload deployment, and upgrade management
  • Operate Docker-based workloads on infrastructure VMs using compose-driven deployments and container health monitoring
  • Maintain container image pipelines and registry infrastructure (Azure Container Registry or AWS ECR)

Cloud Platforms

  • Engineer and maintain infrastructure on AWS and Azure, integrating cloud resources with on-premises systems (DNS, monitoring, identity, networking)
  • Apply cloud cost awareness, security best practices, and IaC principles (IAM, security groups, networking, storage) across AWS and Azure environments

Networking & Core Services

  • Operate and troubleshoot core distributed services including authoritative DNS (BIND9), recursive DNS (Unbound), load balancing (HAProxy), and high-availability VIPs (Keepalived/VRRP)
  • Maintain directory services (OpenLDAP master-replica topology) and AAA infrastructure (FreeRADIUS) used for SSH, VPN, and network device authentication
  • Manage OVS-based network configurations, VLAN topologies, and bonded NIC arrangements across hypervisor fleets

Observability & Security

  • Maintain and extend monitoring infrastructure (Prometheus, Observium) across a global fleet including SNMP polling, metrics collection, and alerting
  • Manage centralised log aggregation pipelines (Fluentbit) and ensure log delivery integrity across DCs
  • Operate runtime security tooling (Falco) and file integrity monitoring (AIDE) in production environments
  • Support PCI-DSS compliance activities including CIS hardening, audit logging (auditd), and participation in control reviews

Reliability & Incident Response

  • Participate in a 24x7 on-call rotation, responding to and leading production incident resolution
  • Conduct blameless post-mortems and drive remediation of root causes through automation and system improvements
  • Define and track SLOs/SLIs for critical infrastructure services
  • Identify and address single points of failure; design and implement HA improvements

Requirements:

  • Strong Linux systems administration skills (RHEL/CentOS preferred) including systemd, networking, storage, kernel tuning, and package management

  • Proficiency with Ansible (or similar tool) for large-scale configuration management, including role design, inventory management, and CI/CD integration

  • Hands-on experience with at least one hypervisor platform, preferably Proxmox VE, Harvester (Kubevirt) or similar (VMware vSphere, KVM)

  • Production experience operating on-premise Kubernetes clusters (rke2, k3s, etc)

  • Practical AWS or Azure experience including compute, networking (VPC/VNet, security groups, DNS), IAM, and managed services

  • Solid understanding of networking fundamentals: VLANs, bonding/LAG, routing, BGP concepts, DNS, load balancing, and firewall rule management

  • Experience with secrets management platforms (HashiCorp Vault or equivalent)

  • Familiarity with PCI-DSS requirements as they apply to infrastructure -- hardening standards (CIS benchmarks), audit logging, access control

  • Experience writing and maintaining CI/CD pipelines (GitLab CI, GitHub Actions, or equivalent)

  • Demonstrable on-call experience and comfort leading incident response in a global production environment

  • 7+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure Engineering role in a production environment

Skills

What Principal SRE jobs ask for — and how much of it you have →
Apply

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available