freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer

Summary

Improve reliability of market-critical trading systems by automating infrastructure, CI/CD, and observability while leading incident response and 24/7 on-call support.

You will improve the reliability of market-critical services through automation, observability, resilience engineering and operational readiness. You will develop CI/CD and infrastructure-as-code solutions, support capacity and performance planning, manage incidents and changes, and provide rostered 24/7 on-call support.

Responsibilities

  • Reduce operational toil through automation
  • Improve observability across logs, metrics and traces
  • Support production readiness and non-functional testing
  • Drive resilience, capacity, incident learning and reliability improvement
  • Design and maintain CI/CD pipelines
  • Develop and maintain infrastructure as code
  • Automate operational tasks, deployments and service recovery
  • Conduct production readiness assessments
  • Support capacity planning and performance engineering
  • Validate failover and recovery and participate in resilience exercises
  • Lead or contribute to post-incident reviews
  • Design observability practices and actionable alerts
  • Provide rostered 24/7 on-call support
  • Perform weekend and after-hours installations and upgrades
  • Manage incidents, problems, releases and changes
  • Undertake business-as-usual team work

Requirements

  • 5+ years of experience in a similar SRE role
  • Experience with incident response, post-incident review and problem management
  • Experience with production readiness, release readiness and operational acceptance
  • Experience with capacity, performance and resilience testing
  • Knowledge of observability design
  • Experience supporting high-availability distributed business-critical platforms
  • Scripting and automation skills using Python, PowerShell and shell scripting
  • AWS experience with EC2, S3, Lambda and RDS
  • Understanding of microservices and containerisation with Docker
  • Kubernetes administration experience
  • CI/CD pipeline experience
  • Experience with CloudWatch, Grafana, Prometheus and OpenTelemetry
  • Database operations experience with Oracle and/or Microsoft SQL Server
  • Linux and Unix administration and troubleshooting experience
  • Microsoft Windows Server 2019-2022 skills
  • Troubleshooting, problem-solving and root cause analysis skills
  • AWS certification at Associate level or above
  • Experience with distributed transactions, high availability and performance-critical systems
  • Networking troubleshooting across DNS, TLS, load balancers, firewalls and TCP/IP

Benefits

  • Hybrid working
  • Flexible working arrangements

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available