Site Reliability Engineer (Splunk, Python, OCI, Dynatrace, RCA, Terraform, Ansible )

Summary

The Site Reliability Engineer will manage and maintain highly available production environments, lead incident management and RCA, and automate infrastructure using Python, Terraform, and Ansible. The role focuses on observability and system reliability using tools like Splunk and Dynatrace within Linux and cloud environments.

Responsibilities

  • Manage and maintain highly available, scalable, and reliable production environments.
  • Lead incident management, root cause analysis (RCA), and problem management activities to ensure service stability.
  • Develop and maintain automation solutions using Python and Shell scripting to improve operational efficiency.
  • Implement and enhance monitoring, logging, and observability using Splunk and other enterprise monitoring tools.
  • Administer, troubleshoot, and optimize Linux servers and production infrastructure.
  • Provide L2/L3 production support and resolve critical application and infrastructure issues within SLA.
  • Automate infrastructure provisioning and configuration management using Infrastructure-as-Code tools.
  • Collaborate with development and infrastructure teams to improve system reliability, performance, and deployment processes.
  • Identify opportunities for process improvement and implement automation to reduce manual effort.
  • Create and maintain operational documentation, runbooks, and best practices for production support.

Requirements

  • Bachelor's degree in Computer Science, Information Technology, or a related field.
  • 8+ years of experience in Site Reliability Engineering (SRE), Production Support, or Infrastructure Operations.
  • Strong hands-on experience with Python automation and Shell scripting.
  • Strong experience in Dynatrace, AppDynamics
  • Hands on experience in Splunk for monitoring, log analysis, troubleshooting, and observability.
  • Strong experience in administering and troubleshooting Linux environments.
  • Experience in incident management, RCA, problem management, and production support for mission-critical systems.
  • Experience with cloud platforms such as AWS and/or Oracle Cloud Infrastructure (OCI).
  • Hands-on experience with Infrastructure-as-Code tools such as Terraform and configuration management tools like Ansible.
  • Experience in SQL, and application performance tuning is preferred.
  • Excellent analytical, troubleshooting, communication, and collaboration skills.

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available