Site Reliability Engineer
Summary
Site Reliability Engineer at Instrumental in Palo Alto: operate, scale, and automate an AWS-based SaaS manufacturing-analytics platform with bi-weekly on-call. Core stack includes AWS (EC2, ECS, RDS), Terraform, CI/CD, Datadog observability, Docker/Kubernetes, and Python/Bash scripting.
- 3–4 years of experience in Site Reliability Engineering, DevOps, Cloud Operations, Platform Engineering, or Systems Engineering supporting production SaaS environments.
- Strong hands-on experience with AWS, including EC2, VPC, IAM, RDS, ECS, and S3.
- Experience managing infrastructure using Terraform or other Infrastructure as Code technologies.
- Experience designing and supporting CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, or similar platforms.
- Strong experience with monitoring and observability tools, preferably Datadog, including dashboards, alerting, logging, and APM.
- Experience with Docker and Kubernetes.
- Scripting experience with Python and/or Bash.
- Experience supporting production environments through an on-call rotation, including incident response and root cause analysis.
- Proven ability to take ownership of production issues and drive them through investigation, remediation, and long-term resolution.
- Dead serious about performance, scalability, and reliability (PSR): You care deeply about how systems behave in the real world and continuously look for ways to make them more reliable, scalable, observable, and supportable.
- Automation, automation, automation: If something is repetitive, manual, or error-prone, your first instinct is to automate it and make it disappear.
- An engineer at heart: You don’t want to repeatedly fight the same fires. You look for the underlying cause and build durable engineering solutions that reduce operational toil and technical debt.
- Strong systems thinker: You understand how infrastructure, applications, networks, deployments, monitoring, and people interact—and can troubleshoot complex production issues across those boundaries.
- Collaborative and reliable: You partner closely with software engineers to make services production-ready, improve operational workflows, and build reliability into systems before they become problems.
- Comfortable with growth and ambiguity: You’re comfortable making good decisions without perfect information and adapting as the platform, customer base, and company scale quickly.
- Experience working in a high-growth B2B SaaS environment.
- Experience implementing SRE practices such as SLIs, SLOs, and error budgets.
- Experience building internal tooling and automation to eliminate operational toil.
- Experience supporting multi-region AWS environments.
- AWS cost optimization or FinOps experience.
- Network, application security, and compliance experience.
- Experience introducing AI tools or processes into engineering and operational workflows.
