Site Reliability Engineer - Intermediate
Summary
You design, build, and run large-scale, fault-tolerant cloud systems, automate remediation, and maintain 24/7 availability using DevSecOps practices and infrastructure-as-code.
You build and operate large-scale, distributed, fault-tolerant systems in a DevSecOps environment. You develop auto-remediation tools, troubleshoot cloud and operational issues, implement infrastructure as code, monitor availability, resolve incidents, and support highly available systems through a 24/7 first-response model.
Responsibilities
- Build and operate large-scale distributed systems
- Collaborate with development and operations teams
- Resolve cloud operations trouble tickets
- Develop and run troubleshooting scripts
- Create auto-remediation tools
- Establish monitoring and alerting for critical systems
- Build infrastructure as code patterns
- Participate in 24/7 incident and problem management
Requirements
- BS degree in Computer Science or a related technical field, or equivalent experience
- 2–5 years of experience in software engineering, systems administration, database administration, and networking
- 1+ years of experience developing or administering software in public cloud
- Experience monitoring infrastructure and application uptime
- Experience with Python, Bash, Java, Go, JavaScript, or Node.js
- Knowledge of systems, storage, networking, security, and databases
- System administration and automation skills
- Experience with Terraform, Chef, Ansible, Docker, or Kubernetes
- Proficiency with CI/CD tooling and practices
- Cloud certification strongly preferred
Benefits
- Healthcare packages
- 401k matching
- Paid time off