Site Reliability Engineer
Site Reliability Engineer
Location: New York City, United States
Department: Infrastructure & Security
Location Type: REMOTE
Employment Type: FULL_TIME
About the Role
What You'll Do
- Infrastructure & Kubernetes Orchestration
- Designing, deploying, and maintaining Kubernetes (EKS) clusters for enterprise-grade availability.
- Optimizing the footprint of infrastructure objects across core AWS services (EC2, RDS, S3) for performance, cost, and reliability.
- Evolving a scalable infrastructure management platform, with the right interfaces and guardrails to maximize engineering agency at minimal cognitive load.
- Defining golden paths for workload orchestration and integration with infrastructure dependencies, ensuring the right way to do things is also the easiest.
- Automation & AI-Assisted Operations
- Providing robust, reusable GitHub Actions components to streamline the software delivery lifecycle.
- Developing internal tools that replace manual operations with intelligent, autonomous systems.
- Exploring and deploying agentic workflows for AI-assisted runbooks that automate complex, error-prone procedures and repetitive tasks.
- Observability & Incident Management
- Driving the evolution of observability practices by maintaining reliable mechanisms to collect the metrics, traces, and logs needed to meet SLOs.
- Leading incident response efforts and facilitating the blameless postmortems that help systematically reduce recovery time (MTTR).
- Defining and monitoring platform-level SLIs and SLOs to ensure it consistently meets rigorous healthcare performance standards.
- Compliance & Collaboration
- Ensuring every piece of infrastructure is continuously compliant with HIPAA and other critical healthcare regulatory requirements.
- Reviewing architectural decisions and technical proposals, asking incisive questions to surface technical and organizational risks before they become incidents.
- Mentoring engineers across the company on reliability best practices and contributing a clinical-safety perspective to cross-functional design reviews.
Why You Might Be a Good Fit
- You design systems, shape processes, and write code, treating provisioning, orchestration, observability, and security as disciplines, not just tools to operate.
- You want to empower the people around you and make that the point of everything you build for them, not a side effect.
- You instinctively dig for the root cause rather than settling for a quick patch.
- You use AI and agentic tools where they help, know where they add risk, and own everything you ship.
- You thrive in fast-paced, safety-critical environments where pragmatism is balanced with technical rigor.
This Might Not Be The Right Fit If...
- You prefer the stability of static infrastructure and settled processes over continually adapting your codebase, tooling, and workflows as systems evolve.
- You are looking for a more specialized role where you can focus on a single discipline or technology.
- You prefer a siloed role that does not involve active collaboration with colleagues from different departments and backgrounds.
- You take pride in being the indispensable go-to person when something needs to be executed, rather than in building systems others can run without you.
Your Qualifications
- 5+ years of experience in SRE or Platform Engineering roles managing production environments at scale.
- Expert technical depth in AWS (EKS, EC2, RDS, S3) and production-grade Kubernetes management.
- Proficiency with modern tooling including Terraform (IaC), Datadog (Observability), Helm (Release Management), and GitHub Actions (CI/CD).
- Solid coding and scripting skills in Python, Bash, or Go.
- Preferred experience building agentic workflows or AI-assisted tooling to drive operational efficiency.
- A "rigor-first" mindset with a dedication to HIPAA-compliant, high-availability architecture.
