Senior Site Reliability Engineer, AWS/Datacenter Hybrid
Cognitiv Senior Site Reliability Engineer, AWS/Datacenter Hybrid
Summary
Senior SRE who owns and scales Cognitiv's AWS cloud infrastructure (compute, networking, security), drives company-wide service management around deployments, monitoring, and disaster recovery, and supports co-located datacenter deployments. Stack centers on AWS/EC2, Terraform, Ansible, Datadog, and some Kubernetes, with on-site hybrid work in Bellevue, WA.
We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv. Our immediate challenge is to scale and harden our AWS environment as we continue expanding our hybrid cloud footprint. Past that, we are a rapidly growing organization and need to continue to move toward industry best practices. This role demands an experienced engineer with the interest to rapidly learn our environment and help drive our long-term service management roadmap.
This role partners closely with our datacenter-focused SRE, so while deep, hands-on AWS expertise is the priority, working familiarity with datacenter operations is important so you can help provide multi-DC coverage when needed.
Location: This position will be in our Bellevue, WA office with a hybrid work schedule of 3 days in office (Mon/Tue/Wed) and 2 days remote (Thursday/Friday).
Responsibilities
- Design, implement, and maintain infrastructure across our AWS environment, serving as the primary owner of our cloud footprint.
- Evaluate our existing AWS architecture (compute, networking, security) and ensure we are set up for long-term scalability and growth.
- Work across engineering and product teams to scope projects tightly to core business requirements.
- Drive engineering-wide efforts to improve company service management around deployments, monitoring, and disaster recovery.
- Support and help maintain our co-located datacenter deployments alongside our datacenter-focused SRE, providing coverage as needed.
Our Stack
- Hosting: AWS and Equinix colocation.
- Monitoring: Datadog (migrating off Prometheus).
- Infrastructure as Code: Terraform and Ansible.
- Compute: Primarily raw EC2 instances and bare-metal machines, with a small footprint of Kubernetes.
Requirements
- Deep knowledge of AWS infrastructure, networking, and management practices.
- 10+ years of experience in operations, software engineering, or as an SRE.
- Working knowledge of modern datacenter practices, with the ability to support multi-DC deployments as needed.
- Proven experience with infrastructure as code
- Proficiency with Python and Bash.
- An independent self-starter who looks at the big picture, takes ownership, and independently seeks out new challenges with creative solutions.
- A constructive, supportive team player, with good communication and interpersonal skills.
Preferred Qualifications
- AWS certifications (e.g., Solutions Architect, SysOps Administrator)
- Experience with hybrid cloud/on-prem solutions.
- Hands-on experience building out datacenters.
- Willingness and interest to travel 1-2 times per quarter.
Salary: $160,000 - $210,000 USD Base Salary + Equity
What We Offer
- Medical, Dental and Vision plan for US employees & Extended Health Benefits for Canadian employees
- 12 weeks paid parental leave + 4 weeks WFH
- Unlimited PTO + Work-From-Anywhere August
- Career development with clear advancement paths
- Equity for all employees
- Hybrid work model & daily team lunch
- Health & wellness stipend + cell phone reimbursement
- 401(k) & RRSP with employer match
- Parking (CA, WA, Vancouver offices) & pre-tax commuter benefits
- Employee Assistance Program
- Comprehensive onboarding (Cognitiv University)
- …and more!
What You’ll Find at Cognitiv
- Festiv – We make work fun with cross-team games, events, and creative team bonding.
- Responsiv – You’ll be close to clients and leadership, influencing real outcomes.
- Inclusiv – Diversity and individuality are celebrated across all levels.
- Inventiv – We reward curiosity and embrace bold ideas.
- Transformativ – We support your growth with training, mentorship, and flexibility.
- Collaborativ – We operate across coasts, connected by purpose and teamwork.
As published by greenhouse
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location
- Preferred First Name optional
- LinkedIn Profile
- Why Cognitiv? Why are you interested in this role with us? written answer · optional
- Do you have 10+ years of total experience working in IT Operations, Software Engineering, or Site Reliability Engineering (SRE)? choose one
- Do you have at least 5+ years of hands-on experience designing, scaling, and troubleshooting complex AWS networking and compute environments (e.g., Direct Connect, VPC peering, EC2)? choose one
- Which cloud platform is your primary area of technical expertise? choose one
- Briefly describe your experience working with physical datacenter or co-location infrastructure (e.g., Equinix, rack/stack, hardware troubleshooting, hybrid networking like AWS Direct Connect/VPNs).
- Do you have at least 3+ years of hands-on experience using Terraform and/or Ansible to manage production infrastructure? choose one
- Which best describes how you use AI in your work today? choose one · optional
- We have a hybrid culture. Are you able to work out of our Bellevue WA office Monday, Tuesday and Wednesday? choose one
- Are you legally eligible to work in the United States today and in the future? choose one
- Do you require sponsorship to work in the United States? choose one
- The base salary for this role is $160,000 - $210,000 USD + Equity (Not Liquid). Does this align with your compensation expectations? choose one