Point your AI agent at freehire and let it find you a job.

Get the CLI →

Hebbia

NewBe an early applicant

Software Engineer Site Reliability

Posted Updated 3 views
Discussion

Summary

Hebbia is hiring a Site Reliability Software Engineer in NYC or San Francisco to own critical production services end to end: writing production code, profiling and rewriting hot paths, leading incident response, and building observability, deployment tooling, and CI/CD. Core stack spans Go/Python/C++/Rust, distributed systems, container orchestration, AWS, and observability tooling.

You will own critical production systems from design through incident response. You will write production code, improve service performance and observability, define reliability objectives, build deployment tooling, and turn incident learnings into durable architectural improvements.

Responsibilities

  • Own critical production services from design and code review through deployment, operation, and incident response
  • Profile, benchmark, and rewrite hot paths to eliminate performance bottlenecks
  • Lead incident response and turn post-mortem findings into code and architecture improvements
  • Build observability frameworks, instrumentation, alerting logic, and debugging tooling
  • Define and enforce SLOs for platform services
  • Own capacity planning and cost-efficiency automation
  • Build internal platforms and deployment tooling
  • Improve CI/CD systems
  • Embed with product engineering teams to co-design reliable systems
  • Partner on infrastructure security through threat modeling, hardening, and compliance tooling

Requirements

  • 5+ years of software development experience writing, shipping, and maintaining production services
  • Proficiency in Go, Python, C++, or Rust
  • Experience as a Production Engineer, SRE, or infrastructure-focused software engineer
  • Deep understanding of distributed systems
  • Container orchestration expertise
  • Experience debugging distributed production failures
  • Knowledge of operating-system concepts
  • Cloud platform experience, preferably AWS
  • Experience building and maintaining observability stacks
  • CI/CD pipeline expertise

Benefits

  • Unlimited PTO
  • Medical, dental, vision, and 401K
  • Daily catered lunch
  • DoorDash dinner credit for late work
  • 3 months of parental leave for non-birthing parents and 4 months for birthing parents
  • $15k lifetime fertility benefit
  • New-hire equity grant

Skills

Apply

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available