Staff Site Reliability Engineer, Ads

Summary

This Staff Site Reliability Engineer role focuses on ensuring the reliability, scalability, and performance of a large-scale advertising technology ecosystem. The position involves hands-on engineering, incident management, and technical leadership to optimize high-QPS, low-latency systems using technologies like Go and Kubernetes.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer, Ads based in United States.

As a Staff Site Reliability Engineer, you will provide technical leadership for reliability across a large-scale advertising technology ecosystem.
You will help ensure critical systems remain highly available, scalable, performant, and resilient as traffic and business demands grow.
Your scope will span ad serving, auctions, targeting, reporting, measurement, and billing, with direct impact on revenue and customer experience.
You will partner with engineering leaders and multiple technical teams to establish reliability roadmaps, architecture standards, and operational practices.
The role combines hands-on engineering, automation, incident leadership, observability, and long-term resilience initiatives.
You will also mentor engineers and influence technical decisions across a broad organization.
This is an opportunity to shape reliability strategy for high-QPS, low-latency systems where operational excellence directly drives business outcomes.

Accountabilities

  • Lead reliability initiatives across multiple advertising technology domains, including ad serving, auctions, targeting, reporting, measurement, and billing.
  • Partner with engineering leadership to define and execute a roadmap focused on reliability, scalability, operational excellence, and developer productivity.
  • Design and build platforms, tooling, automation, and engineering solutions that improve system resilience and developer efficiency at scale.
  • Lead architecture reviews and influence technical decisions affecting critical, revenue-generating distributed systems.
  • Participate in on-call rotations, lead complex production investigations, and coordinate cross-functional responses to major incidents.
  • Identify systemic reliability risks and drive durable engineering solutions that strengthen platform resilience.
  • Establish and monitor reliability metrics for critical customer and advertiser journeys, including campaign creation, ad delivery, auction participation, reporting, attribution, and billing.
  • Improve operational maturity through SLOs, automation, incident management, observability, and performance optimization.
  • Troubleshoot complex issues across modern distributed system stacks and drive root-cause analysis toward long-term improvements.
  • Mentor engineers and provide technical leadership across multiple teams and initiatives.
  • Influence roadmap and investment decisions by ensuring reliability requirements are incorporated into product and infrastructure planning.
  • Requirements

    • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or a related discipline, with experience operating large-scale distributed systems.
    • Strong experience evolving and supporting high-traffic, user-facing production environments.
    • Deep expertise in distributed systems, scale engineering, cloud-native architectures, and highly available system design.
    • Strong software engineering capabilities in a general-purpose backend language such as Go.
    • Solid understanding of observability practices and technologies, including metrics, logging, tracing, and alerting.
    • Proven experience improving reliability through SLOs, automation, incident management, performance optimization, and operational best practices.
    • Demonstrated ability to diagnose and resolve complex issues across modern distributed technology stacks.
    • Strong cross-functional leadership skills, with the ability to influence technical direction and drive operational improvements across multiple teams.
    • Excellent written and verbal communication skills and the ability to collaborate effectively with engineering leadership and technical stakeholders.
    • Experience with Kubernetes, cloud infrastructure, and large-scale distributed systems is highly valued.
    • Experience with advertising technology or other revenue-critical platforms is a plus, particularly ad serving, real-time auctions, budget pacing, campaign delivery, measurement, attribution, or billing systems.
    • Experience operating high-QPS, low-latency services where performance directly affects business outcomes is advantageous.
    • Familiarity with large-scale data technologies such as Kafka, ClickHouse, Spark, Flink, or BigQuery is a plus.
    • Experience partnering with Product, Data Science, and advertising engineering teams is beneficial.
    • Exposure to machine learning inference or recommendation systems operating at scale is also valued.
    • Benefits

      • Comprehensive health benefits, including medical, dental, and vision coverage.
      • 401(k) program with employer matching.
      • Equity compensation in the form of restricted stock units, subject to the position offered.
      • Flexible vacation policy and global company days off.
      • 4+ months of paid parental leave.
      • Family planning support.
      • Workspace benefits and support for a home office.
      • Personal and professional development funds.
      • Paid volunteer time off.
      • Flexible-first workforce approach with remote-friendly work arrangements.
      • Reasonable accommodations available for qualified candidates with disabilities and disabled veterans.
      • Base salary range of $217,000–$303,900 USD, with final compensation determined by factors such as skills, depth of experience, relevant credentials, level, and location.
      • Certain positions may also include additional variable compensation depending on the role.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available