Sr. Director, Back-End Engineering
Summary
The Senior Director of Site Reliability Engineering will lead Coupang's company-wide strategy for system reliability, resilience, and autonomous operations. This role involves building and scaling a world-class SRE organization, defining engineering standards, and implementing AI-assisted incident management across large-scale distributed systems.
Sr. Director, Site Reliability Engineering
Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define and lead company-wide reliability, resilience, scalability, and operational excellence. This leader will transform reliability from a collection of team-specific practices into platform mechanisms that services inherit by tier, while advancing incident response toward an intelligent, AI-assisted, and increasingly autonomous operating model. We are looking for a visionary, industry-recognized technology leader who has previously conceived, built, and scaled a comparable SRE, production engineering, resilience, or autonomous-operations organization at a leading global technology company. The successful candidate must combine deep technical credibility with the organizational leadership required to align executives, influence architecture across the company, and build a world-class leadership bench.
Key Responsibilities
- Set a bold, multi-year vision for company-wide reliability, resilience, and autonomous operations, and translate that vision into an executable roadmap with measurable business outcomes.
- Define and own the SRE strategy, operating model, engineering standards, and reliability governance across Coupang.
- Build platform mechanisms that allow services to inherit reliability requirements based on service tier rather than recreate them independently.
- Lead initiatives that materially improve availability, resilience, scalability, performance, and operational readiness.
- Partner with engineering, product, infrastructure, security, finance, and business leaders to align reliability investments with customer and business priorities.
- Own executive reliability metrics, including availability, detection and recovery performance, change risk, incident recurrence, capacity readiness, and operational toil.
- Build and scale a world-class SRE organization capable of influencing engineering practices across the company.
Reliability Strategy, SLOs & Engineering Governance
- Establish and evolve service-tier definitions, SLOs, SLAs, error budgets, reliability scorecards, and objective certification mechanisms such as RBD/RBO.
- Create clear reliability requirements for Tier 0, Tier 1, and Tier 2 services, including redundancy, load testing, disaster recovery, observability, and incident response.
- Ensure reliability governance is embedded in architecture, development, release, and production operations rather than applied as a final review.
- Drive systematic reduction of recurring incidents, reliability risks, operational debt, and unsafe change patterns.
- Influence company-wide architecture for graceful degradation, fault isolation, load shedding, circuit breaking, and failure containment.
Incident Management & Autonomous Operations
- Transform incident management into a fast, disciplined, data-driven, and increasingly autonomous operating model.
- Enable AI-assisted detection, event correlation, triage, escalation, root-cause drafting, remediation recommendations, and selected guardrailed auto-remediation.
- Improve incident command, on-call quality, escalation mechanisms, communication, post-incident learning, and corrective-action completion.
- Reduce noisy alerts, manual on-call work, repeated diagnosis, and time spent coordinating across fragmented systems.
- Use incident and telemetry data to continuously improve platform standards, testing, capacity models, and engineering roadmaps.
Disaster Recovery, Resilience & Capacity
- Own the strategy and execution model for disaster recovery, regional resilience, availability-zone loss, capacity-constrained recovery, and critical business continuity.
- Build reusable DR and failover mechanisms that services inherit from the platform rather than implement as bespoke projects.
- Establish objective RPO/RTO targets, automated readiness gates, regular game days, fault injection, and evidence-based recovery certification.
- Drive proactive and intelligent capacity management using forecasting, reservations, workload prioritization, and automated response to demand and failure scenarios.
- Partner with compute, traffic, networking, storage, and application leaders to enable safe zone evacuation, regional failover, and surge readiness.
Observability, Testing & Reliability Intelligence
- Partner with Observability and TestOps leaders to integrate logs, metrics, traces, continuous profiling, testing, and incident intelligence into one reliability feedback loop.
- Ensure every critical service has actionable telemetry, meaningful SLOs, release-quality signals, and production-readiness evidence.
- Use production incidents and operational patterns to drive targeted integration, load, resilience, and regression testing.
- Establish executive reliability dashboards that provide trusted views of service health, risk, capacity, and operational effectiveness.
Talent Leadership & Organization
- Lead multiple layers of SRE leaders, including senior managers, directors, principal engineers, and senior individual contributors.
- Own organizational design, global hiring strategy, leadership development, succession planning, and the creation of a strong leadership bench.
- Attract exceptional SRE, distributed systems, resilience, incident-management, and capacity-engineering talent from best-in-class technology organizations.
- Build an empowered organization with clear accountability, strong technical judgment, high execution velocity, and a company-wide perspective.
- Act as a force multiplier by mentoring technical and organizational leaders and raising reliability capabilities across engineering.
Technical Leadership & Architecture
- Own reliability architecture decisions across large-scale distributed systems and cloud-native infrastructure.
- Define resilient patterns for redundancy, failover, traffic management, data recovery, workload prioritization, and dependency isolation.
- Guide architecture reviews and platform standards for safe scaling, fault tolerance, and operational simplicity.
- Balance availability, customer impact, engineering velocity, cost, and operational complexity in major technical decisions.
- Maintain sufficient technical depth to challenge assumptions, guide principal engineers, and make high-quality decisions during critical incidents.
Execution & Impact
- Deliver measurable improvements in availability, time to detect, time to mitigate, time to recover, incident recurrence, change-failure rate, and on-call burden.
- Create disciplined operating rhythms, milestones, ownership models, and quarterly targets for strategic reliability programs.
- Increase adoption of common reliability mechanisms and reduce team-specific implementations and manual operations.
- Demonstrate business impact through improved customer experience, reduced outage exposure, stronger peak readiness, and more efficient use of infrastructure capacity.
- Build credibility through predictable delivery, transparent risk management, and objective evidence of reliability improvement.
Essential Qualifications
- Leadership experience in a best-in-class SRE, production engineering, infrastructure reliability, or cloud operations organization at hyperscaler, major cloud provider, global marketplace, leading fintech, or similarly scaled technology company.
- Experience building an SRE practice comparable in maturity to leading industry organizations, rather than operating a traditional support or operations function renamed as SRE.
- Experience with Kubernetes, service mesh, cloud-native platforms, traffic engineering, and large-scale capacity management.
- Experience with chaos engineering, fault injection, regional resilience, and automated disaster recovery.
- Experience building AI-assisted operations, incident intelligence, predictive reliability, or self-healing systems.
- 15+ years of experience in software engineering, infrastructure engineering, distributed systems, or site reliability engineering.
- 8+ years leading large-scale, multi-layer engineering organizations, including senior managers, directors, and senior individual contributors.
- Demonstrated experience personally defining the vision and leading the architecture, build-out, launch, and scaled adoption of a company-wide SRE, reliability, resilience, or autonomous-operations program.
- Prior experience building reliability systems and operating practices for high-scale, high-availability, customer-critical distributed systems.
- Deep expertise in SLOs, error budgets, observability, incident management, disaster recovery, capacity planning, and resilience engineering.
- Proven ability to lead through major incidents while also creating durable mechanisms that prevent recurrence.
- Recognized as a visionary technology and organizational leader who can influence executive stakeholders, align multiple engineering organizations, and attract exceptional talent.
- Proven ability to convert long-term strategy into measurable execution and company-wide adoption.
As published by greenhouse
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location
- LinkedIn Profile optional
- 쿠팡 주식회사와 그 계열사 및 자회사*(이하 합칭하여 ‘회사’라 함)는 본 동의서에 명시된 개인정보 항목 이외의 개인정보를 수집 및 이용하지 않으니 이력서 제출시 동의서 외의 정보는 제출하지 않도록 유의하시기 바랍니다. choose one
- 1. 국적이 어떻게 되십니까? choose one
- 2. 대한민국에서 합법적으로 근무를 하기 위한 비자가 필요합니까? choose one
- 1. (필수) i) 지원자 본인 확인 및 자격 확인, ii) 배경 및 평판조회, iii) 입사지원, iv) 지원 내역 및 합격 여부 확인, v) 지원자 의사소통, vi) 채용 관련 요청 및 문의 처리, vii) 불공정 채용 방지를 목적으로 성명, 연락처, 이메일 주소, 국적, 학력사항, 경력사항, 자격사항, 국가유공자 여부, 친인척의 성명/관계/소속팀, 지원내역을 수집 및 이용합니다. 수집한 개인정보는 채용 결과 통지일로부터 6개월 보관 후 삭제합니다. 본 동의를 거부할 수 있으나, 거부 시 지원이 불가능합니다. choose one
- 2. (선택) 상시 채용 진행 목적으로 성명, 국적, 연락처, 이메일 주소, 학력사항, 경력사항, 자격사항, 국가유공자 여부, 친인척의 성명/관계/소속팀, 지원내역을 수집 및 이용합니다. 수집한 개인정보는 5년 보관 후 삭제합니다. 본 동의를 거부할 수 있으나, 거부 시 상시 채용이 진행되지 않습니다. choose one
- 3. (선택) i) 면접 과정 녹화, ii) 면접 음성 전사 및 기록, iii) 면접 결과 요약, iv) 면접 과정 및 평가결과 검증, v) 채용 절차 개선, vi) 다른 채용 공고에 대한 지원자 적합성 확인, vii) 과거 면접 기록 확인 목적으로 면접 과정의 녹화 영상 및 음성 전사 기록을 수집 및 이용할 수 있습니다. 수집한 개인정보는 채용 결과 통지일로부터 4년 보관 후 삭제합니다. 본 동의를 거부할 수 있으며, 거부하더라도 채용 과정 진행에는 문제가 없습니다. choose one
- 4. (선택) 쿠팡 계열사 및 자회사* 중 당초 지원한 회사를 제외한 회사에게 유사 포지션 채용 검토 및 진행을 목적으로 성명, 국적, 연락처, 이메일 주소, 학력사항, 경력사항, 자격사항, 국가유공자 여부, 친인척의 성명/관계/소속팀, 지원내역을 채용 지원 시점에 제공합니다. 그에 따라 개인정보가 해외로 이전될 수 있음을 알려드립니다. 해외 이전시에는 안전한 전송수단을 통해 전송하며, 제공한 개인정보는 5년 보관 후 삭제합니다. 본 동의를 거부할 수 있으나, 거부 시 지원한 이외 회사에 대한 상시 채용이 진행되지 않습니다. choose one
- 5. (선택) 쿠팡과 쿠팡 계열사 및 자회사*가 위 동의를 통해 수집한 연락처, 이메일을 이용하여 전화, 이메일, SMS 등 전송 매체를 통해 채용알림을 발송할 것입니다. choose one
- 6. (필수) 지원자는 AIM Screening Limited(U.S., privacy@sterlingcheck.com)에게 i) 학력 조회, ii) 자격사항 조회, iii) 근무경력 조회 및 평판 조회, iv) 국제 제재 대상 인물 여부 조회, v) 글로벌 매체, 국제 제재 대상 명단, 사법 기관 또는 금융 감독기관 등에서 발표한 거래제한 대상자 명단 등을 대상으로 문제 소지 여부 조회, vi) 수사 경력 또는 범죄 경력 조회(외국인인 지원자) 목적으로 배경 조회 시 성명, 이메일 주소, 지원 포지션, 지원 법인, 국적을 제공합니다. 해외 이전시에는 안전한 네트워크를 이용해 전송하며, 제공한 개인정보는 제공목적 달성 시까지 보유합니다. 본 동의를 거부할 수 있으나, 거부 시 지원이 불가능합니다. choose one
- 7. (선택) 본인은 동의한 범위에 대해 회사가 아래와 같이 평판조회를 실시하는 것에 동의합니다. 대면면접을 통과한 후보자 중 회사에 지원한 한국인(*)에 대해 평판조회를 추가로 실시할 수 있으며, 회사는 평판조회의 방법을 마련하고 있습니다. choose one
- 1. 본인은 친인척과 쿠팡 그룹의 계열사(이하 쿠팡 그룹사)에서 함께 근무하는 경우, 서로의 업무에 영향을 미칠 수 있는 상황이 잠재적으로 이해상충을 초래할 수 있음을 이해합니다. 따라서 이러한 상황은 People Operations의 사전 승인을 취득해야 합니다. 상기 내용을 기반하여 본인은 다음의 사항을 확인합니다. 나는 쿠팡 그룹사에 재직 중이거나 쿠팡 그룹사의 이사회에 재임 중인 친인척이 전혀 없습니다. choose one
- - 친인척의 이름 / Name of Relative(s): optional
- - 본인과의 관계 / Relationship: optional
- - 재직 중인 조직이나 부서 / Working Org/Team: optional
- 1. 공직자윤리법 제17조 등에 따라 퇴직일로부터 3년간 취업심사대상기관에의 취업이 제한되는 취업심사대상자에 해당합니까? choose one
- 2. 취업심사대상자에 해당하는 경우, 후보자는 관련 법령에 따라 관할 공직자윤리위원회로부터 취업가능결정 또는 취업승인결정을 받아 입사 전까지 심사결과 등 관련 서류를 회사에 제출하여야 하고, 이를 제출하지 않거나 그 기재에 허위 또는 거짓이 있는 경우 및 위원회로부터 취업제한결정 내지 취업불승인결정을 받은 경우 채용이 취소될 수 있음에 동의합니다. choose one · optional
- 「국가유공자 등 예우 및 지원에 관한 법률」 등에 따른 취업지원 대상자(보훈)에 해당하십니까? choose one
- 채용 지원 또는 법령상 의무를 준수하기 위해 최소한의 개인정보를 수집합니다. choose one
- Document Policy choose one