Sr. Director, Back-End Engineering

Coupanginternal · Seoul, South Korea

Location
Seoul, South Korea
Funding
N/A
Posted
Aug 23, 2026

Coupanginternal is hiring a Sr. Director, Back-End Engineering based in Seoul, South Korea. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.

Apply directly at Coupanginternal

Role details

Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define and lead company-wide reliability, resilience, scalability, and operational excellence. This leader will transform reliability from a collection of team-specific practices into platform mechanisms that services inherit by tier, while advancing incident response toward an intelligent, AI-assisted, and increasingly autonomous operating model. We are looking for a visionary, industry-recognized technology leader who has previously conceived, built, and scaled a comparable SRE, production engineering, resilience, or autonomous-operations organization at a leading global technology company. The successful candidate must combine deep technical credibility with the organizational leadership required to align executives, influence architecture across the company, and build a world-class leadership bench. Key Responsibilities Set a bold, multi-year vision for company-wide reliability, resilience, and autonomous operations, and translate that vision into an executable roadmap with measurable business outcomes. Define and own the SRE strategy, operating model, engineering standards, and reliability governance across Coupang. Build platform mechanisms that allow services to inherit reliability requirements based on service tier rather than recreate them independently. Lead initiatives that materially improve availability, resilience, scalability, performance, and operational readiness. Partner with engineering, product, infrastructure, security, finance, and business leaders to align reliability investments with customer and business priorities. Own executive reliability metrics, including availability, detection and recovery performance, change risk, incident recurrence, capacity readiness, and operational toil. Build and scale a world-class SRE organization capable of influencing engineering practices across the company. Reliability Strategy, SLOs & Engineering Governance Establish and evolve service-tier definitions, SLOs, SLAs, error budgets, reliability scorecards, and objective certification mechanisms such as RBD/RBO. Create clear reliability requirements for Tier 0, Tier 1, and Tier 2 services, including redundancy, load testing, disaster recovery, observability, and incident response. Ensure reliability governance is embedded in architecture, development, release, and production operations rather than applied as a final review. Drive systematic reduction of recurring incidents, reliability risks, operational debt, and unsafe change patterns. Influence company-wide architecture for graceful degradation, fault isolation, load shedding, circuit breaking, and failure containment. Incident Management & Autonomous Operations Transform incident management into a fast, disciplined, data-driven, and increasingly autonomous operating model. Enable AI-assisted detection, event correlation, triage, escalation, root-cause drafting, remediation recommendations, and selected guardrailed auto-remediation. Improve incident command, on-call quality, escalation mechanisms, communication, post-incident learning, and corrective-action completion. Reduce noisy alerts, manual on-call work, repeated diagnosis, and time spent coordinating across fragmented systems. Use incident and telemetry data to continuously improve platform standards, testing, capacity models, and engineering roadmaps. Disaster Recovery, Resilience & Capacity Own the strategy and execution model for disaster recovery, regional resilience, availability-zone loss, capacity-constrained recovery, and critical business continuity. Build reusable DR and failover mechanisms that services inherit from the platform rather than implement as bespoke projects. Establish objective RPO/RTO targets, automated readiness gates, regular game days, fault injection, and evidence-based recovery certification. Drive proactive and intelligent capacity management using forecasting, reservations, workload prioritization, and automated response to demand and failure scenarios. Partner with compute, traffic, networking, storage, and application leaders to enable safe zone evacuation, regional failover, and surge readiness. Observability, Testing & Reliability Intelligence Partner with Observability and TestOps leaders to integrate logs, metrics, traces, continuous profiling, testing, and incident intelligence into one reliability feedback loop. Ensure every critical service has actionable telemetry, meaningful SLOs, release-quality signals, and production-readiness evidence. Use production incidents and operational patterns to drive targeted integration, load, resilience, and regression testing. Establish executive reliability dashboards that provide trusted views of service health, risk, capacity, and operational effectiveness. Talent Leadership & Organization Lead multiple layers of SRE leaders, including senior managers, directors, principal engineers, and senior individual contributors. Own organizational design, global hiring strategy, leadership development, succession planning, and the creation of a strong leadership bench. Attract exceptional SRE, distributed systems, resilience, incident-management, and capacity-engineering talent from best-in-class technology organizations. Build an empowered organization with clear accountability, strong technical judgment, high execution velocity, and a company-wide perspective. Act as a force multiplier by mentoring technical and organizational leaders and raising reliability capabilities across engineering. Technical Leadership & Architecture Own reliability architecture decisions across large-scale distributed systems and cloud-native infrastructure. Define resilient patterns for redundancy, failover, traffic management, data recovery, workload prioritization, and dependency isolation. Guide architecture rev

More roles at Coupanginternal

Search all engineering roles