Graphcore logo

Site Reliability Engineering Lead

Graphcore · Austin, Texas, United States

Location
Austin, Texas, United States
Funding
$682M pre-acquisition
Posted
Sep 29, 2026

Graphcore is hiring a Site Reliability Engineering Lead based in Austin, Texas, United States. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.

Apply directly at Graphcore

Role details

About Graphcore How often do you get the chance to build a technology that transforms the future of humanity? Graphcore products have set the standard in made-for-AI compute hardware and software, gaining global attention and industry acclaim. Now we are developing the next generation of artificial intelligence compute with systems that will allow AI researchers to develop more advanced models, help scientists unlock exciting new discoveries, and power companies around the world as they put AI at the heart of their business. Graphcore recently joined SoftBank Group – bringing large and ongoing investment from one of the world’s leading backers of innovative AI companies. Job Summary We are seeking an experienced Site Reliability Engineering leader to build and lead a new SRE organization responsible for the production operation of a rapidly scaling AI supercomputing platform. The environment combines highly customized compute, high-performance networking, storage and supporting infrastructure, and will grow through multiple phases of deployment. This is a rare opportunity to establish the reliability function for a new platform from the ground up. The platform and its operational model are being developed in parallel and will ultimately support a 24x7x365 production service with stringent availability requirements. You will take the SRE organization from initial formation through production launch, stabilization and scale. This includes hiring and developing the team, defining the operating model, establishing production readiness and incident-management practices, and ensuring reliability and operability are engineered into the platform from the outset. SRE is responsible for the operational capability required to run the platform reliably in production, while partnering with engineering teams that remain accountable for the reliability and operability of the systems they build. This is not a purely managerial position. During the development and early production phases, the SRE Manager will be expected to work directly with engineering teams, develop a deep understanding of the platform, and participate in troubleshooting and incident response. Over time, success will increasingly mean building the people, processes, automation, tooling, and operational discipline that allow the organization to operate effectively without depending on you for day-to-day escalation. Responsibilities and Duties Build the SRE Organization: Build and develop the team from its initial formation through full 24x7x365 production operations, including defining roles, interviewing and hiring team members, establishing career expectations, and developing future technical leaders. Mentor engineers and team leads, develop successors, and build an organization capable of operating effectively without depending on any single individual. Work with leadership to forecast staffing requirements as the platform grows from initial deployment through full production scale. Establish the Production Operating Model: Define the operating model for an SRE organization, including staffing and coverage model, escalation paths, on-call responsibilities, incident management, handoffs, production access, change management, and operational readiness requirements. Establish clear operational interfaces with Datacenter Operations, engineering teams, vendors, and other service owners. Establish and continuously improve production readiness standards, runbooks, operational procedures, failure-mode documentation, escalation processes, and incident response practices with a strong emphasis on automation and engineering over manual operational work. Develop training, cross-training, simulation, and production incident-response exercises to ensure the team can operate independently and confidently. Engineer Reliability Into the Platform: Embed with platform engineering teams during development to gain deep technical knowledge of the system and ensure reliability, serviceability, observability, and operational requirements are incorporated into the platform before production. Lead the development of SLOs, operational health indicators, alerting standards, incident severity definitions, and reliability reporting appropriate for a large-scale production infrastructure service. Build a culture in which recurring operational problems are engineered out through automation, improved observability, better platform design, and elimination of unnecessary toil. Ensure the SRE organization can rapidly diagnose and mitigate issues across compute, networking, storage, and supporting infrastructure. Develop strong technical depth within the team while maintaining access to specialist expertise in critical areas such as high-performance networking, storage, observability, and platform scheduling. Lead Production Operations: Lead or participate in major production incidents as necessary, particularly during platform development, launch, and early production. Establish a blameless post-incident review process focused on identifying systemic improvements and ensuring corrective actions are completed. Serve as the senior operational authority for the SRE organization and represent production reliability concerns in engineering and leadership discussions. Required Skills and Experience Significant experience leading or building an SRE, Production Engineering, Infrastructure Reliability, or comparable function supporting large-scale, highly available production infrastructure. Experience taking a new or rapidly evolving platform through production readiness, launch, stabilization, and ongoing operation. Experience building and operating sustainable 24x7x365 production support or on-call organizations. Strong understanding of modern Site Reliability Engineering principles, including SLOs, incident management, observability, automation, toil reduction, capacity management, and production readiness. Strong incident leadership experience, including managing high-sev

More roles at Graphcore

Search all engineering roles