Staff Site Reliability Engineer
Crunchyroll · Los Angeles, California, United States
- Location
- Los Angeles, California, United States
- Salary
- $210,500 - $263,100 USD
- Experience
- 12+ years
- Funding
- N/A
- Posted
- Sep 9, 2026
Crunchyroll is hiring a Staff Site Reliability Engineer based in Los Angeles, California, United States. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.
Apply directly at CrunchyrollRole details
About Crunchyroll Founded by fans, Crunchyroll delivers the art and culture of anime to a passionate community. We super-serve over 100 million anime and manga fans across 200+ countries and territories, and help them connect with the stories and characters they crave. Whether that experience is online or in-person, streaming video, theatrical, games, merchandise, events and more, it’s powered by the anime content we all love. Join our team, and help us shape the future of anime! About the role We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US and play a critical role in advancing the reliability, scalability, performance, and security of Crunchyroll's consumer-facing data platforms. As a senior technical leader, you will partner closely with Engineering, Data, Infrastructure, Product, and Security teams to design and operate resilient cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability, incident management, automation, capacity planning, disaster recovery, and operational excellence while helping teams adopt modern SRE practices such as SLIs, SLOs, and error budgets. The ideal candidate combines deep expertise in large-scale distributed systems with a strong sense of ownership, collaboration, and service leadership. You are passionate about building highly reliable platforms, eliminating operational toil through automation, and enabling engineering teams to move quickly and safely. In addition, you will champion SecOps best practices by driving vulnerability management, supporting penetration testing initiatives, improving security observability, strengthening cloud and Kubernetes security controls, and ensuring operational readiness for emerging threats. This is a unique opportunity to shape reliability and security engineering practices across CDI while helping build a world-class data and insights ecosystem that enables informed decision-making throughout Crunchyroll. Core Areas of Responsibility Reliability Engineering : Define, measure, and continuously improve the reliability, availability, and performance of CDI platforms through SLIs, SLOs, and error budgets. Operational Excellence : Establish and drive best practices for incident management, root cause analysis, postmortems, and service ownership across engineering teams. Observability & Monitoring : Build and evolve comprehensive monitoring, logging, tracing, and alerting capabilities to enable proactive issue detection and rapid resolution. Automation : Identify operational inefficiencies and develop automation, self-service capabilities, and self-healing mechanisms to improve engineering productivity. Platform Scalability : Design and optimize cloud-native infrastructure and services to support growing business demands while maintaining performance and cost efficiency. Infrastructure Engineering : Drive Infrastructure as Code (IaC), platform standardization, and deployment automation to improve consistency, reliability, and operational agility. Capacity Planning & Performance : Lead capacity planning and performance optimization initiatives to ensure platforms can scale predictably and efficiently. Disaster Recovery & Resilience : Develop and regularly validate disaster recovery, backup, and business continuity strategies to ensure platform resiliency. Security Operations (SecOps) : Partner with Crunchyroll's security team to integrate security controls, operational risk management, and security best practices into platform operations and engineering workflows. Vulnerability Management : Own the triage and remediation of identified vulnerabilities across infrastructure, platform, container, and application security vulnerabilities through established Crunchyroll vulnerability management processes. Penetration Testing & Security Remediation : Support penetration test scoping activities by providing technical context on CDI platforms. Own the triage, prioritization, and remediation of resulting findings to drive timely resolution and strengthen platform security posture. Cloud & Kubernetes Security : Implement and maintain secure cloud, container, and Kubernetes environments following least-privilege, defense-in-depth, and Zero Trust principles. Cross-Functional Leadership : Collaborate with Engineering, Data, Product, Infrastructure, and Security teams to drive reliability, scalability, and security initiatives across CDI. Mentorship & Engineering Excellence : Mentor engineers and champion a culture of operational excellence, reliability, ownership, continuous improvement, and security awareness. About You We get excited about candidates like you, because… 12+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, or related disciplines, with a proven track record of operating and scaling production-critical systems. Deep expertise in Kubernetes and GCP , including the design, deployment, and operation of highly available, cloud-native platforms at scale. Strong Infrastructure as Code (IaC) experience , preferably with Terraform, and a commitment to automation, standardization, and operational efficiency. Solid foundation in Linux systems administration, networking, and distributed systems , with the ability to troubleshoot complex production issues across multiple layers of the technology stack. Proficiency in one or more programming and scripting languages , such as Go, Python, Java, or Shell, with a focus on automation and platform engineering. Hands-on experience with modern observability platforms and practices , including Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent monitoring and telemetry solutions. Demonstrated expertise in incident management, service reliability, capacity planning, performance optimization, and operational excellence , including the implementation of SLIs, SLOs, and error budgets. Strong understanding o
More roles at Crunchyroll
- Senior Site Reliability Engineer2 days agoHyderabad, Telangana, India
- Software Engineer III, Playback Services18-08-2026San Francisco, CA, United States$16,900 - $205,000 USD
- Software Engineer II, Frontend10-08-2026San Francisco, CA, United States$135,100 - $168,900 USD
- Software Engineer III, Service Monetization03-08-2026Los Angeles, California, United States$152,400 - $190,500 USD
- Principal Engineer, Engineering Practices28-07-2026San Francisco, CA, United States$260,400 - $325,500 USD
- Staff Software Engineer21-07-2026Hyderabad, Telangana, India
- Staff AI Engineer13-07-2026Dallas, Texas, United States
- Senior AI Engineer02-07-2026Dallas, Texas, United States
- Senior Software Engineer - Android, Partner Engineering30-06-2026Hyderabad, Telangana, India
- Senior Engineering Manager, Service Monetization24-06-2026Los Angeles, California, United States$218,500 - $273,100 USD
- Software Engineer I, AndroidEarly Career09-06-2026Mexico City, Mexico City, Mexico
- Staff Software Engineer, AI/ML09-06-2026Hyderabad, Telangana, India