09 Sep
|
Crunchyroll
|
Hyderabad
09 Sep
Crunchyroll
Hyderabad
Job Summary
We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data &
- Insights (CDI) in India and play a critical role in advancing the reliability, scalability, performance, and security of Crunchyroll's consumer-facing data platforms. As a senior technical leader, you will partner closely with Engineering, Data, Infrastructure, Product, and Security teams to design and operate resilient cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability, incident management, automation, capacity planning, disaster recovery, and operational excellence while helping teams adopt modern SRE practices such as SLIs, SLOs, and error budgets. The ideal candidate combines deep expertise in large-scale distributed systems with a strong sense of ownership, collaboration, and service leadership. You are passionate about building highly reliable platforms, eliminating operational toil through automation, and enabling engineering teams to move quickly and safely. In addition, you will champion SecOps best practices by driving vulnerability management, supporting penetration testing initiatives, improving security observability, strengthening cloud and Kubernetes security controls, and ensuring operational readiness for emerging threats. This is a unique opportunity to shape reliability and security engineering practices across CDI while helping build a world-class data and insights ecosystem that enables informed decision-making throughout Crunchyroll.
Responsibilities
- Reliability Engineering: Define, measure, and continuously improve the reliability, availability, and performance of CDI platforms through SLIs, SLOs, and error budgets.
- Operational Excellence: Establish and support best practices for incident management, root cause analysis, postmortems, and service ownership across engineering teams.
- Observability &
- Monitoring
: Build and evolve comprehensive monitoring, logging, tracing, and alerting capabilities to enable proactive issue detection and rapid resolution.
- Automation: Develop automation, self-service capabilities, and self-healing mechanisms to improve engineering productivity.
- Platform Scalability: Design and optimize cloud-native infrastructure and services to support growing business demands while maintaining performance and cost efficiency.
- Infrastructure Engineering:
Implement Infrastructure as Code (IaC), platform standardization, and deployment automation to improve consistency, reliability, and operational agility.
- Capacity Planning &
- Performance
: Participate in capacity planning and performance optimization initiatives to ensure platforms can scale predictably and efficiently.
- Disaster Recovery &
- Resilience
: Develop and regularly validate disaster recovery, backup, and business continuity strategies to ensure platform resiliency.
- Security Operations (SecOps): Partner with Crunchyroll's security team to integrate security controls, operational risk management, and security best practices into platform operations and engineering workflows.
- Vulnerability Management: Support the triage and remediation of identified vulnerabilities across infrastructure, platform, container, and application security vulnerabilities through established Crunchyroll vulnerability management processes.
- Penetration Testing &
- Security Remediation
: Support penetration test scoping activities by providing technical context on CDI platforms. Own the triage, prioritization, and remediation of resulting findings to drive timely resolution and strengthen platform security posture.
- Cloud &
- Kubernetes Security
: Implement and maintain secure cloud, container, and Kubernetes environments following least-privilege, defense-in-depth, and Zero Trust principles.
- Cross-Functional Collaboration: Collaborate with Engineering, Data, Product, Infrastructure, and Security teams to drive reliability, scalability, and security initiatives across CDI.
About You
We get excited about candidates like you, because
- 8+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, or related disciplines, with a proven track record of operating and scaling production-critical systems.
- Strong experience with Kubernetes and GCP, including deploying, operating, and troubleshooting cloud-native applications and services at scale.
- Excellent Infrastructure as Code (IaC) Experience in implementing IaC solutions, preferably using Terraform, to improve automation, consistency, and operational efficiency.
- Systems &
- Networking Fundamentals
: Solid understanding of Linux systems administration, networking fundamentals, and distributed systems concepts
- Programming &
- Automation
: Proficiency in one or more programming or scripting languages such as Go, Python, Java, or Shell
- Observability &
- Monitoring
: Experience with up-to-date observability platforms, including Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent monitoring and telemetry solutions.
- Service Reliability &
- Operations
: Working knowledge of incident management, service reliability practices, capacity planning, performance optimization, and operational excellence, including the use of SLIs and SLOs.
- Platform Security: Experience supporting cloud and platform security initiatives, including container security, Kubernetes security, CI/CD security, vulnerability remediation, and secure infrastructure operations.
- Security Best Practices: Familiarity with industry-standard security frameworks and practices, including OWASP Top 10, Identity and Access Management (IAM), secrets management, SSDLC, and security-by-design principles.
- Collaboration &
- Communication
: Strong communication, collaboration, and problem-solving skills, with the ability to work effectively across teams and contribute to reliability and operational excellence initiatives.
About the Team
The Center for Data and Insights (CDI) is a service-oriented, horizontal organization uniquely positioned within the company to serve as the trusted, unbiased source of timely, data-driven insights for Crunchyroll. Our vision is to inspire, support, and guide our stakeholders to be data-aware and build the systems of intelligence to discover insights and act on them. We have built a highly functional organization that truly believes in being a responsive partner, with the utmost curiosity, unwavering accountability, and the courage to lead with actions.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Sr Site Reliability Engineer (Hyderabad)
🏢 Crunchyroll
📍 Hyderabad