Senior Software Engineer, SRE (Bengaluru)

Senior Software Engineer, SRE (Bengaluru)

03 Sep
|
Roku
|
Bengaluru

03 Sep

Roku

Bengaluru

Teamwork makes the stream work.

Roku is changing how the world watches TV

Roku is the #1 TV streaming platform in the U.S., Canada, and Mexico, and we've set our sights on powering every television in the world. Roku pioneered streaming to the TV. Our mission is to be the TV streaming platform that connects the entire TV ecosystem. We connect consumers to the content they love, enable content publishers to build and monetize large audiences, and provide advertisers unique capabilities to engage consumers.

From your first day at Roku, you'll make a valuable - and valued - contribution. We're a fast-growing public company where no one is a bystander. We offer you the opportunity to delight millions of TV streamers around the world while gaining meaningful experience across a variety of disciplines.

What does the team work on?

The Platform Infrastructure team ensures that all Roku systems run smoothly. These systems support over 100M+ users and billions in transaction revenue per year. We are a group of highly skilled infrastructure and software engineers who help build and operate systems at internet scale, including Platform (Kubernetes, Istio, Envoy, operators, and more) and Observability (OSS/CNCF-supported observability projects). We engage with multiple teams to achieve company-impacting results.

What is the role?

We are seeking a talented and experienced SRE (Site Reliability Engineering) Senior Software Engineer to help architect, build, and operate large-scale systems that stay reliable, secure, and cost-effective at internet scale. The ideal candidate takes end-to-end ownership of outcomes—treating reliability, security, cost, operability, and supportability as part of the job, not just delivering code. They bring calm, decisive incident leadership, separating mitigation from root-cause investigation and running blameless reviews that produce lasting improvements. Strong judgment and prioritization are essential, balancing roadmap delivery against operational debt, security, and compliance while clearly explaining trade-offs. This engineer pairs deep technical depth in distributed systems with broad systems thinking, and turns ambiguous objectives into executable roadmaps, epics, and backlogs. Just as important is the ability to build influence through credibility and sound reasoning, coach other engineers, and raise the operational capability of the whole team. If you enjoy solving intriguing system challenges, are innovative at heart, and thrive on making a measurable impact across teams, this role might be a great fit for you.

How will I use AI at Roku?

At Roku, we don't just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability.

We value curious, adaptable builders, who can show how they've used AI, agents, or automation to move faster, improve quality, and scale their impact. Robust candidates know how to frame problems, guide AI-assisted work, check the output, and learn quickly. Above all, they bring curiosity, adaptability, and sound judgment.

What are the responsibilities of the role?

Ownership & Incident Leadership

- Take responsibility for service outcomes end to end, including reliability, security, cost, operability, and supportability
- Lead major incidents with composure when information is incomplete, separating mitigation from root-cause investigation and communicating impact, status, risks, and next steps without speculation
- Facilitate comprehensive, blameless post-incident reviews that identify root causes and contributing factors, and follow corrective actions through to completion
- Track incident trends to surface systemic issues and prioritize reliability improvements
- Implement chaos engineering, game days, and disaster recovery exercises to validate resilience and build confidence in recovery procedures

SRE Process & Principles Implementation

- Establish and evolve SRE principles, frameworks, and methodologies across the organization, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets
- Manage Error Budgets as a data-driven mechanism for balancing feature velocity against reliability, and facilitate risk-tolerance conversations between engineering and product teams
- Use SLOs, error budgets, incident data, and operational metrics to guide where the team invests

Reliability Engineering & Infrastructure

- Reduce toil by identifying repetitive operational work and eliminating it through infrastructure-as-code, automation frameworks,



and intelligent tooling, aiming to keep toil below 50% of team time
- Prefer sustainable corrective action over repeated manual intervention
- Implement capacity planning that ensures adequate headroom to meet SLOs during peak traffic, load spikes, and degraded states, including predictive models and automated scaling

Observability, Monitoring & Reporting

- Build observability systems that provide deep visibility into service health, performance, and user experience using the Four Golden Signals and USE/RED methodologies
- Create SRE dashboards and reporting that give real-time visibility into SLO compliance, error budget consumption, and reliability metrics, including executive-level reporting on trends and incident impact
- Establish actionable, symptom-based alerting aligned with SLOs, tuning thresholds to reduce noise while ensuring critical issues trigger appropriate responses
- Treat an alert without ownership or a usable runbook as an incomplete control

Collaboration, Influence & Leadership

- Partner with development teams to design reliability in from the start, conducting design reviews focused on failure modes, scalability, observability, and operational concerns
- Build alignment across product, application, security, compliance, and infrastructure teams, managing disagreements with evidence, impact, and trade-offs
- Gain support through credibility, relationships, and sound reasoning rather than title or tenure, and lead cross-functional work where participants do not report to you
- Coach engineers by delegating meaningful ownership, giving specific and timely feedback, and creating opportunities for others to lead projects and incidents
- Manage project priorities using error budgets as a decision-making framework, ensuring reliability work is prioritized alongside feature development

Delivery & Execution Discipline

- Convert roadmap objectives into clear epics, stories, milestones, owners, and acceptance criteria, identifying dependencies and risks early
- Track outcomes rather than activity, and adjust scope transparently when incidents or unplanned work affect delivery
- Finish work completely, including testing, documentation, monitoring, runbooks, rollout, and operational handoff

Operational Excellence & Continuous Improvement

- Identify and eliminate performance bottlenecks through analysis of metrics, traces, and profiles, and optimize resources, configurations, and auto-scaling for SLO compliance
- Drive continuous improvement by analyzing SLO violations, incident trends, and toil metrics to champion the reliability roadmap and technical-debt reduction
- Maintain a culture of documentation and knowledge sharing through runbooks, operational guides, architecture documentation, and disaster recovery procedures
- Track and report on SRE metrics, including SLO compliance, error budget consumption, Mean Time To Detection (MTTD), Mean Time To Resolution (MTTR), toil percentage, and reliability improvement velocity

Security & Risk Awareness

- Incorporate least privilege, secure defaults, auditability, and compliance into designs, and express risk in terms stakeholders can understand
- Avoid normalizing manual changes, undocumented exceptions, or permanent emergency access

On-call & Reliability

- Participate in a 24x7 on-call rotation and be available to work with global teams during critical outages

What experience would help someone be successful in this role at Roku?
- Preferably 12+ years of experience in DevOps/SRE roles, with demonstrated expertise implementing SRE principles, SLO/SLI frameworks, and error budget policies in production environments
- Deep experience with observability and monitoring platforms such as Prometheus, Grafana, Datadog, or New Relic, including building custom dashboards, alerts, and SLO-based monitoring
- Strong background in incident management, including experience as an Incident Commander, conducting blameless postmortems, and turning incident learnings into systemic improvements
- Strong understanding of distributed systems and reliability engineering, including failure modes, fault tolerance patterns, circuit breakers, bulkheads, rate limiting,



and graceful degradation
- Experience with several of the following: Kubernetes, Docker, and service mesh technologies such as Istio, Envoy, Linkerd, Solo, and Elastic Container Service (ECS)
- Experience in cloud-focused software development, preferably in Go, Python, or other object-oriented languages
- Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, or CloudFormation
- Experience with Continuous Integration/Continuous Delivery (CI/CD) automation, including GitLab pipelines and related tooling
- Strong hands-on experience with cloud platforms such as AWS, GCP, or Azure
- Proven track record of delivering scalable, high-performance infrastructure solutions in fast-paced, dynamic environments
- Demonstrated ability to communicate clearly with both technical and non-technical stakeholders and to work effectively across cross-functional teams
- Demonstrated influence without authority—improving outcomes across multiple teams and establishing standards that continued without your direct involvement
- Intellectual humility and curiosity, including comfort saying "I don't know, but here's how I'd find out," and changing position when presented with better evidence
- Self-driven and detail-oriented, with the ability to understand complex distributed systems and identify reliability risks proactively
- Certifications in relevant technologies, such as Certified Kubernetes Administrator (CKA), AWS Certified DevOps Engineer, or Certified Information Systems Security Professional (CISSP), are preferred
- BS degree in Computer Science or equivalent

#LI-SK8 What's Roku's approach to hybrid working?

Roku fosters an inclusive and collaborative environment where teams generally work in the office Monday through Thursday. Fridays are generally flexible for remote work, except for employees whose specific roles or assigned office location require five days' a week attendance.

What are some of the benefits?

Roku is committed to offering a diverse range of benefits as part of our compensation package to support our employees and their families. Our comprehensive benefits include global access to mental health and financial wellness support and resources. Local benefits include statutory and voluntary benefits which may include healthcare (medical, dental, and vision), life, accident, disability, commuter, and retirement options (401(k)/pension). Employees are supported in taking time off, in accordance with local leave policies and other personal needs to support their evolving work and life needs. It's important to note that not every benefit is available in all locations or for every role. For details specific to your location, please consult with your recruiter.

Accommodations

Roku welcomes applicants of all backgrounds and provides reasonable accommodations and adjustments in accordance with applicable law. If you require reasonable accommodation at any point in the hiring process, please direct your inquiries to [email protected].

What should I know about Roku's culture?

Roku is a great place for people who want to work in a fast-paced environment where everyone is focused on the company's success rather than their own. We try to surround ourselves with people who are great at their jobs, who are easy to work with, and who keep their egos in check. We appreciate a sense of humor. We believe a fewer number of very talented folks can do more for less cost than a larger number of less talented teams. We're independent thinkers with big ideas who act boldly, move fast and accomplish extraordinary things through collaboration and trust. In short, at Roku you'll be part of a company that's changing how the world watches TV.

We have a unique culture that we are proud of. We think of ourselves primarily as problem-solvers, which itself is a two-part idea. We come up with the solution, but the solution isn't real until it is built and delivered to the customer. That penchant for action gives us a pragmatic approach to innovation, one that has served us well since 2002.

To learn more about Roku, our global footprint, and how we've grown, visit https://www.weareroku.com/factsheet.

By providing your information, you acknowledge that you want Roku to contact you about job roles, that you have read Roku's Applicant Privacy Notice, and understand that Roku will use your information as described in that notice. If you do not wish to receive any communications from Roku regarding this role or similar roles in the future, you may unsubscribe at any time by emailing [email protected].

📌 Senior Software Engineer, SRE (Bengaluru)
🏢 Roku
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior software engineer, sre (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior software engineer, sre (bengaluru) / bengaluru