13 Sep
|
Roku
|
Bengaluru
Job Summary
We are seeking a Senior SRE / Platform Engineer (Infrastructure) to join the Ads Business Automation team. This is a product-embedded infrastructure role where you will actively contribute to application and ETL code while owning critical platform capabilities such as cloud infrastructure, Kubernetes environments, CI/CD systems, observability, networking, and security. This role blends software engineering, infrastructure engineering, and reliability, and is ideal for engineers who enjoy building and operating systems end-to-end as part of a product team rather than working in a centralized operations function.
About the team
The Ads Customer Interfaces team builds full-stack web applications, APIs, UIs, and ETL pipelines that serve as the primary interface between Roku s advertising platform and both internal and external customers. Our mission is to deliver a best-in-class user experience by simplifying complex business workflows, enabling users to focus on their customers rather than tedious processes. As a member of this team, you will work closely with product, data, and engineering partners to build reliable, scalable systems that power Roku s advertising business.
About the role
What you ll be doing
- Write clean, maintainable Python for services, automation, and infrastructure tooling.
- Design, migrate and operate Kubernetes clusters (EKS / GKE) in production.
- Lead cluster upgrades, workload migrations, autoscaling and capacity planning.
- Implement safe deployment strategies (rolling, canary, blue/green).
- Manage Infra as Code (Terraform or equivalent) and fully checked into version control.
- Operate multi-environment (dev/stage/prod) and multi-region setups on AWS and/or GCP.
- Build and maintain CI/CD pipelines for large monorepos.
- Support deployments for Web applications , Background workers and ETL and batch pipelines
- Improve release safety, rollback mechanisms, and developer velocity.
- Design and maintain telemetry across services using Metrics, logs, and traces, Grafana, Prometheus, OpenTelemetry or equivalent
- Set up PagerDuty alerts, on-call workflows, and incident response processes.
- Define and track SLIs, SLOs, and service health indicators
- Support GenAI-powered services used in Ads automation.
- Implement observability for LLM systems, including: Latency, throughput, error rates and Infrastructure-level reliability and capacity signals
We are excited to have you if you have:
- 8+ years of experience in SRE, Platform Engineering, Infrastructure or Backend Engineering roles.
- Solid Python proficiency for services, pipelines, and automation.
- Hands-on production experience with Kubernetes.
- Experience working with AWS or GCP.
- Linux systems
- Networking (DNS, VPCs, routing)
- Distributed systems
- Experience with CI/CD pipelines, monorepos, and infrastructure-as-code.
- Experience building or operating observability and alerting systems.
- Experience with Airflow or large-scale ETL/data systems.
- Experience supporting GenAI / ML infrastructure in production.
- Prior on-call ownership for production systems.
Key Skills
- Python
- Kubernetes
- AWS
- GCP
- Linux
- Networking
- CI/CD
- Observability
- Airflow
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Senior Software Engineer - Devops (Bengaluru)
🏢 Roku
📍 Bengaluru