10 Sep
|
Media.net
|
Bengaluru
10 Sep
Media.net
Bengaluru
Job Description
Senior Manager - Site Reliability Engineer
n
n
About the Role
n
We are looking for a Senior Manager - Site Reliability Engineering to lead a team of 8+ Site Reliability Engineers within our core platform engineering organization. The team supports a high-throughput, low-latency distributed platform that processes more than 500 billion requests daily, with stringent performance requirements where every millisecond matters.
n
As a leader, you will own the reliability charter for this platform: setting technical direction, growing engineers, partnering with product and development leadership, and ensuring the infrastructure is resilient, scalable, and always available.
n
The team is also evolving rapidly toward agentic AI. You should be well-versed in the contemporary AI ecosystem and capable of driving the adoption of AI-powered tools, workflows, and automation frameworks that improve engineering productivity, operational excellence, and platform reliability.
n
n
What You'll Do People Leadership & Team Management
n
• Lead, mentor, and grow a team of 8-10 SREs across varying experience levels, fostering a culture of ownership, learning, and psychological safety.
n
• Drive AI adoption, AIOps, intelligent automation, and agentic AI initiatives. • Conduct regular 1:1s, performance reviews, and career development conversations; define clear growth paths for individual contributors.
n
• Drive hiring: partner with recruiting to define role requirements, lead interviews, and build a diverse, highperforming SRE team.
n
• Manage team workload, priorities, and on-call rotations to prevent burnout while maintaining operational excellence.
n
• Champion SRE best practices, including error budgets, SLOs/SLIs/SLAs, and blameless post-mortems, and instill a data-driven reliability culture
n
n
Infrastructure Strategy & Ownership
n
• Own the end-to-end reliability roadmap for a high-scale, real-time platform, including load balancers, data stores, CI/CD pipelines, observability stacks, and networking layers.
n
• Define and drive multi-quarter infrastructure programs focused on resilience, scalability, performance, and cost efficiency to support more than 500 billion requests daily.
n
• Develop and enforce policies and procedures that improve platform stability; represent the team in architectural reviews and cross-functional planning.
n
• Establish infrastructure-as-code standards using tools such as Terraform, Puppet, or Ansible and ensure consistent adoption across the team.
n
• Bring deep, hands-on technical expertise in Kubernetes, including cluster design, multi-tenancy, networking, autoscaling, upgrades, and operational best practices at scale.
n
• Drive architectural ownership across the SRE landscape, making high-judgment decisions on platform design, service topology, and infrastructure trade-offs.
n
• Define and evolve the cloud and container strategy, including Kubernetes, service mesh, CI/CD, and observability, to support rapid product growth and engineering velocity.
n
• Lead capacity planning, performance engineering, and resilience initiatives, including disaster recovery, chaos engineering, and multi-region availability.
n
n
Cross-Functional Collaboration
n
• Partner closely with software engineering leads to set quality and performance benchmarks; ensure production-readiness criteria are met before every deployment.
n
• Participate in system design reviews, providing leadership-level input on infrastructure reliability, security, and operability.
n
• Serve as the primary point of escalation for high-severity incidents; coordinate cross-team response and ensure transparent communication to stakeholders.
n
• Collaborate with product management to balance feature velocity against reliability investments, using error budgets as a shared language.
n
n
Tooling, Automation & Performance
n
• Drive the vision for internal SRE tooling, including monitoring, alerting, deployment automation, failure detection, and self-healing systems.
n
• Champion full-stack performance optimization, from request handling and network paths to database query tuning,
with a focus on sub-millisecond latency targets.
n
• Ensure robust observability across all services using Prometheus, Grafana, or the ELK stack; establish dashboards and alerts that provide real-time operational insight.
n
• Oversee CI/CD pipeline health using tools such as Jenkins and Argo CD, and continuously improve deployment speed, safety, and rollback capabilities.
n
n
Who Should Apply Required Qualifications
n
• B.Tech, M.Tech, or equivalent degree in Computer Science, Information Technology, or a related field.
n
• 13+ years of overall experience in SRE, infrastructure engineering, or DevOps, including at least 2 years in a people management role leading teams of 8+ engineers.
n
• Proven track record of building and scaling SRE teams supporting high-volume, latency-sensitive distributed systems.
n
• Strong technical foundation in networking concepts, including TCP/IP, routing, and SDN, as well as modern software architectures.
n
• Hands-on proficiency in Python, Go, or Ruby, with a strong automation mindset.
n
• Deep experience with Kubernetes and container orchestration; familiarity with GCP or AWS, in addition to operating large-scale co-location data centers.
n
• Strong understanding of agentic AI, AIOps, MCP, A2A, and AI automation frameworks.
n
• Demonstrated ability to define and operate against SLOs, error budgets, and reliability metrics at scale.
n
• Ability to independently own complex problem statements, set team priorities, and drive solutions end to end.
n
n
Preferred Skills & Tool Expertise
n
• Infrastructure as Code: Terraform, Puppet, or Ansible.
n
• Monitoring & Logging: Prometheus, Grafana, and the ELK stack.
n
• CI/CD Pipelines: Jenkins and Argo CD.
n
• Databases: MySQL, Redis, Aerospike, HBase, or similar technologies.
n
• Web Proxies & Service Networking: Envoy and Nginx.
n
• Version Control: Git-based workflows at team scale.
n
• Experience operating high-throughput, low-latency platforms or large-scale real-time distributed systems is a strong plus.
n
• Familiarity with incident management frameworks, PagerDuty ecosystems, and a blameless post-mortem culture.
📌 Senior Manager - Site Reliability Engineer|-0246 (Bengaluru)
🏢 Media.net
📍 Bengaluru