Lead Platform Engineer - Observability (Bengaluru)

Lead Platform Engineer - Observability (Bengaluru)

10 Aug
|
Sony India Software Centre
|
Bengaluru

10 Aug

Sony India Software Centre

Bengaluru

Lead Platform Engineer (Observability Focus)

About The Role We are looking for a Senior Platform Engineer with deep expertise in observability, cloud-native infrastructure, and large-scale distributed systems. This role is highly hands-on and focuses on designing, building, and operating reliable, observable, and scalable platforms running on Kubernetes, with a strong preference for AWS, and having GCP knowledge would be an edge.

Key Responsibilities Reliability & Operations

- Design, implement, and maintain highly available and resilient systems in Kubernetes-based environments
- Define and enforce SLOs, SLIs, and error budgets
- Lead incident response, RCA, and postmortems
- Drive reliability improvements through automation

Observability (Core Focus)

- Architect and operate observability platforms for metrics, logging, tracing, and alerting
- Work with Prometheus, Alertmanager, Grafana, Splunk, Cribl, Datadog
- Establish actionable alerting standards

Cloud & Platform Engineering

- Build and manage infrastructure on AWS.
- Operate Kubernetes clusters (EKS preferred)
- Deploy services using Helm, ArgoCD and Argo rollout
- Manage containerized workloads using Docker and containerd

Automation & Tooling

- Robust Python skills with emphasis on reliability, automation, and observability tooling
- Develop automation and tooling using Python
- Create internal reliability and monitoring tools
- Integrate CI/CD pipelines with observability and reliability checks

Collaboration & Leadership

- Mentor junior engineers
- Influence architecture decisions
- Collaborate across engineering teams

Required Qualifications

- 6+ years of relevant experience in SRE, DevOps, or Platform Engineering
- Strong Python skills with experience building production-grade automation and tooling
- Strong programming experience in Python
- Production experience with Kubernetes
- Strong observability fundamentals
- Experience with Helm, ArgoCD,



Argo Rollout and Docker
- Experience with AWS cloud
- Strong Linux and networking fundamentals
- Familiarity with the SDLC

Preferred Qualifications

- Experience with OpenTelemetry and Observability tools
- Experience with Kubernetes package manager (helm) and deployment (ArgoCD / Argo Rollout)
- Multi-cluster or multi-region Kubernetes experience
- Service mesh (Istio) and API Gateway (Kong) experience
- Infrastructure-as-Code (Terraform preferred)
- Cloud cost optimization experience

Technology Stack Programming & Automation: Python (strong proficiency, production-grade tooling and automation) Containerization & Orchestration: Docker, AWS EKS

Packaging & Deployment: Helm, ArgoCD, Argo Rollout

Observability & Monitoring: Observability & Monitoring: Prometheus, Alertmanager, Grafana, OpenTelemetry, Datadog, Splunk, Cribl, Edge Collectors, AWS Cloud Watch

Platforms: AWS (primary and preferred), Google Cloud Platform (good to have)

CI/CD & DevOps: Git-based CI/CD pipelines, release automation, reliability checks

Infrastructure as Code: Terraform (preferred)

Operating Systems & Networking: Linux, TCP/IP, DNS, load balancing

Project Details / What You’ll Work On Build and operate a centralized observability platform for metrics, logs, traces, and alerting across Kubernetes workloads using above mentioned tooling for services running in AWS, on-prem, and GCP (good to have)

Contribute to the o11y center-of-excellence guiding teams to create their SLOs, SLIs,



reduce MTTR and collaborate with developers to move toward and Observability Driven Development mindset. Support observability for services running on our Kubernetes ecosystem.

Develop Python-based automation and tooling for observability, SLO reporting, incident response, and operational workflows

Lead incident response for production issues, conduct blameless postmortems, and drive long-term reliability improvements

Optimize platform scalability, performance, and cloud cost efficiency with a strong focus on GCP and AWS.

Act as a technical leader, influencing architecture and mentoring teams on reliability and observability best practices

AWS, Sre, Elk

Reliability & Operations

- Design, implement, and maintain highly available and resilient systems in Kubernetes-based environments
- Define and enforce SLOs, SLIs, and error budgets
- Lead incident response, RCA, and postmortems
- Drive reliability improvements through automation

Observability (Core Focus)
- Architect and operate observability platforms for metrics, logging, tracing, and alerting
- Work with Prometheus, Alertmanager, OpenTelemetry, Grafana, Loki / ELK / OpenSearch
- Implement cloud-native monitoring (GCP Cloud Monitoring & Logging preferred)
- Establish actionable alerting standards

Cloud & Platform Engineering
- Build and manage infrastructure on GCP (preferred) or AWS
- Operate Kubernetes clusters (GKE preferred)
- Deploy services using Helm
- Manage containerized workloads using Docker

Automation & Tooling
- Robust Python skills with emphasis on reliability, automation, and observability tooling
- Develop automation and tooling using Python
- Create internal reliability and monitoring tools
- Integrate CI/CD pipelines with observability and reliability checks

Collaboration & Leadership
- Mentor junior engineers
- Influence architecture decisions
- Collaborate across engineering teams

📌 Lead Platform Engineer - Observability (Bengaluru)
🏢 Sony India Software Centre
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead platform engineer - observability (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: lead platform engineer - observability (bengaluru) / bengaluru