Design, implement, and maintain highly available and resilient systems in Kubernetes-based environments
Define and enforce SLOs, SLIs, and error budgets
Lead incident response, RCA, and postmortems
Drive reliability improvements through automation
Observability (Core Focus)
Architect and operate observability platforms for metrics, logging, tracing, and alerting
Work with Prometheus, Alertmanager, OpenTelemetry, Grafana, Loki / ELK / OpenSearch
Implement cloud-native monitoring (GCP Cloud Monitoring & Logging preferred)
Establish actionable alerting standards
Cloud & Platform Engineering
Build and manage infrastructure on GCP (preferred) or AWS
Operate Kubernetes clusters (GKE preferred)
Deploy services using Helm
Manage containerized workloads using Docker
Automation & Tooling
Solid Python skills with emphasis on reliability, automation, and observability tooling
Develop automation and tooling using Python
Create internal reliability and monitoring tools
Integrate CI/CD pipelines with observability and reliability checks
Collaboration & Leadership
Mentor junior engineers
Influence architecture decisions
Collaborate across engineering teams
📌 Lead Platform Engineer (Bengaluru)
🏢 Sony India Software Centre
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.