Lead I - DevOps Engineering (Bengaluru)

Lead I - DevOps Engineering (Bengaluru)

08 Sep
|
UST
|
Bengaluru

08 Sep

UST

Bengaluru

Details

- Locations: Bengaluru
- Minimum Experience: 5
- Maximum Experience: 8
- Mandatory Skills: Observability, Telemetry, Cloud, Containers, Infrastructure as Code, Automation, CI/CD

Job Summary

Job Title: Platform Engineer

As a Platform Engineer, you will build, operate, and continuously improve enterprise monitoring and observability platforms that enable reliable service delivery. You will design scalable telemetry pipelines for metrics, logs, and traces; improve data quality and effectiveness; and develop dashboards that provide clear insight into service health and performance. Working closely with application, infrastructure, security, compliance, and incident-response teams, you will automate platform operations, establish monitoring standards, and improve detection-to-diagnosis workflows. Your work will help reduce noise, accelerate incident recovery, and promote consistent adoption of observability practices across cloud and on-premises environments.

Responsibilities

- Engineer and operate enterprise monitoring and observability platforms such as Dynatrace, Datadog, New Relic, Splunk, Elastic, Prometheus, and Grafana.
- Design and maintain telemetry collection using agents, exporters, collectors, and API-based integrations across applications, servers, networks, databases, containers, storage, and cloud services.
- Manage the performance, capacity, high availability, security, and lifecycle upgrades of monitoring platforms and supporting infrastructure.
- Implement platform security controls, including role-based access control, single sign-on, encryption, secrets management, and auditability.
- Establish ing standards for severity, thresholds, anomaly detection, suppression, correlation, deduplication, and notification routing.
- Reduce fatigue by tuning signals, removing redundant rules, and creating actionable s with links to dashboards, runbooks, and diagnostic context.
- Build standardized dashboards for service health, application performance, infrastructure capacity, availability, and SLO/SLA tracking.
- Partner with service owners to define and implement service-level indicators, service-level objectives, and error budgets where applicable.
- Report operational trends such as availability, mean time to detect, mean time to restore, recurring incidents, and platform adoption.
- Automate platform configuration, deployment, and service onboarding using infrastructure-as-code, configuration-management, and CI/CD tools.
- Integrate monitoring platforms with Jira,



CMDB systems, Microsoft Teams, Slack, PagerDuty, Opsgenie, and other incident-management tools.
- Develop reusable s, dashboards, templates, policies, and self-service onboarding patterns.
- Define monitoring standards covering minimum telemetry requirements, metadata, tagging, naming conventions, and data governance.
- Maintain technical documentation, operational procedures, troubleshooting guides, and onboarding runbooks.
- Participate in incident investigations and post-incident reviews to improve detection, prevent recurrence, and strengthen monitoring coverage.
- Collaborate with Security and Compliance teams to meet privacy, data-retention, regulatory, and audit requirements.
- Participate in a scheduled 24 7 on-call rotation supporting monitoring infrastructure and services.

Skills and Tools Required

- Three to seven or more years of experience in monitoring and observability, systems or platform engineering, site reliability engineering, or IT operations.
- Hands-on experience with at least one enterprise observability platform, such as Dynatrace, Datadog, New Relic, Splunk, Elastic, Prometheus, or Grafana.
- Strong understanding of metrics, logs, traces, telemetry pipelines, and distributed-system troubleshooting.
- Experience with OpenTelemetry concepts, collectors, and instrumentation best practices.
- Experience administering and troubleshooting Linux and Windows systems in large-scale environments.
- Working knowledge of networking fundamentals, including DNS, TCP/IP, firewalls, load balancers, and common infrastructure components.
- Experience with cloud-monitoring services such as Azure Monitor and Log Analytics, AWS CloudWatch, or Google Cloud Operations.
- Knowledge of container and Kubernetes observability using technologies such as Prometheus, Grafana, Tempo, Loki, and OpenTelemetry.
- Scripting and automation skills using Python, PowerShell, or Bash.
- Experience with infrastructure-as-code and configuration-management tools such as Terraform, Bicep, CloudFormation, or Ansible.
- Familiarity with version control, CI/CD pipelines, automated deployments,



and repeatable build practices.
- Experience integrating observability platforms with ITSM, CMDB, event-management, and incident-response workflows.
- Understanding of SRE principles, including SLIs, SLOs, error budgets, reliability engineering, MTTD, and MTTR.
- Robust ability to distinguish meaningful signals from noise and design s that support effective action.
- Systems-level troubleshooting skills across application, platform, infrastructure, database, and network layers.
- Excellent written and verbal communication skills, including the ability to work with technical and non-technical stakeholders.
- A customer-focused approach to self-service enablement, documentation, standardization, and platform adoption.
- The ability to prioritize effectively, work in a fast-paced operational environment, and handle incident escalation calmly.
- A bachelor s degree in Computer Science, Information Systems, or a related field, or equivalent professional experience.

Qualifications

- Bachelor s Degree in Computer Science or a related field

Preferred Certifications

- Monitoring-platform certification from vendors such as Datadog, Dynatrace, or New Relic.
- Cloud certification for Microsoft Azure, Amazon Web Services, or Google Cloud.
- ITIL Foundation certification.

Join Sony as a Platform Engineer

Join Sony as a Platform Engineer and help establish reliable, scalable, and actionable observability capabilities across the organization. You will work within an inclusive, collaborative, and global technology community while contributing to service reliability, operational efficiency, and continuous improvement.

Education Qualification

Project Details

The Monitoring Platform Engineering team builds and operates scalable, secure, and reliable monitoring solutions for enterprise applications and infrastructure across cloud and on-premises environments. The role involves managing telemetry pipelines, dashboards, s, and platform integrations while automating deployment and service onboarding through IaC and CI/CD. The project aims to reduce noise, improve incident detection and resolution, and strengthen service reliability and operational efficiency.

Shift Timings

9am - 5pm

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Lead I - DevOps Engineering (Bengaluru)
🏢 UST
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead i - devops engineering (bengaluru) / bengaluru