Role Description
Job Code 15023 Locations Bengaluru Minimum Experience 5 Maximum Experience 8 Mandatory Skills Observability,Telemetry,Cloud,Containers,Infrastructure as Code,Automation,CI/CD Skill to Evaluate Observability,Telemetry,Cloud,Containers,Infrastructure as Code,Automation,CI/CD Experience 5 to 8 Years Location Bengaluru Job Title: Platform Engineer : As a Platform Engineer, you will build, operate, and continuously improve enterprise monitoring and observability platforms that enable reliable service delivery. You will design scalable telemetry pipelines for metrics, logs, and traces; improve data quality and effectiveness; and develop dashboards that provide clear insight into service health and performance. Working closely with application, infrastructure, security, compliance, and incident-response teams, you will automate platform operations, establish monitoring standards, and improve detection-to-diagnosis workflows.
Your work will help reduce noise, accelerate incident recovery, and promote consistent adoption of observability practices across cloud and on-premises environments. Key Responsibilities:
- Engineer and operate enterprise monitoring and observability platforms such as Dynatrace, Datadog, Current Relic, Splunk, Elastic, Prometheus, and Grafana.
- Design and maintain telemetry collection using agents, exporters, collectors, and API-based integrations across applications, servers, networks, databases, containers, storage, and cloud services.
- Manage the performance, capacity, high availability, security, and lifecycle upgrades of monitoring platforms and supporting infrastructure.
- Implement platform security controls, including role-based access control, single sign-on, encryption, secrets management, and auditability.
- Establish ing standards for severity, thresholds, anomaly detection, suppression, correlation, deduplication, and notification routing.
- Reduce fatigue by tuning signals, removing redundant rules, and creating actionable s with links to dashboards, runbooks, and diagnostic context.
- Build standardized dashboards for service health, application performance, infrastructure capacity, availability, and SLO/SLA tracking.
- Partner with service owners to define and implement service-level indicators, service-level objectives, and error budgets where applicable.
- Report operational trends such as availability, mean time to detect, mean time to restore, recurring incidents, and platform adoption.
- Automate platform configuration, deployment, and service onboarding using infrastructure-as-code, configuration-management, and CI/CD tools.
- Integrate monitoring platforms with Jira, CMDB systems, Microsoft Teams, Slack, PagerDuty, Opsgenie, and other incident-management tools.
- Develop reusable s, dashboards, templates, policies, and self-service onboarding patterns.
- Define monitoring standards covering minimum telemetry requirements, metadata, tagging, naming conventions, and data governance.
- Maintain technical documentation, operational procedures, troubleshooting guides, and onboarding runbooks.
- Participate in incident investigations and post-incident reviews to improve detection, prevent recurrence, and strengthen monitoring coverage.
- Collaborate with Security and Compliance teams to meet privacy, data-retention, regulatory, and audit requirements.
- Participate in a scheduled 24×7 on-call rotation supporting monitoring infrastructure and services. Skills and Tools Required:
- Three to seven or more years of experience in monitoring and observability, systems or platform engineering, site reliability engineering, or IT operations.
- Hands-on experience with at least one enterprise observability platform, such as Dynatrace, Datadog, New Relic, Splunk, Elastic, Prometheus, or Grafana.
- Strong understanding of metrics, logs, traces, telemetry pipelines, and distributed-system troubleshooting.
- Experience with OpenTelemetry concepts, collectors, and instrumentation best practices.
- Experience administering and troubleshooting Linux and Windows systems in large-scale environments.
- Working knowledge of networking fundamentals, including DNS, TCP/IP, firewalls, load balancers, and common infrastructure components.
- Experience with cloud-monitoring services such as Azure Monitor and Log Analytics, AWS CloudWatch, or Google Cloud Operations.
- Knowledge of container and Kubernetes observability using technologies such as Prometheus, Grafana, Tempo, Loki, and OpenTelemetry.
- Scripting and automation skills using Python, PowerShell, or Bash.
- Experience with infrastructure-as-code and configuration-management tools such as Terraform, Bicep, CloudFormation, or Ansible.
- Familiarity with version control, CI/CD pipelines, automated deployments, and repeatable build practices.
- Experience integrating observability platforms with ITSM, CMDB, event-management, and incident-response workflows.
- Understanding of SRE principles, including SLIs, SLOs, error budgets, reliability engineering, MTTD, and MTTR.
- Strong ability to distinguish meaningful signals from noise and design s that support effective action.
- Systems-level troubleshooting skills across application, platform, infrastructure, database, and network layers.
- Excellent written and verbal communication skills, including the ability to work with technical and non-technical stakeholders.
- A customer-focused approach to self-service enablement, documentation, standardization, and platform adoption.
- The ability to prioritize effectively, work in a fast-paced operational environment, and handle incident escalation calmly.
- A bachelor’s degree in Computer Science, Information Systems, or a related field, or equivalent professional experience.
Preferred
Certifications:
- Monitoring-platform certification from vendors such as Datadog, Dynatrace, or New Relic.
- Cloud certification for Microsoft Azure, Amazon Web Services, or Google Cloud.
- ITIL Foundation certification. Join Sony as a Platform Engineer and help establish reliable, scalable, and actionable observability capabilities across the organization. You will work within an inclusive, collaborative, and global technology community while contributing to service reliability, operational efficiency, and continuous improvement.
Education Qualificaiton Bachelor’s Degree in Computer Science or a related field Project Details The Monitoring Platform Engineering team builds and operates scalable, secure, and reliable monitoring solutions for enterprise applications and infrastructure across cloud and on-premises environments. The role involves managing telemetry pipelines, dashboards, s, and platform integrations while automating deployment and service onboarding through IaC and CI/CD. The project aims to reduce noise, improve incident detection and resolution, and strengthen service reliability and operational efficiency.
Shift
Timings 9am - 5PM
Skills Observability, OpenTelemetry, Cloud Infrastructure, CI/CD
📌 Lead I - DevOps Engineering (Bengaluru)
🏢 UST
📍 Bengaluru