16 Sep
|
MetLife
|
Hyderabad
Hi,
Greetings!
Hope you are doing well!
Please have a look at the below and let us know your interest.
Site Reliability Engineer
Location- Hyderabad
Mode- Hybrid
Interview mode- Face to Face
Date- 26th September 2026
Role Overview A Site Reliability Engineer (SRE) will be responsible for ensuring highly available, scalable, performant, and reliable data platforms, pipelines, and services. The role will drive operational excellence, strengthen observability, improve incident response, reduce toil through automation, and collaborate with engineering, data, cloud, and operations teams to deliver resilient services and measurable business outcomes.
Key Responsibilities
- Ensure availability, performance, scalability, and reliability of data platforms, pipelines, and services through proactive monitoring, troubleshooting, and timely issue resolution.
- Design, implement, and operate large-scale data systems in partnership with engineering teams, using automation and tooling to streamline operations.
- Develop scripts, utilities, and reusable automation to reduce manual effort, minimize errors, improve efficiency, and support repeatable platform operations.
- Build and continuously improve monitoring, alerting, dashboards, and observability signals for early detection, faster response, and improved platform health.
- Participate in incident response, service restoration, root cause analysis, postmortems, and corrective actions to strengthen resilience and prevent recurring issues.
- Maintain transparent documentation, runbooks, processes, and operational procedures while promoting knowledge sharing and alignment with governance,
controls, and production support standards.
Candidate Qualifications
- 4+ years of experience in production support, DevOps, infrastructure, cloud operations, data platform operations, or software engineering.
- Experience supporting business-critical systems and working within incident, problem, and change management processes.
- Strong scripting and automation skills using Python, PowerShell, Bash, Spark, or equivalent technologies.
- Bachelor's degree in computer science, engineering, or equivalent practical experience.
- Exposure to hybrid cloud platforms, Azure-hosted services, regulated enterprise environments, insurance, banking, or financial services is preferred.
- Business proficiency in English and Japanese language skills is preferred.
Tech Stack
- Programming & Automation: Python, Spark, Bash, and PowerShell for automation, diagnostics, data processing, and operational tooling.
- Azure Data & Cloud Platforms: Azure Data Lake, Data Factory, Synapse Analytics, Azure SQL, Cosmos DB, Databricks, and hybrid production support environments.
- Observability & Reliability: Azure Monitor, Application Insights, Log Analytics, Splunk, AppDynamics, ELK, SLIs, SLOs, SLAs, incident management, RCA, and toil reduction.
- DevOps & Cloud-Native Operations: Azure DevOps, GitHub, Docker, Kubernetes, AKS, CI/CD, deployments, scalability, and platform reliability.
- Operations, Resilience & Collaboration: ServiceNow, ITSM, runbooks, disaster recovery, multi-region architecture, AI-assisted engineering tools such as GitHub Copilot or M365 Copilot, and cross-functional collaboration.
📌 SRE (Hyderabad)
🏢 MetLife
📍 Hyderabad