Site Reliability Technical Operations Manager (Pune)

Site Reliability Technical Operations Manager (Pune)

02 Oct
|
Johnson Controls
|
Pune

02 Oct

Johnson Controls

Pune

Summary:The Site Reliability Engineering team at Johnson Controls is seeking a Technical Reliability & Support Manager to lead cloud product reliability, operational support, and production stability across global cloud applications and platforms. This role will be responsible for managing day-to-day technical operations, L2/L3 support coordination, incident response, service reliability governance, and continuous improvement for cloud products hosted across platforms such as Azure (primarily), Google Cloud, Ali Cloud, and other cloud environments.The role will partner closely with Engineering, SRE, Security, Observability, and external support partners to ensure production issues are resolved quickly, recurring problems are eliminated, support processes are standardized, and operational risks are proactively identified and addressed.Primary Duties:Lead technical operations and support management for cloud products across global environmentsManage day-to-day L2/L3 support activities, incident response, escalations, and production issue resolutionOwn service reliability governance for assigned cloud products, including availability, incident trends, MTTR, recurring issues, and operational risksPartner with Engineering, SRE, Security, Observability and Platform teams to improve service stability and operational readinessDrive incident management, problem management, RCA/PCA reviews, and corrective action tracking for production issuesEnsure timely communication during incidents, including stakeholder updates, executive summaries, customer-impact statements, and resolution updatesEstablish and manage support processes aligned with ITIL practices, including incident, problem, change, service request, and escalation managementDefine, track, and report operational KPIs such as availability, MTTR, incident volume,



severity trends, backlog, service requests, alert noise, and SLA performanceIdentify recurring operational pain points and drive permanent fixes through engineering backlog, automation, monitoring improvements, and process enhancementsCollaborate with Observability teams to ensure critical application, infrastructure, database, and integration components are properly monitored and alertedSupport implementation and adoption of SLIs, SLOs, SLAs, error budgets, and reliability / supportability scorecards for cloud productsEnsure production readiness for new product releases, migrations, infrastructure changes, and platform transformationsLead operational reviews with internal teams, vendors, and external support partners to ensure accountability and continuous improvementDrive automation of manual support tasks, ticket workflows, reporting, and operational runbooks to improve efficiency and consistencyEnsure support activities, incidents, changes, and action items are properly documented and tracked in tools such as Jira, ServiceNow, Confluence, or equivalent platformsManage support handoffs across global teams and ensure transparent ownership, escalation paths, and communication protocolsProvide technical leadership during high-severity incidents, major outages, migrations, and critical customer-impacting eventsEnsure compliance with security, audit, operational,



and IT governance standards for cloud product supportMaintain operational documentation, SOPs, runbooks, escalation matrices, support models, and knowledge base articlesMentor support engineers and technical teams on reliability practices, incident handling, RCA quality, and operational excellenceQualifications:10+ years of experience in technical operations, production support, SRE, cloud support, or reliability engineering.Strong experience managing cloud-hosted applications and support operations.Good knowledge of Azure (preferred), with exposure to AWS, GCP, or Ali Cloud.Experience with incident management, problem management, change management, and operational governance.Familiarity with cloud-native technologies, microservices, APIs, databases, containers, and Kubernetes.Understanding of observability tools such as Grafana, Datadog, ELK, Logz.Io, Azure Monitor, or similar platforms.Knowledge of reliability practices including SLIs, SLOs, SLAs, monitoring, alerting, and automation.Strong troubleshooting skills across applications, infrastructure, networking, databases, and cloud services.Experience with Jira, Confluence, or similar platforms.Excellent leadership, communication, stakeholder management, and vendor coordination skills.Ability to lead high-severity incidents and drive operational excellence in a fast-paced environment.Mandatory Skills:Cloud Operations / SRE leadershipL2 production support (good to have L2 Production Support)Incident & escalation managementRCA/PCA and problem managementAzure cloud and cloud-native platformsObservability, monitoring & KPI reportingITIL-based support processesAutomation and runbook developmentExecutive communication & stakeholder managementLeadership in high-pressure production environments

📌 Site Reliability Technical Operations Manager (Pune)
🏢 Johnson Controls
📍 Pune

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability technical operations manager (pune) / pune