Tower Lead - Microsoft Azure - IaaS, Azure Kubernetes Service
Noida, Uttar Pradesh
Job Summary
Cloud SME
Key Responsibilities
Advanced Incident Management & Escalation (Tier 3)High-Severity Resolution: Act as the final internal escalation point for critical (Severity-1 and Severity-2) production incidents across the Azure platform [L3 Resource Overview].Deep Diagnostic Triage: Troubleshoot deeply complex platform anomalies using Kusto Query Language (KQL), Azure Monitor, Log Analytics, and Application Insights [L3 Resource Overview].Vendor Collaboration: Serve as the technical lead when collaborating with Microsoft Premier/Unified Support teams, ensuring swift resolution of platform-level bugs [Operational Skill Set & Prerequisites].2. Deep Platform Troubleshooting & RemediationKubernetes & Container Operations: Diagnose complex cluster issues within Azure Kubernetes Service (AKS), including container network interface (CNI) IP exhaustion, pod eviction loops, and core DNS failures [Platform & Core Services Mastery].Advanced Networking Resolution: Debug hybrid networking failures across ExpressRoute circuits, Azure Virtual WAN, VPN gateways, and complex Application Gateway / WAF configurations [Packet Analysis, Platform & Core Services Mastery].Database & PaaS Optimization: Investigate and remediate transaction deadlocks, connection pooling limits, and performance degradation in Azure SQL Managed Instance, Cosmos DB, and Azure Cache for Redis [Platform & Core Services Mastery].3. Proactive Engineering & Root Cause Analysis (RCA)Blameless RCAs:
Conduct exhaustive post-incident investigations, authoring high-quality Root Cause Analysis (RCA) documents to ensure identical failures never reoccur [Root Cause Analysis (RCA) & Resilience].Infrastructure Patching: Implement emergency configuration patches by safely updating modular Infrastructure as Code (IaC) configurations using Terraform or Azure Bicep [Root Cause Analysis (RCA) & Resilience].Disaster Recovery Executions: Lead operational readiness testing and manage high-pressure, live database and VM failovers using Azure Site Recovery (ASR) during regional outages [Root Cause Analysis (RCA) & Resilience].
Skill Requirements
Core Platform Mastery: Deep, hands-on administrative command over Azure Virtual Machines, Virtual Machine Scale Sets (VMSS), Azure Storage Accounts, and Azure App Services [PaaS & Data Services, Up-to-date App Hosting, Platform & Core Services Mastery].Advanced Networking Diagnostics: Proficient use of Azure Network Watcher, packet capture tools, Wireshark, and traffic analytics to trace distributed routing packet drops [Packet Analysis].Identity & Governance Triage: Experience diagnosing access failures, service principal expiration issues, and permission blocks within Microsoft Entra ID (Azure AD) and Azure Policy [Identity & Governance].Scripting & Automation: Strong ability to read, write, and execute troubleshooting scripts using PowerShell, Bash, or Python to automate emergency triage tasks [Automation].
Other Requirements
NA
📌 Tower Lead - Microsoft Azure - IaaS, Azure Kubernetes Service (Noida)
🏢 HCLTech
📍 Noida