Key Skills: itsm, Linux, Observability, Azure, Troubleshooting, Kubernetes Roles and Responsibilities: Analyze and resolve complex application and infrastructure incidents escalated from monitoring and first-line support teams. Manage day-to-day IT operations activities including incidents, problems, changes, vulnerabilities, and release coordination. Support server maintenance activities and coordinate infrastructure upgrades and patching with internal teams and providers. Participate in Major Incident Management activities and drive technical resolution during critical outages. Operate and maintain monitoring and observability tooling to improve service visibility and operational stability. Skills Required: ~7 to 8 years of relevant experience in application support or IT operations.
~ Strong hands-on experience in Microsoft Azure, Linux systems, and application support in a production environment. ~ Working knowledge of Azure Kubernetes Service, infrastructure operations, and enterprise monitoring platforms. ~ Ability to troubleshoot incidents using observability tools, logs, dashboards, and SQL-based analysis. ~ Experience working in ITIL-based support processes with SLA/KPI-driven operations and ticket management in JIRA. Positive to Have: ~ Exposure to Kubernetes in containerized environments. Education: Any graduation or post-graduation in IT or related discipline.