I
Required Skills & Experience
Technical Skills
Strong experience with Microsoft Azure cloud services.
Experience managing enterprise infrastructure and production environments.
Expertise in incident management, problem management, and root cause analysis.
Hands-on experience with monitoring and observability tools such as Splunk, Dynatrace, Datadog, Azure Monitor, AppDynamics, or similar.
Experience with disaster recovery, backup, business continuity, and resilience planning.
Strong understanding of DevOps and Site Reliability Engineering practices.
Knowledge of infrastructure automation and scripting using PowerShell, Python, Terraform, Ansible, or similar tools.
Experience with vulnerability management and security best practices.
Leadership Skills
Experience leading infrastructure, cloud operations, or SRE teams.
Robust stakeholder management and communication skills.
Ability to manage multiple priorities in a fast-paced workplace.
Experience working with globally distributed teams.
Infrastructure & Operations Management
Own the engineering, availability, resiliency, performance, monitoring, and capacity planning of Audit portfolio applications.
Ensure 99.0% or higher availability for critical business applications through proactive monitoring and operational excellence.
Manage and support cloud and on-premises infrastructure environments.
Drive operational improvements to enhance reliability, scalability, and performance.
Site Reliability & Production Support
Proactively monitor production environments to identify and prevent failures, performance issues, and capacity constraints.
Lead incident response activities, problem management, root cause analysis (RCA), and service restoration efforts.
Participate in and oversee a 24x7x365 follow-the-sun support model.
Develop and implement monitoring, alerting, and observability strategies.
Cloud & DevOps
Apply DevOps and SRE best practices to streamline infrastructure provisioning, operations, automation, and change management.
Collaborate with Cloud Platform Engineering teams to maintain secure and scalable cloud settings.
Support application deployments, infrastructure upgrades, and lifecycle management activities.
Drive infrastructure automation using scripting and tooling
📌 Infra & Cloud Engineer Hyderabad (India)
🏢 Xoriant
📍 India