06 Sep
|
Cloudxtreme
|
Pune
Who are we looking for?
Application Support Engineer with overall experience of 4+ years of experience in development and supporting Complex and critical large scale distributed systems and extensive hands-on experience in handling production failures & driving root cause analysis and remediation.
Primary Responsibilities:
Proactively manage production outages and performance issues through quality analysis and rapid resolution.
Lead incident management, ensuring clear communication with users, application owners, and senior stakeholders.
Perform root cause analysis (RCA) and document post-mortems for incidents.
Identify recurring patterns, conduct thorough post-mortems, and implement permanent fixes to prevent reoccurrence.
Collaborate with engineering teams to automate ing and recovery processes, reducing manual intervention.
Drive continuous improvement by minimizing manual tasks through automation and dynamic monitoring solutions.
Build and enhance runbooks to streamline operational workflows.
Work with development teams to improve application observability and enable faster MTTD (Mean Time to Detect) and MTTR (Mean Time to Resolve).
Identify automation prospects in s and processes, and partner with engineering to implement them.
Possess solid understanding of deployment methodologies with hands-on experience in production deployments.
Skilled in instrumentation, monitoring, ing, and incident response using tools like AppDynamics, Splunk, ThousandEyes, and ITRS.
Technical Skills:
Over 4+ years of IT experience, including 3+ years in L2 application support.
Proficient in managing large-scale production systems involving load balancing, distributed systems, microservices, monitoring, and configuration management.
Robust hands-on experience in troubleshooting:
Application failures and performance degradation
Code-level issues and cloud platform incidents
Batch, infrastructure, database, and network failures
Skilled in ITIL practices: Event, Incident, Release, Problem, and Knowledge Management.
Experienced in production deployments using CI/CD tools.
Solid understanding of SLIs, SLOs, and handling burn rate s.
Working knowledge of SQL and NoSQL databases.
Proficient in monitoring and ing tools like AppDynamics, Splunk, ThousandEyes, and ITRS.
Expertise in cloud technologies, with a preference for PCF (Pivotal Cloud Foundry).
Expertise in Scripting technologies like Shell, Powershell, Python.
📌 Production Support / Application Support Pune
🏢 Cloudxtreme
📍 Pune