27 Sep
|
Cloudxtreme
|
Pune
Who are we looking for?
Application Support Engineer with overall experience of 4+ years of experience in development and supporting Complex and critical large scale distributed systems and extensive hands-on experience in handling production failures & driving root cause analysis and remediation.
Primary Responsibilities:
- Proactively manage production outages and performance issues through quality analysis and rapid resolution.
- Lead incident management, ensuring clear communication with users, application owners, and senior stakeholders.
- Perform root cause analysis (RCA) and document post-mortems for incidents.
- Identify recurring patterns, conduct thorough post-mortems, and implement permanent fixes to prevent reoccurrence.
- Collaborate with engineering teams to automate ing and recovery processes, reducing manual intervention.
- Drive continuous improvement by minimizing manual tasks through automation and dynamic monitoring solutions.
- Build and enhance runbooks to streamline operational workflows.
- Work with development teams to improve application observability and enable faster MTTD (Mean Time to Detect) and MTTR (Mean Time to Resolve).
- Identify automation opportunities in s and processes, and partner with engineering to implement them.
- Possess strong understanding of deployment methodologies with hands-on experience in production deployments.
- Skilled in instrumentation, monitoring, ing, and incident response using tools like AppDynamics, Splunk, ThousandEyes, and ITRS.
Technical Skills:
- Over 4+ years of IT experience, including 3+ years in L2 application support.
- Proficient in managing large-scale production systems involving load balancing, distributed systems, microservices, monitoring, and configuration management.
- Robust hands-on experience in troubleshooting:
- Application failures and performance degradation
- Code-level issues and cloud platform incidents
- Batch, infrastructure, database, and network failures
- Skilled in ITIL practices: Event, Incident, Release, Problem, and Knowledge Management.
- Experienced in production deployments using CI/CD tools.
- Solid understanding of SLIs, SLOs, and handling burn rate s.
- Working knowledge of SQL and NoSQL databases.
- Proficient in monitoring and ing tools like AppDynamics, Splunk, ThousandEyes, and ITRS.
- Expertise in cloud technologies, with a preference for PCF (Pivotal Cloud Foundry).
- Expertise in Scripting technologies like Shell, Powershell, Python.
📌 Production Support / Application Support (Pune)
🏢 Cloudxtreme
📍 Pune