Incident Engineer (Mumbai)

Incident Engineer (Mumbai)

03 Sep
|
Razorpay
|
Mumbai

03 Sep

Razorpay

Mumbai

Razorpay —

Incident Engineer

Bangalore (In-Office)Experience: 1-3 YearsAn Incident Engineer @ Razorpay is well-grounded — smart, quality focussed, and a product thinker. Engineering creates a significant impact across different areas considering the scale of our software product outreach. You’re also expected to influence the culture of the company and help shape it the right way.We are looking for an Incident Engineer to join our Incident Management team. This role is primarily focused on Problem Management — owning the RCA lifecycle, managing post-incident reviews, and driving continuous improvement. You will also handle developer-facing outages and, as the process demands, provide rotational support for production incidents impacting external customers. This is a rotational role with 24/7 coverage as needed.ResponsibilitiesMonitoring & Proactive DetectionResponsible for monitoring all major metrics via various monitoring tools and following the major incident management process in restoring major impacting incidents.Respond to reported service incidents, identifying the cause, and initiating the incident management process.Proactively identify high-impact scenarios from monitoring tools and engage stakeholders to avert any potential impacts.Prioritize incidents according to their urgency and impact on the business.Execute strategies for proactive monitoring and move the support from reactive to proactive.Effective measurement strategies for impacts averted or avoided (P-1 Avoids).Incident Resolution & FacilitationTriage, facilitate, and drive all major issues to resolution.Follow the escalation matrix and engage stakeholders over bridge calls; in parallel,



send incident communications as per defined timelines.Serve as the point of contact for all incidents and ensure effective implementation of the Incident Management process.Represent the first stage of escalation for incidents.Manage developer-facing outages (CI/CD, DevStack, build pipelines, internal tooling) and coordinate with platform teams for quick resolution.Provide rotational production incident response for external customer-impacting incidents, including 24/7 shift coverage when assigned, based on process demands.Problem Management, Documentation & Knowledge ManagementOwn the end-to-end RCA lifecycle — assign, track, review, and close post-incident RCAs within defined timelines.Schedule and conduct RCA review meetings with engineering teams; ensure action items are validated and closed.Prepare weekly problem management reports summarizing RCA status, recurring patterns, and improvement trends.Conduct Post-Incident Management Reviews, manage Problem Management, track key metrics, and drive RCA assignment & documentation on identified action items.Produce and maintain documents outlining incident protocols (e.g., handling cybersecurity threats, correcting server failures).Monitor incidents to ensure Service Level Agreements are met and respected.Ensure closure of all resolved and end-user confirmed incident records.Conduct brown bag sessions on Incident Management Process and educate/train stakeholders.Provide guidance to the Incident Process Coordinators.AI & Automation InitiativesLeverage AI and automation to improve incident and problem management.Propose and execute initiatives such as automated RCA summaries, intelligent alert correlation, predictive detection, and AI-assisted reporting.Identify automation opportunities to enhance incident and problem management workflows.Process ImprovementExecute continuous process improvement initiatives where process performance, activities, roles and responsibilities, policies, procedures, and supporting technology are reviewed and enhanced where applicable.Drive continuous improvement — move from reactive to proactive monitoring, and measure impacts averted (P-1 Avoids).Stakeholder Communication & ReportingCommunicate with stakeholders and leadership teams for major issues with timely updates during the lifecycle of the incident.Plan and coordinate all activities required to perform, monitor, and report on the process.Translate findings into business reports and presentations.RequirementsMinimum of 1–2 years of overall experience in the IT industry, with at least 1 year as Incident Manager or Senior Incident Engineer.ITIL Foundation certification (mandatory); Expert Level preferred.Experience working with Enterprise Command Center or NOC teams.Strong verbal and written communication skills.Proficiency in data management tools.Solid data/information literacy skills; ability to perform root cause analysis and problem solving.Technical and analytical skills with the ability to translate findings into business reports and presentations.High degree of self-motivation and a can-do demeanor.Open to working shifts including 24/7 rotational coverage and on-call support as necessary.Eager to adopt AI tools and identify automation opportunities to enhance incident and problem management workflows.

📌 Incident Engineer (Mumbai)
🏢 Razorpay
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: incident engineer (mumbai) / mumbai

Subscribe to this job alert:

Get the latest job offers by email for: incident engineer (mumbai) / mumbai