05 Sep
|
Razorpay
|
Bengaluru
05 Sep
Razorpay
Bengaluru
Job Description
n
Razorpay — Incident Engineer | Bangalore (In-Office)
n
Experience: 1-3 Years
n
An Incident Engineer @ Razorpay is well-grounded — smart, quality focussed, and a product thinker. Engineering creates a significant impact across different areas considering the scale of our software product outreach. You’re also expected to influence the culture of the company and help shape it the right way.
n
We are looking for an Incident Engineer to join our Incident Management team. This role is primarily focused on Problem Management — owning the RCA lifecycle, managing post-incident reviews, and driving continuous improvement. You will also handle developer-facing outages and, as the process demands, provide rotational support for production incidents impacting external customers. This is a rotational role with 24/7 coverage as needed.
n
Responsibilities
n
Monitoring & Proactive Detection
n
n
- Responsible for monitoring all major metrics via various monitoring tools and following the major incident management process in restoring major impacting incidents.
n
- Respond to reported service incidents, identifying the cause, and initiating the incident management process.
n
- Proactively identify high-impact scenarios from monitoring tools and engage stakeholders to avert any potential impacts.
n
- Prioritize incidents according to their urgency and impact on the business.
n
- Execute strategies for proactive monitoring and move the support from reactive to proactive.
n
- Effective measurement strategies for impacts averted or avoided (P-1 Avoids).
n
n
Incident Resolution & Facilitation
n
n
- Triage, facilitate, and drive all major issues to resolution.
n
- Follow the escalation matrix and engage stakeholders over bridge calls; in parallel, send incident communications as per defined timelines.
n
- Serve as the point of contact for all incidents and ensure effective implementation of the Incident Management process.
n
- Represent the first stage of escalation for incidents.
n
- Manage developer-facing outages (CI/CD, DevStack, build pipelines, internal tooling) and coordinate with platform teams for quick resolution.
n
- Provide rotational production incident response for external customer-impacting incidents, including 24/7 shift coverage when assigned, based on process demands.
n
n
Problem Management, Documentation & Knowledge Management
n
n
- Own the end-to-end RCA lifecycle — assign, track, review, and close post-incident RCAs within defined timelines.
n
- Schedule and conduct RCA review meetings with engineering teams; ensure action items are validated and closed.
n
- Prepare weekly problem management reports summarizing RCA status, recurring patterns, and improvement trends.
n
- Conduct Post-Incident Management Reviews, manage Problem Management, track key metrics, and drive RCA assignment & documentation on identified action items.
n
- Produce and maintain documents outlining incident protocols (e.g., handling cybersecurity threats, correcting server failures).
n
- Monitor incidents to ensure Service Level Agreements are met and respected.
n
- Ensure closure of all resolved and end-user confirmed incident records.
n
- Conduct brown bag sessions on Incident Management Process and educate/train stakeholders.
n
- Provide guidance to the Incident Process Coordinators.
n
n
AI & Automation Initiatives
n
n
- Leverage AI and automation to improve incident and problem management.
n
- Propose and execute initiatives such as automated RCA summaries, intelligent alert correlation, predictive detection, and AI-assisted reporting.
n
- Identify automation opportunities to enhance incident and problem management workflows.
n
n
Process Improvement
n
n
- Execute continuous process improvement initiatives where process performance, activities, roles and responsibilities, policies, procedures, and supporting technology are reviewed and enhanced where applicable.
n
- Drive continuous improvement — move from reactive to proactive monitoring, and measure impacts averted (P-1 Avoids).
n
n
Stakeholder Communication & Reporting
n
n
- Communicate with stakeholders and leadership teams for major issues with timely updates during the lifecycle of the incident.
n
- Plan and coordinate all activities required to perform, monitor, and report on the process.
n
- Translate findings into business reports and presentations.
n
n
Requirements
n
n
- Minimum of 1–2 years of overall experience in the IT industry, with at least 1 year as Incident Manager or Senior Incident Engineer.
n
- ITIL Foundation certification (mandatory); Expert Level preferred.
n
- Experience working with Enterprise Command Center or NOC teams.
n
- Solid verbal and written communication skills.
n
- Proficiency in data management tools.
n
- Strong data/information literacy skills; ability to perform root cause analysis and problem solving.
n
- Technical and analytical skills with the ability to translate findings into business reports and presentations.
n
- High degree of self-motivation and a can-do demeanor.
n
- Open to working shifts including 24/7 rotational coverage and on-call support as necessary.
n
- Eager to adopt AI tools and identify automation opportunities to enhance incident and problem management workflows.
n
n
📌 Incident Engineer (Bengaluru)
🏢 Razorpay
📍 Bengaluru