05 Aug
|
HCL Technologies
|
Pune
05 Aug
HCL Technologies
Pune
Senior Site Reliability Engineer Lead
Experience: Not Available to Not Available years
Location: Pune, India
Skills: MQ, NATS/Event Broker, Python, Bash, Java, Linux, monitoring tools, logging tools, alerting tools, distributed systems, messaging platforms, automation, reliability engineering, SRE best practices
Job Summary
Overview
The Reactive Systems Architecture (RSA) organization is responsible for operating Mastercard’s enterprise messaging and event‑driven platforms, including MQ, NATS/Event Broker (EB). These platforms are foundational to Mastercard’s critical business flows and support high‑volume, low‑latency, and highly regulated workloads. We are seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, scalability, and operational excellence of these platforms. This role blends production operations, distributed systems expertise, automation, and incident leadership, with responsibilities that scale based on job level.
Role
As a Site Reliability Engineer supporting MQ, NATS/Event Broker, you will be responsible for the stability and resilience of Mastercard’s messaging backbone. You will partner closely with application teams, platform engineering, infrastructure, and security teams to reduce operational risk, improve system reliability, and ensure issues are detected and resolved before customer impact. This is a production‑focused engineering role, not an application development role.
Key Responsibilities
Ensure high availability, performance, and resilience of MQ, NATS/Event Broker platforms across environments.
Participate in on‑call rotations and provide hands‑on support during production incidents.
Lead or contribute to incident triage, mitigation, and service restoration.
Perform root cause analysis (RCA) and drive corrective and preventive actions to closure.
Design, implement, and maintain monitoring, alerting, and dashboards to enable proactive detection.
Support and govern production changes, including upgrades, patching, certificate renewals,
and configuration changes.
Assess operational readiness for changes and ensure rollback and validation plans are in place.
Automate operational tasks and workflows to reduce manual effort and improve recovery times.
Partner with application teams to support onboarding, scaling, and operational best practices.
Create and maintain runbooks, SOPs, and operational documentation.
Contribute to continuous improvement of reliability, observability, and operational processes.
Skill Requirements
Required Skills and Experience
Experience supporting mission‑critical production systems with on‑call responsibility.
Solid understanding of distributed systems and messaging platforms.
Hands‑on experience with MQ, NATS/Event Broker, or similar middleware technologies.
Experience with monitoring, logging, and alerting tools.
Proficiency in at least one scripting or programming language (e.g., Python, Bash, Java).
Solid knowledge of Linux, networking fundamentals, and system troubleshooting.
Ability to troubleshoot complex, multi‑component issues under pressure.
Preferred Qualifications
Experience operating enterprise‑scale messaging or event‑driven platforms.
Familiarity with clustering, replication, persistence, and high‑availability patterns.
Experience working in regulated environments with strong change management practices.
Exposure to automation, reliability engineering, or SRE best practices.
Corporate Security Responsibility
Every person working for, or on behalf of, Mastercard is responsible for information security. All activities involving access to Mastercard information assets must comply with Mastercard policies and standards.
Career Development
This role supports growth across SRE levels (SRE I through Principal), with increasing expectations around:
Scope and complexity of platform ownership
Incident leadership and RCA quality
Automation and systemic reliability improvements
Influence and mentorship across teams
Other Requirements
Relevant certifications in Site Reliability Engineering (SRE) or Cloud Services are a plus.
📌 Senior Site Reliability Engineer Lead (Pune)
🏢 HCL Technologies
📍 Pune