Sr. Site Reliability Engineer (Chennai)

Sr. Site Reliability Engineer (Chennai)

15 Sep
|
Xoom
|
Chennai

15 Sep

Xoom

Chennai

Job Summary

This is an incident command role. You'll direct application and infrastructure teams during incidents making work assignments, prioritizing troubleshooting paths, and authorizing critical actions like rollbacks and regional failovers. You need the technical depth to rapidly read Infrastructure as Code, Kubernetes manifests, and CI/CD configurations to make informed decisions under pressure. Meet Our Team The Site Health Engineering (Command Center) team serves as the operational and technical authority during PayPal's most critical incidents. We are a team of experienced infrastructure and reliability professionals who blend deep technical expertise with sound judgment under pressure, directing cross-functional engineering efforts across PayPal's core platforms and family of brands, including Venmo, Xoom, Zettle, and Braintree.

Essential Responsibilities

- Delivers complete solutions spanning all phases of the Software Development Lifecycle (SDLC) (design, implementation, testing, delivery and operations), based on definitions from more senior roles.
- Advises immediate management on project-level issues
- Guides junior engineers
- Operates with little day-to-day supervision, making technical decisions based on knowledge of internal conventions and industry best practices
- Applies knowledge of technical best practices in making decisions

Minimum Qualifications

- 3+ years relevant experience and a Bachelor's degree OR Any equivalent combination of education and experience.

Your Way to Impact

Rather than building or maintaining infrastructure day-to-day, our team is entrusted with a broader mandate: safeguarding the reliability, resiliency,



and availability of some of the world's most heavily trafficked financial platforms. We hold final decision-making authority during high-severity incidents, partner closely with executive leadership on post-incident learnings, and drive the tooling and processes that continuously strengthen our incident response capabilities. You'll also regularly interface with executive leadership during critical incidents and post-mortems, and drive implementation of tooling that advances the Command Centers capabilities.

Site Resiliency & Infrastructure Management

- Proactively identify and address vulnerabilities in cloud (AWS, GCP, Azure) and on-premises infrastructure
- Review Infrastructure as Code changes for reliability risks as part of change approval process
- Identify architectural anti-patterns in Kubernetes deployments and cloud migrations
- Conduct regular disaster recovery drills and readiness tests before major events (Thanksgiving, Cyber 5, peak shopping seasons)
- Participate in situation room activities for recent product rollouts
- Drive site resilience projects to enhance system reliability and uptime
- Implement automated monitoring solutions to detect single points of failure
- Lead new datacenter and CDN certification initiatives

Incident Management & Response





- Act as incident commander with final decision authority -- directing engineering teams, authorizing rollbacks, and commanding regional failovers
- Direct application and infrastructure teams during incidents by making work assignments and prioritizing troubleshooting paths
- Rapidly assess incidents by reading Infrastructure as Code (Terraform, CloudFormation), Kubernetes manifests, and CI/CD configurations
- Give final authorization for critical actions including production rollbacks, regional failovers, and emergency changes
- Interface with executive leadership during critical incidents and post-mortems to provide technical guidance and impact assessments
- Identify when incidents stem from teams deviating from established cloud-native patterns
- Command cross-functional teams during high-severity incidents affecting PayPal core and brand platforms (Venmo, Xoom, Zettle, Braintree)
- Lead blameless postmortem sessions and contribute to Root Cause Analysis (RCA) processes
- Drive continuous improvement initiatives based on incident learnings
- Serve as the primary technical escalation point during critical incidents
- Accelerate incident response times through standardized playbooks and automated workflows
- Coordinate cross-functional teams during high-severity incidents affecting PayPal core and brand platforms (Venmo, Xoom, Zettle, Braintree)

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Sr. Site Reliability Engineer (Chennai)
🏢 Xoom
📍 Chennai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sr. site reliability engineer (chennai) / chennai