Job Title: Site Reliability Engineer
Location: Remote / India
Experience Level: Senior (6–8 years)
Role Type: Contract (Fixed Term)
Working Hours: Primarily in US timezone, details TBD
Position Overview:
We are looking for a Senior Site Reliability Engineer to join our global SRE team. This is a hands-on individual contributor role for an experienced engineer who can take ownership of high-severity incidents, and drive reliability improvements across a fast-moving fintech infrastructure. You will serve in the on-call rotation, work closely with global engineering teams to maintain system health and improve the resilience of our platforms.
Key Responsibilities:
On-Call & Incident Command: Serve in the on-call rotation and act as incident lead for incidents: drive triage, coordinate cross-team response, communicate status to stakeholders, and own the incident through resolution. Engage management or service owners when incidents require architectural decisions or business-level trade-offs. Incident Mitigation:
Independently triage and resolve production alerts across AWS infrastructure and application layers. Apply sound judgment to distinguish transient issues from systemic failures and act accordingly.
Monitoring & Alerting Strategy: Own and evolve Datadog monitoring standards: design monitors, dashboards, and SLO/SLI frameworks, and drive signal-to-noise improvements to reduce alert fatigue across the team.
Runbook & SOP Authorship: Create runbooks from scratch for current failure modes, update existing SOPs to reflect operational learnings, and lead post-mortems with structured root cause analysis and actionable follow-ups.
Reliability Initiatives: Lead initiatives to reduce toil and operational inefficiency: identify systemic weaknesses, design and implement automation or process improvements, and see them through to adoption.
Root Cause Analysis: Lead RCA investigations for infrastructure and application-level failures in the AWS environment. Produce clear, ac