We are looking for a Mid-Level Site Reliability Engineer to join our global SRE team. This role goes beyond reactive incident response — you will handle incidents independently, contribute to reliability improvements, and help reduce operational toil across a fast-moving fintech infrastructure. You will work closely with global engineering teams to maintain system health and improve the resilience of our platforms.
Key Responsibilities
- On-Call & Incident Response:Independently handle incidents during scheduled shifts, coordinating communication with relevant teams and driving issues to resolution. Escalate complex or novel issues to senior engineers with clear context and initial findings.
- Monitoring & Alerting:Build and improve Datadog monitors and dashboards following team standards. Help reduce alert fatigue by identifying noisy alerts and proposing tuning improvements.
- Runbook & SOP Authorship:Create runbooks from scratch for recent failure modes, update existing SOPs to reflect operational learnings, and contribute meaningfully to post-mortems with structured root cause analysis and action items.
- Reliability Initiatives:Proactively identify sources of toil and operational inefficiency. Propose and implement automation or process improvements that reduce manual intervention and improve system resilience.
- Deployment Support:Monitor CI/CD pipelines during deployments, flag reliability risks, and initiate rollbacks following established procedures when stability is at risk.
Qualifications
- Experience:4-6 years of hands-on experience in SRE, DevOps, or Platform Engineering roles.
- AWS Expertise:Deep working knowledge of Amazon ECS, IAM, VPC, ALB/NLB, RDS, S3, MSK, ElastiCache, Lambda, CloudWatch, and an awareness of cost optimization practices.
- Infrastructure as Code:Proficiency