02 Sep
|
BayOne Solutions
|
Hyderabad
02 Sep
BayOne Solutions
Hyderabad
We are seeking a Senior Site Reliability Engineer with 5 years of experience to lead the reliability engineering strategy for our core banking infrastructure and distributed transaction processing applications. You will design self-healing architectures, establish SLOs/SLAs, lead major incident resolution, and drive zero-downtime architecture for mission-critical financial platforms.
Key Responsibilities
- Reliability Architecture & Design: Partner with software architects to design fault-tolerant, multi-region distributed banking systems capable of processing millions of daily transactions.
- SLO & Error Budget Management: Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets alongside product managers to balance feature velocity with system stability.
- Advanced Automation & IaC: Implement Infrastructure as Code (IaC) using Terraform, Ansible, and Kubernetes operators. Build automated self-healing mechanisms to resolve known failure modes without human intervention.
- Incident Leadership & Postmortems: Lead response for high-priority (P1/P2) banking outages. Facilitate blameless post-mortems and enforce long-term root cause remediations.
- Capacity & Chaos Engineering: Forecast system growth, conduct load testing under peak banking hours,
and run Chaos Engineering experiments (Gremlin, Chaos Mesh) to uncover hidden vulnerabilities.
- Compliance & Security: Ensure infrastructure complies with banking regulatory frameworks (PCI-DSS, SOC2, Central Bank guidelines) and lead automated compliance auditing tools.
Required Qualifications & Skills
- Education: Btech or Bachelor’s or Master’s degree in Computer Science, Engineering, or equivalent practical experience.
- Deep Experience: 5 years in SRE, DevOps, or Software Engineering, with at least 2 years in banking, fintech, or high-volume transactional environments.
- Orchestration & Cloud: Expert-level knowledge of Kubernetes (CKA certified preferred) and public/hybrid cloud enterprise architectures (AWS/Azure/GCP).
- Software Development: Solid programming skills in Go, Python, or Java for building internal SRE tooling, CLI utilities, and automated controllers.
- Data Stores: Understanding of high-availability relational (PostgreSQL, Oracle) and distributed non-relational databases (Cassandra, Redis, Kafka).
- Observability: Mastery of ELK/EFK stack, OpenTelemetry, Prometheus, Cortex/Thanos, and distributed tracing (Jaeger/Zipkin).
📌 Senior Site Reliability Engineer (Hyderabad)
🏢 BayOne Solutions
📍 Hyderabad