02 Sep
|
BayOne Solutions
|
Hyderabad
02 Sep
BayOne Solutions
Hyderabad
We are seeking a Senior Site Reliability Engineer with 5+ years of experience to lead the reliability engineering strategy for our core banking infrastructure and distributed transaction processing applications. You will design self-healing architectures, establish SLOs/SLAs, lead major incident resolution, and drive zero-downtime architecture for mission-critical financial platforms.
Key Responsibilities
- Reliability Architecture &
- Design:
Partner with software architects to design fault-tolerant, multi-region distributed banking systems capable of processing millions of daily transactions.
- SLO &
- Error Budget Management:
Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets alongside product managers to balance feature velocity with system stability.
- Advanced Automation &
- IaC:
Implement Infrastructure as Code (IaC) using Terraform, Ansible, and Kubernetes operators. Build automated self-healing mechanisms to resolve known failure modes without human intervention.
- Incident Leadership &
- Postmortems:
Lead response for high-priority (P1/P2) banking outages. Facilitate blameless post-mortems and enforce long-term root cause remediations.
- Capacity &
- Chaos Engineering:
Forecast system growth, conduct load testing under peak banking hours,
and run Chaos Engineering experiments (Gremlin, Chaos Mesh) to uncover hidden vulnerabilities.
- Compliance &
- Security:
Ensure infrastructure complies with banking regulatory frameworks (PCI-DSS, SOC2, Central Bank guidelines) and lead automated compliance auditing tools.
Required Qualifications &
- Skills
- Education: Btech or
Bachelor’s or Master’s degree in Computer Science, Engineering, or equivalent practical experience.
- Deep Experience:
5+ years in SRE, DevOps, or Software Engineering, with at least 2 years in banking, fintech, or high-volume transactional environments.
- Orchestration &
- Cloud:
Expert-level knowledge of Kubernetes (CKA certified preferred) and public/hybrid cloud enterprise architectures (AWS/Azure/GCP).
- Software Development:
Robust programming skills in Go, Python, or Java for building internal SRE tooling, CLI utilities, and automated controllers.
- Data Stores:
Understanding of high-availability relational (PostgreSQL, Oracle) and distributed non-relational databases (Cassandra, Redis, Kafka).
- Observability:
Mastery of ELK/EFK stack, OpenTelemetry, Prometheus, Cortex/Thanos, and distributed tracing (Jaeger/Zipkin).
📌 Senior Site Reliability Engineer (Hyderabad)
🏢 BayOne Solutions
📍 Hyderabad