09 Sep
|
HireTalent
|
Bengaluru
09 Sep
HireTalent
Bengaluru
The position is hands-on and establishes SRE best practices, automation standards and observability capabilities for Spire. The role contributes to the long-term evolution of reliability engineering at EquiLend, starting with Spire and potentially expanding to the wider product base over time.
Key Responsibilities
- Build and maintain the reliability, availability and performance posture of the Spire platform through application of SRE principles and software engineering practices
- Establish and embed operational standards, governance models and reliability processes across engineering and operations teams
- Design, implement and optimise observability capabilities - including metrics, logs, traces and alerting - to proactively identify performance risk and capacity constraints
- Define and maintain SLIs, SLOs and SLAs for critical Spire services and ensure alignment across engineering and product functions
- Lead capacity planning activities including utilisation analysis, performance modelling and forecasting to support future growth
- Develop automation for deployments, infrastructure provisioning, monitoring, testing and operational workflows to reduce manual effort and ensure repeatability
- Implement and maintain CI/CD pipelines, infrastructure-as-code tooling and structured release processes to support secure and reliable delivery
- Conduct root cause analysis for production incidents, identify systemic issues and deliver long-term corrective actions to prevent reoccurrence
- Drive performance optimisation initiatives by analysing system behaviour, identifying bottlenecks and implementing effective remediation
- Strengthen operational readiness by reviewing proposed changes, assessing reliability impact and establishing safeguards for high-risk deployments
- Develop runbooks, architectural artefacts and operational documentation to support consistent and safe platform operations
- Collaborate with cross-functional stakeholders—including engineering, product and operations—to align reliability requirements with business priorities and support post-incident review processes
- Contribute to the future definition and maturation of the SRE capability within EquiLend, supporting adoption of best practice across platforms as the function evolves
Required Experience and Skills
- 8+ years of experience in site reliability engineering, systems engineering or software engineering within complex, high-availability environments
- Strong hands-on background in cloud platforms such as AWS, Azure or GCP (we primarily use AWS)
- Expertise in Kubernetes, container orchestration and distributed systems
- Proficiency in programming languages such as Python or Go for automation, tooling and system optimisation
- Proven experience building observability platforms using tools such as Prometheus, Grafana, ELK, Datadog or similar technologies
- Demonstrated capability in capacity planning, scalability modelling, system performance analysis and proactive reliability management
- Strong background in incident response and operational stability, including structured root cause analysis and implementation of preventative measures
- Experience operating large-scale systems within financial services or similarly regulated, high-throughput environments
- Understanding of security, compliance and resilience principles relevant to financial technology platforms
- Effective communication and stakeholder engagement skills suitable for cross-functional collaboration
- Proactive mindset with the ability to identify risks, anticipate operational challenges and implement forward-looking solutions
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender, identity, national origin, disability, or protected veteran status.
📌 Specialist Site Reliability Engineer (Bengaluru)
🏢 HireTalent
📍 Bengaluru