07 Aug
|
Ig Group
|
Bengaluru
07 Aug
Ig Group
Bengaluru
Job Title
Senior Platform SRE
Job Description
The Platform SRE team is the engine of IG's reliability programme. We sit within Infrastructure & Operations, working across IG's hybrid estate of on-premises HashiCorp Nomad and AWS.
We are not a reactive ops team. We build the platform, standards, and tooling that make reliability the default for every engineering team at IG. Through the SRE Guild, we connect with Domain SREs and Reliability Champions across the organisation, setting the bar and lifting it together.
Your team
The Platform SRE team is the engine of IG's reliability programme. We sit within Infrastructure & Operations, working across IG's hybrid estate of on-premises HashiCorp Nomad and AWS.
We are not a reactive ops team. We build the platform, standards, and tooling that make reliability the default for every engineering team at IG. Through the SRE Guild, we connect with Domain SREs and Reliability Champions across the organisation, setting the bar and lifting it together.
Your role in the Teams Success
You will be a hands-on technical contributor at the heart of the Platform SRE team, owning pieces of the reliability platform that hundreds of engineers depend on. You will work at the intersection of software engineering, observability, and systems reliability, turning reliability from a reactive concern into a proactive engineering discipline.
You will partner with Platform Engineering, product teams, and Reliability Champions to define what valuable looks like in production and then make it the default. You will contribute to the SRE Guild, mentor engineers across the organisation, and when things go wrong, you will be on the call helping to mitigate, understand, and prevent a repeat.
What youll do
- Build and own the reliability platform
- Implement comprehensive monitoring and observability using OpenTelemetry and distributed tracing. Maintain SLO, error budgets, and burn-rate tracking.
- Establish and maintain 24/7 operational readiness including automated deployments, blue/green releases, and zero-downtime patching strategies
- Engineer self-healing capabilities: auto-remediation, error-budget-gated rollback, and automated traffic rerouting
- Design and run chaos experiments across the AWS estate, turning severe-but-plausible failure scenarios into engineering improvements
- Build automation tools and CI/CD pipelines that embed reliability practices, while applying software engineering discipline including version control, code reviews, and testing
- Contribute to the SRE AI agent, IG's agentic tooling for incident investigation and reliability review, built on AWS frontier models
- Mentor junior SREs and Reliability Champions on reliability patterns and production engineering discipline
Set and uphold standards
- Author and evolve the SRE standards that underpin the Guild: SLO methodology, error budget policy, observability instrumentation guide, and Production Readiness Review (PRR) checklist
- Mentor developers on reliability patterns including circuit breakers, retry logic, and fault tolerance
- Work with development teams and Reliability Champions to design SLOs on customer journeys rather than per-service
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Senior Platform SRE Professional (Bengaluru)
🏢 Ig Group
📍 Bengaluru