- Ensure overall platform reliability, availability, and performance.
- Drive continuous improvements to reduce incidents and operational risks.
- Design and implement SLI/SLO frameworks.
- Monitor service health and proactively address burn-rate violations.
Production Support
- Lead Sev1 and Sev2 incident investigations.
- Drive service restoration and stakeholder communication.
- Manage major incident bridges and technical war rooms.
- Perform Root Cause Analysis and post-mortem reviews.
Application Operations
- Troubleshoot complex application, infrastructure, database, and network issues.
- Support authentication and authorization services globally.
- Monitor business-critical transaction flows.
- Build advanced monitoring dashboards and synthetic health checks.
Deployments & Release Management
- Support production deployments and releases.
- Lead deployment readiness assessments.
- Ensure successful rollout and rollback execution.
- Drive release governance processes.
Automation & Engineering Excellence
- Automate operational processes and runbooks.
- Implement self-healing and auto-remediation capabilities.
- Identify toil reduction opportunities.
- Improve MTTD and MTTR metrics.
Agile Delivery
- Participate in sprint planning.
- Own and deliver assigned epics and user stories.
- Review technical solutions and implementation approaches.
- Ensure operational readiness for new platform capabilities.
Leadership & Team Management (20%)
Team Leadership
- Lead, mentor, and coach SRE engineers.
- Conduct technical reviews and guidance sessions.
- Support career development initiatives.
- Build a culture of operational excellence.
Delivery Governance
- Ensure SLA, SLO, and operational commitments are consistently achieved.
- Monitor service delivery metrics.
- Review team performance and workload distribution.
- Drive capacity and resource planning.
Stakeholder Management
- Act as primary escalation point for critical incidents.
- Communicate effectively with business, engineering, and executive leadership.
- Manage client expectations during outages and major events.
Continuous Improvement
- Drive service improvement initiatives.
- Lead automation programs.
- Improve reliability maturity across application portfolios.
- Contribute to organizational SRE best practices.
📌 Site Reliability Engineering Lead (Application SRE Lead) (Bengaluru)
🏢 Hirexa Solutions
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.