Site Reliability Engineering Lead (Application SRE Lead) (Bengaluru)

Site Reliability Engineering Lead (Application SRE Lead) (Bengaluru)

29 Aug
|
Hirexa Solutions
|
Bengaluru

29 Aug

Hirexa Solutions

Bengaluru

Details

Job Title*

Site Reliability Engineering Lead (Application SRE Lead)

Role Type*

Individual Contributor (80%) + People Leadership (20%)

Experience Level*

Min: 8 Years

Location*

Hyderabad

Work Type*

Full-time

Shift Requirement*

24x7 Production Support and On-call Rotation

Coding Requirement*

Mandatory

Technical Requirements

SRE & Production Engineering

- 8+ years of experience in Application Support, Site Reliability Engineering, Production Engineering, or Platform Operations.
- Minimum 2-4 years of experience managing technical teams in large-scale production environments.
- Hands-on experience supporting highly available, business-critical applications.
- Strong understanding of:
- SRE principles
- Reliability Engineering
- Operational Excellence
- Error Budgets
- Service Level Indicators (SLIs)
- Service Level Objectives (SLOs)
- Availability Engineering
- Capacity Planning

- Expertise managing production incidents, service outages, and major incident bridges.

Authentication & Security Platforms

Strong understanding and operational support experience in:

- OAuth 2.0
- OpenID Connect (OIDC)
- SAML
- Multi-Factor Authentication (MFA)
- Session Management
- Token Lifecycle Management
- API Authentication
- Okta
- Transmit Security
- Identity and Access Management (IAM)

Application Troubleshooting

Hands-on troubleshooting expertise involving:

- Java applications
- .NET Core applications
- REST APIs
- Microservices
- Batch processing applications
- Application latency issues
- Memory leaks
- Thread contention
- Authentication failures
- Service-to-service communication issues

Cloud Platforms

Experience managing applications hosted on:

- Google Cloud Platform (Preferred)
- Microsoft Azure
- Amazon Web Services
- PCF (Pivotal Cloud Foundry)

Areas of expertise:

- Application hosting
- Networking
- Cloud security
- Service reliability
- Monitoring
- Disaster recovery

Containers & Microservices

Strong experience with:

- Kubernetes
- Docker
- Helm
- Service Mesh concepts
- Container troubleshooting
- Pod lifecycle management
- Ingress controllers




- Service discovery
- Scaling strategies

Observability & Monitoring

Hands-on expertise in:

- Splunk
- AppDynamics
- Datadog
- Grafana
- ThousandEyes
- ITRS Geneos
- MoogSoft
- AppMetrics

Experience with:

- Log Analytics
- Distributed Tracing
- Metrics Monitoring
- Synthetic Monitoring
- Alert Tuning
- Dashboard Creation
- Event Correlation

Database & Messaging

Strong understanding of:

- MongoDB
- PostgreSQL
- SQL Query Optimization
- Database Performance Tuning
- Replication Issues
- Database Connectivity Troubleshooting
- Kafka Monitoring and Troubleshooting

Automation & Scripting

Strong coding and automation skills using:

- Shell Scripting
- Python
- PowerShell
- Go (Preferred)
- Java (Preferred)

Experience automating:

- Operational runbooks
- Deployment validation
- Monitoring
- Incident remediation
- Service recovery procedures

CI/CD & Release Management

Experience with:

- Harness
- Bamboo
- Bitbucket Pipelines
- GitHub Actions
- Jenkins
- Azure DevOps

Strong understanding of:

- CI/CD
- Release Engineering
- Deployment Strategies
- Blue-Green Deployments
- Canary Deployments
- Rollback Procedures

ITIL & ITSM

Robust experience in:

- Incident Management
- Problem Management
- Change Management
- Release Management
- Knowledge Management

Tools:

- ServiceNow
- Remedy
- JIRA

Key Responsibilities

Delivery & Technical Contribution (80%)

Reliability Engineering

- Ensure overall platform reliability, availability, and performance.
- Drive continuous improvements to reduce incidents and operational risks.
- Design and implement SLI/SLO frameworks.
- Monitor service health and proactively address burn-rate violations.

Production Support





- Lead Sev1 and Sev2 incident investigations.
- Drive service restoration and stakeholder communication.
- Manage major incident bridges and technical war rooms.
- Perform Root Cause Analysis and post-mortem reviews.

Application Operations

- Troubleshoot complex application, infrastructure, database, and network issues.
- Support authentication and authorization services globally.
- Monitor business-critical transaction flows.
- Build advanced monitoring dashboards and synthetic health checks.

Deployments & Release Management

- Support production deployments and releases.
- Lead deployment readiness assessments.
- Ensure successful rollout and rollback execution.
- Drive release governance processes.

Automation & Engineering Excellence

- Automate operational processes and runbooks.
- Implement self-healing and auto-remediation capabilities.
- Identify toil reduction opportunities.
- Improve MTTD and MTTR metrics.

Agile Delivery

- Participate in sprint planning.
- Own and deliver assigned epics and user stories.
- Review technical solutions and implementation approaches.
- Ensure operational readiness for new platform capabilities.

Leadership & Team Management (20%)

Team Leadership

- Lead, mentor, and coach SRE engineers.
- Conduct technical reviews and guidance sessions.
- Support career development initiatives.
- Build a culture of operational excellence.

Delivery Governance

- Ensure SLA, SLO, and operational commitments are consistently achieved.
- Monitor service delivery metrics.
- Review team performance and workload distribution.
- Drive capacity and resource planning.

Stakeholder Management

- Act as primary escalation point for critical incidents.
- Communicate effectively with business, engineering, and executive leadership.
- Manage client expectations during outages and major events.

Continuous Improvement

- Drive service improvement initiatives.
- Lead automation programs.
- Improve reliability maturity across application portfolios.
- Contribute to organizational SRE best practices.

📌 Site Reliability Engineering Lead (Application SRE Lead) (Bengaluru)
🏢 Hirexa Solutions
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineering lead (application sre lead) (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineering lead (application sre lead) (bengaluru) / bengaluru