Senior Site Reliability Engineer (Bengaluru)

Senior Site Reliability Engineer (Bengaluru)

17 Sep
|
Laerdal Medical
|
Bengaluru

17 Sep

Laerdal Medical

Bengaluru

Role & responsibilities

We are looking for a Senior Site Reliability Engineer (SRE) to drive system reliability, scalability, and operational excellence across our production environments.

This is a hands-on, high-impact Individual Contributor role where you will:

- Own production reliability and availability
- Define and enforce SRE practices and standards
- Lead incident response and operational excellence
- Coach and guide engineers on building reliable systems

You will play a critical role in ensuring systems are measurable, observable, resilient, and scalable.

Core SRE Responsibilities

Reliability Engineering & System Health
- Ensure high availability, reliability, and fault tolerance of distributed systems.

- Define and implement

- Service Level Indicators (SLIs)

- Service Level Objectives (SLOs)

- Error Budgets
- Continuously improve system reliability using data-driven approaches.
- Track and improve key reliability metrics:

- Availability

- Error rates

- Latency

Observability, Monitoring & APM
- Build and maintain end-to-end observability systems:

- Metrics

- Logs

- Distributed tracing

- Implement effective

- Monitoring
- Alerting (low noise, high signal)
- Use APM tools to track:

- Latency

- Error rates

- Service dependencies

Scalability & Performance Engineering
- Ensure systems scale efficiently under load.

- Perform

- Capacity planning
- Load testing & performance tuning

- Optimize for

- Latency

- Throughput

- Resource utilization

Automation & Toil Reduction
- Identify and eliminate manual, repetitive operational tasks (toil).

- Build automation for

- Incident response

- Remediation workflows

- System recovery




- Improve engineering productivity through self-healing systems.

Security & Reliability Integration
- Ensure systems follow secure-by-design principles.

- Integrate reliability with

- Security controls

- Compliance requirements
- Partner with security teams on incident handling and risk mitigation.

Coaching & Technical Leadership
- Mentor engineers on:

- SRE principles

- Production readiness

- Reliability engineering
- Drive adoption of SRE best practices across teams.
- Influence architecture for resilient system design.

Operational Excellence & Incident Management
- Own and lead production incident management (L2/L3 support).
- Act as Incident Commander during critical outages.
- Drive blameless postmortems (RCA) and ensure action closure.

- Improve
- Mean Time to Detect (MTTD)
- Mean Time to Recover (MTTR)
- Mean Time Between Failures (MTBF)

Establish solid on-call practices, runbooks, and escalation mechanisms. Key Reliability Metrics (Must Have Experience)

Candidates must have hands-on experience defining, measuring, and improving:

- Availability (%)
- Mean Time Between Failures (MTBF)
- Mean Time to Recovery (MTTR)
- Mean Time to Failure (MTTF)
- Error Rates / Failure Rates
- SLO Compliance
- Error Budget Management
- Customer Impact Metrics (user-facing reliability)

CI/CD & Automated Testing




- Ensure reliable and safe deployments through:
- Automated testing (unit, integration, reliability testing)

- Automated deployment pipelines
- Improve deployment success rate and rollback mechanisms.
- Support progressive delivery (canary, blue/green deployments).

Disaster Recovery & Resilience
- Design and validate Disaster Recovery (DR) strategies.

- Define and test

- RTO (Recovery Time Objective)

- RPO (Recovery Point Objective)
- Conduct failure testing / chaos engineering where applicable.

Required Qualifications
- 510 years of experience, with minimum 3+ years in a dedicated SRE role

- Strong experience in

- Production systems ownership
- Incident management & on-call

- Root Cause Analysis (RCA)
- Hands-on experience with:
- Reliability metrics (SLO/SLI/Error Budgets)
- Observability & monitoring systems

- Performance and scalability engineering

- Strong experience with

- AWS cloud

- Kubernetes (EKS) in production

- DevOps and Platform Engineering
- Proficiency in one programming language (Go, Python, Java)

Proven ability to handle and resolve high-severity production incidents

Preferred candidate profile

- Experience in chaos engineering / resilience testing
- Strong background in APM tools (Datadog, Prometheus, Grafana, etc.)
- Experience defining organization-wide SRE practices
- Experience mentoring engineers and leading reliability initiatives

Ideal Candidate Profile We are looking for engineers who:
- Have owned production systems end-to-end
- Have led incidents and handled outages
- Have defined and worked with SLOs and error budgets
- Think in terms of system reliability, not just infrastructure

📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Laerdal Medical
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru