17 Sep
|
Laerdal Medical
|
Bengaluru
17 Sep
Laerdal Medical
Bengaluru
Role & responsibilities
We are looking for a Senior Site Reliability Engineer (SRE) to drive system reliability, scalability, and operational excellence across our production environments.
This is a hands-on, high-impact Individual Contributor role where you will:
- Own production reliability and availability
- Define and enforce SRE practices and standards
- Lead incident response and operational excellence
- Coach and guide engineers on building reliable systems
You will play a critical role in ensuring systems are measurable, observable, resilient, and scalable.
Core SRE Responsibilities
Reliability Engineering & System Health
- Ensure high availability, reliability, and fault tolerance of distributed systems.
- Define and implement
- Service Level Indicators (SLIs)
- Service Level Objectives (SLOs)
- Error Budgets
- Continuously improve system reliability using data-driven approaches.
- Track and improve key reliability metrics:
- Availability
- Error rates
- Latency
Observability, Monitoring & APM
- Build and maintain end-to-end observability systems:
- Metrics
- Logs
- Distributed tracing
- Implement effective
- Monitoring
- Alerting (low noise, high signal)
- Use APM tools to track:
- Latency
- Error rates
- Service dependencies
Scalability & Performance Engineering
- Ensure systems scale efficiently under load.
- Perform
- Capacity planning
- Load testing & performance tuning
- Optimize for
- Latency
- Throughput
- Resource utilization
Automation & Toil Reduction
- Identify and eliminate manual, repetitive operational tasks (toil).
- Build automation for
- Incident response
- Remediation workflows
- System recovery
- Improve engineering productivity through self-healing systems.
Security & Reliability Integration
- Ensure systems follow secure-by-design principles.
- Integrate reliability with
- Security controls
- Compliance requirements
- Partner with security teams on incident handling and risk mitigation.
Coaching & Technical Leadership
- Mentor engineers on:
- SRE principles
- Production readiness
- Reliability engineering
- Drive adoption of SRE best practices across teams.
- Influence architecture for resilient system design.
Operational Excellence & Incident Management
- Own and lead production incident management (L2/L3 support).
- Act as Incident Commander during critical outages.
- Drive blameless postmortems (RCA) and ensure action closure.
- Improve
- Mean Time to Detect (MTTD)
- Mean Time to Recover (MTTR)
- Mean Time Between Failures (MTBF)
Establish solid on-call practices, runbooks, and escalation mechanisms. Key Reliability Metrics (Must Have Experience)
Candidates must have hands-on experience defining, measuring, and improving:
- Availability (%)
- Mean Time Between Failures (MTBF)
- Mean Time to Recovery (MTTR)
- Mean Time to Failure (MTTF)
- Error Rates / Failure Rates
- SLO Compliance
- Error Budget Management
- Customer Impact Metrics (user-facing reliability)
CI/CD & Automated Testing
- Ensure reliable and safe deployments through:
- Automated testing (unit, integration, reliability testing)
- Automated deployment pipelines
- Improve deployment success rate and rollback mechanisms.
- Support progressive delivery (canary, blue/green deployments).
Disaster Recovery & Resilience
- Design and validate Disaster Recovery (DR) strategies.
- Define and test
- RTO (Recovery Time Objective)
- RPO (Recovery Point Objective)
- Conduct failure testing / chaos engineering where applicable.
Required Qualifications
- 510 years of experience, with minimum 3+ years in a dedicated SRE role
- Strong experience in
- Production systems ownership
- Incident management & on-call
- Root Cause Analysis (RCA)
- Hands-on experience with:
- Reliability metrics (SLO/SLI/Error Budgets)
- Observability & monitoring systems
- Performance and scalability engineering
- Strong experience with
- AWS cloud
- Kubernetes (EKS) in production
- DevOps and Platform Engineering
- Proficiency in one programming language (Go, Python, Java)
Proven ability to handle and resolve high-severity production incidents
Preferred candidate profile
- Experience in chaos engineering / resilience testing
- Strong background in APM tools (Datadog, Prometheus, Grafana, etc.)
- Experience defining organization-wide SRE practices
- Experience mentoring engineers and leading reliability initiatives
Ideal Candidate Profile We are looking for engineers who:
- Have owned production systems end-to-end
- Have led incidents and handled outages
- Have defined and worked with SLOs and error budgets
- Think in terms of system reliability, not just infrastructure
📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Laerdal Medical
📍 Bengaluru