06 Sep
|
Sparix Global
|
India
06 Sep
Sparix Global
India
Location
- Hybrid /Banglore
Time
- 7:30 AM-5 PM
""" Summary A Site Reliability Engineer (SRE) applies software engineering practices to operations to ensure services are reliable, scalable, and efficient. The SRE partner with product and platform teams to define SLIs/SLOs, automate operations, runbook and incident engineering, and drive long‑term reliability improvements.
Core Responsibilities
· Service Reliability: Define SLIs/SLOs and monitor error budgets; drive actions when SLOs are at risk.
· Incident Management: Lead on‑call rotations, perform incident response, run postmortems, and implement corrective actions.
· Automation &
- Tooling: Build automation for deployment, remediation, scaling, and runbook tasks to reduce manual toil.
· Observability: Design and maintain metrics, logs, and distributed tracing to support rapid diagnosis and capacity planning.
· Performance &
- Capacity: Run load tests, capacity planning, and tuning to meet performance targets.
· Resilience Engineering: Implement canaries, chaos testing, circuit breakers, and progressive rollouts.
· Platform Improvement: Collaborate with dev teams to productionize features, harden services, and reduce operational risk.
· Knowledge Sharing: Produce runbooks, run regular reliability reviews, and coach teams on best practices.
Required Skills &
- Experience
· Core: Strong programming/scripting (Python, Go, or equivalent),
systems fundamentals (Linux), networking, and debugging skills.
· Observability: Experience with metrics platforms, logging, and tracing (Prometheus, Grafana, ELK/Opentelemetry or equivalent).
· Cloud &
- Infra: Familiarity with cloud platforms (AWS/Azure/GCP), containers, orchestration (Kubernetes), and IaC (Terraform/ARM).
· Operational Practice: Proven incident leadership, postmortem facilitation, and on‑call experience.
· Automation: CI/CD pipelines, deployment automation, and infrastructure automation experience.
· Soft skills: Explicit communicator, calm under pressure, collaborative, and able to drive cross‑team change.
Preferred Qualifications
· Bachelor's in computer science, Engineering, or equivalent experience.
· Experience defining SLIs/SLOs and using error budget processes.
· Prior SRE, production operations, or systems engineering role.
· Certifications or courses in cloud, Kubernetes, or reliability engineering (optional).
Success Metrics (KPIs)
· Achieve and sustain defined SLO targets for owned services.
· Mean Time To Detect (MTTD) and Mean Time To Recover (MTTR) improvements quarter over quarter.
· Reduction in manual toil hours via automation.
· Number and quality of postmortem action items closed within SLA.
· Error budget burn rate and governance adherence."""
📌 SRE (India)
🏢 Sparix Global
📍 India