Site Reliability Engineer (India)

Site Reliability Engineer (India)

16 Aug
|
Terminus Technologies
|
India

16 Aug

Terminus Technologies

India

Freelancing Opportunity

Site Reliability Engineer (SRE)

Experience: 5+ Years

Job Overview

We are looking for an experienced Site Reliability Engineer (SRE) to improve system reliability, observability, performance, and operational efficiency. The ideal candidate will have strong production engineering experience and hands-on expertise in monitoring, alerting, incident management, and reliability engineering.

Key Responsibilities

- Build and maintain Grafana dashboards for system and application monitoring.
- Manage monitoring and alerting using Prometheus, New Relic, and PagerDuty.
- Review Critical User Journeys (CUJs) and identify reliability, performance, and availability risks.
- Define and monitor SLIs, SLOs, SLAs, and error budgets.
- Develop and maintain detailed runbooks for incident response, troubleshooting, and operational procedures.
- Participate in incident management, RCA, postmortems, and reliability improvement initiatives.
- Improve alert quality, reduce alert noise, and strengthen production observability.
- Collaborate with engineering and platform teams to improve production stability and performance.
- Leverage AI/LLM tools such as Claude for troubleshooting, automation, documentation, and operational workflows.
- Automate repetitive operational and troubleshooting tasks wherever possible.

Required Skills

- 5+ years of experience in SRE, Production Engineering, or similar reliability-focused roles.
- Strong hands-on experience with Grafana and Prometheus.




- Experience with New Relic and PagerDuty or equivalent observability/incident-management tools.
- Strong understanding of monitoring, alerting, incident response, and RCA.
- Good understanding of SLI, SLO, SLA, error budgets, and reliability engineering principles.
- Experience creating and maintaining runbooks and operational documentation.
- Experience troubleshooting production systems and performance issues.
- Practical experience using AI/LLM tools such as Claude or ChatGPT for engineering productivity and automation.
- Strong communication, analytical, and problem-solving skills.

Good to Have

- Experience with AWS, Azure, or GCP.
- Experience with Kubernetes and Docker.
- Knowledge of OpenTelemetry and distributed tracing.
- Experience with Python, Bash, or similar scripting for automation.
- Experience with synthetic monitoring and Critical User Journey monitoring.
- Experience implementing observability and reliability standards across multiple services.

Ideal Candidate

A hands-on SRE who can monitor, troubleshoot, respond to incidents, perform RCA, improve observability, define reliability metrics, and automate operational processes in a production workplace.

Pay: ₹10,000.00 - ₹20,000.00 per month

Benefits:

- Flexible schedule

Education:

- Bachelor's (Required)

Experience:

- Site Reliability Engineer (SRE): 5 years (Required)

Shift availability:

- Night Shift (Required)
- Day Shift (Required)

Work Location: Remote

📌 Site Reliability Engineer (India)
🏢 Terminus Technologies
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (india) / india