Site Reliability Engineer (India)

Site Reliability Engineer (India)

10 Sep
|
Sigma Allied Services
|
India

10 Sep

Sigma Allied Services

India

7yrs to 10yrs

Pan India- Hybrid

Note: UI is preferred

SRE JD: Site Reliability Engineer (SRE) GPD Operations

Job Summary

The GPD Ops SRE Engineer will be responsible for improving the reliability, scalability, observability, and operational excellence of Government Digital Portals (GPD). This role blends strong software engineering skills, SRE best practices, and UIcentric operational tooling to reduce toil, automate Ops workflows, and deliver highavailability, productionready systems, especially during critical events such as AEP.

The engineer will work closely with scrum teams, platform engineers, and Ops leadership to implement endtoend monitoring, automation, dashboards, and reliability improvements, aligned with GPDs SRE and AIenabled Ops roadmap

Key Responsibilities:

SRE & Reliability Engineering

- Own availability, performance, resiliency, and reliability of GPD applications and platforms.

Programming & Automation

- Design and develop automation and reliability tooling using a primary programming language (Python preferred).
- Build automation for: Daily and release checkouts
- Health validations
- Certificate and dependency checks
- Incident triage and remediation workflows

- Apply engineeringfirst approach to Ops by reducing manual toil and increasing repeatability.

Observability & Monitoring

- Implement and enhance endtoend observability across application, infrastructure, and network layers.
- Build and maintain Dashboards using tools such as Dynatrace, Splunk, Elastic, Azure Monitor, and custom telemetry pipelines.




- Ensure monitoring supports deepdive triage, live Ops visibility, and executivelevel reporting.
- Partner with engineering teams during feature grooming, demos, and releases to ensure production readiness
- Ability to participate effectively in WAR Rooms.

UIDriven Ops & Dashboards

- Design and develop UIbased Ops tools, dashboards, and visualizations to improve:
- Incident triage
- Release health visibility , Operational metrics reporting

- Collaborate with frontend teams to implement SPAbased Ops dashboards and internal tooling with strong UX principles.
- Translate complex Ops data into clear, actionable visual insights.

Required Qualifications:

- BS Degree in Computer Science or related experience
- 3-5 + years on UI and Python/Java/J2ee experience

PRIMARY RESPONSIBILITIES

MUST HAVE

- Strong proficiency in at least one programming language (Python/Java/UI)
- Experience with Azure cloud (preferred) or any other cloud platform (AWS, GCP).
- Knowledge on analyzing logs from Dynatrace, SPlunk, Azure Kubernetes , Elastic and triaging issues. Priority will be on Dynatrace, SPlunk and Kubernetes

VALUABLE TO HAVE

- Experience with Relational or NOSQL databases
- Ability to review Functional specifications and interpret into Technical specifications
- Experience with Version control tools like GIT.
- Hands on experience on CI/CD tools like GitHub, Jenkins, Docker etc.
- Incident Management and Support
- Oncall rotation and war room handling

📌 Site Reliability Engineer (India)
🏢 Sigma Allied Services
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (india) / india