10 Sep
|
Sigma Allied Services
|
India
10 Sep
Sigma Allied Services
India
7yrs to 10yrs
Pan India- Hybrid
Note: UI is preferred
SRE JD: Site Reliability Engineer (SRE) GPD Operations
Job Summary
The GPD Ops SRE Engineer will be responsible for improving the reliability, scalability, observability, and operational excellence of Government Digital Portals (GPD). This role blends strong software engineering skills, SRE best practices, and UIcentric operational tooling to reduce toil, automate Ops workflows, and deliver highavailability, productionready systems, especially during critical events such as AEP.
The engineer will work closely with scrum teams, platform engineers, and Ops leadership to implement endtoend monitoring, automation, dashboards, and reliability improvements, aligned with GPDs SRE and AIenabled Ops roadmap
Key Responsibilities:
SRE & Reliability Engineering
- Own availability, performance, resiliency, and reliability of GPD applications and platforms.
Programming & Automation
- Design and develop automation and reliability tooling using a primary programming language (Python preferred).
- Build automation for: Daily and release checkouts
- Health validations
- Certificate and dependency checks
- Incident triage and remediation workflows
- Apply engineeringfirst approach to Ops by reducing manual toil and increasing repeatability.
Observability & Monitoring
- Implement and enhance endtoend observability across application, infrastructure, and network layers.
- Build and maintain Dashboards using tools such as Dynatrace, Splunk, Elastic, Azure Monitor, and custom telemetry pipelines.
- Ensure monitoring supports deepdive triage, live Ops visibility, and executivelevel reporting.
- Partner with engineering teams during feature grooming, demos, and releases to ensure production readiness
- Ability to participate effectively in WAR Rooms.
UIDriven Ops & Dashboards
- Design and develop UIbased Ops tools, dashboards, and visualizations to improve:
- Incident triage
- Release health visibility , Operational metrics reporting
- Collaborate with frontend teams to implement SPAbased Ops dashboards and internal tooling with strong UX principles.
- Translate complex Ops data into clear, actionable visual insights.
Required Qualifications:
- BS Degree in Computer Science or related experience
- 3-5 + years on UI and Python/Java/J2ee experience
PRIMARY RESPONSIBILITIES
MUST HAVE
- Strong proficiency in at least one programming language (Python/Java/UI)
- Experience with Azure cloud (preferred) or any other cloud platform (AWS, GCP).
- Knowledge on analyzing logs from Dynatrace, SPlunk, Azure Kubernetes , Elastic and triaging issues. Priority will be on Dynatrace, SPlunk and Kubernetes
VALUABLE TO HAVE
- Experience with Relational or NOSQL databases
- Ability to review Functional specifications and interpret into Technical specifications
- Experience with Version control tools like GIT.
- Hands on experience on CI/CD tools like GitHub, Jenkins, Docker etc.
- Incident Management and Support
- Oncall rotation and war room handling
📌 Site Reliability Engineer (India)
🏢 Sigma Allied Services
📍 India