12 Sep
|
Sigma Allied Services
|
Bengaluru
12 Sep
Sigma Allied Services
Bengaluru
Site Reliability Engineer (SRE) GPD Operations -
7yrs to 10yrs
Pan India- Hybrid
Note: UI is preferred
SRE JD: Site Reliability Engineer (SRE) GPD Operations
Job Summary The GPD Ops SRE Engineer will be responsible for improving the reliability, scalability, observability, and operational excellence of Government Digital Portals (GPD). This role blends strong software engineering skills, SRE best practices, and UIcentric operational tooling to reduce toil, automate Ops workflows, and deliver highavailability, productionready systems, especially during critical events such as AEP.The engineer will work closely with scrum teams, platform engineers, and Ops leadership to implement endtoend monitoring, automation, dashboards, and reliability improvements, aligned with GPD’s SRE and AIenabled Ops roadmap
Key Responsibilities
SRE & Reliability Engineering
Own availability, performance, resiliency, and reliability of GPD applications and platforms.Programming & AutomationDesign and develop automation and reliability tooling using a primary programming language (Python preferred).
Build automation for: Daily and release checkoutsHealth validationsCertificate and dependency checksIncident triage and remediation workflowsApply engineeringfirst approach to Ops by reducing manual toil and increasing repeatability.Observability & MonitoringImplement and enhance endtoend observability across application, infrastructure, and network layers.Build and maintain Dashboards using tools such as Dynatrace, Splunk, Elastic,
Azure Monitor, and custom telemetry pipelines.Ensure monitoring supports deepdive triage, live Ops visibility, and executivelevel reporting.Partner with engineering teams during feature grooming, demos, and releases to ensure production readinessAbility to participate effectively in WAR Rooms. UIDriven Ops & DashboardsDesign and develop UIbased Ops tools, dashboards, and visualizations to improve:Incident triageRelease health visibility , Operational metrics reportingCollaborate with frontend teams to implement SPAbased Ops dashboards and internal tooling with strong UX principles.Translate complex Ops data into transparent, actionable visual insights.
Required Qualifications
BS Degree in Computer Science
Experience
3-5 + years on UI and Python/Java/J2ee experience
PRIMARY RESPONSIBILITIES MUST HAVE
Strong proficiency in at least one programming language (Python/Java/UI)Experience with Azure cloud (preferred) or any other cloud platform (AWS, GCP).Knowledge on analyzing logs from Dynatrace, SPlunk, Azure Kubernetes , Elastic and triaging issues. Priority will be on Dynatrace, SPlunk and Kubernetes GOOD TO HAVEExperience with Relational or NOSQL databasesAbility to review Functional specifications and interpret into Technical specificationsExperience with Version control tools like GIT.Hands on experience on CI/CD tools like GitHub, Jenkins, Docker etc.Incident Management and SupportOncall rotation and war room handling
📌 Site Reliability EngineerH): PAN INDIA (Bengaluru)
🏢 Sigma Allied Services
📍 Bengaluru