28 Sep
|
Robotico Digital
|
India
28 Sep
Robotico Digital
India
Key Responsibilities
Monitoring Framework Ownership
- Own the design, implementation, and ongoing management of the Atlas5 monitoring framework across all 5 layers:
– Layer 1: Infrastructure (VMs, compute, networking)
– Layer 2: Data Tier (SQL DB, Cosmos DB, Elastic Pools, Microsoft Fabric)
– Layer 3: Application (App Services, Function Apps, BRE)
– Layer 4: Integration (EDA, Logic Apps, APIM)
– Layer 5: File Exchange (FTP/SFTP, file drop, inbound feeds)
- Define and configure alert thresholds, severity levels, and escalation paths for all layers
- Ensure every alert triggers a Jira incident ticket and DevOps channel notification within defined SLA
- Enforce mandatory RCA completion for all P1-Critical incidents before ticket closure
Tool Selection & Implementation
- Evaluate, recommend, and implement the right monitoring toolset for each layer — Azure Monitor, Prometheus, Grafana, New Relic, SigNoz, Zabbix, Dynatrace, or equivalent
- Build a cost-effective hybrid monitoring stack that balances coverage, operational overhead, and cost vs. native Azure Monitor
- Deploy and configure monitoring agents across all Atlas5 production infrastructure
- Integrate monitoring tools with Jira Service Management for automated incident creation
- Build and maintain unified monitoring dashboards for real-time 5-layer status visibility
SRE Practice & Incident Management
- Define and maintain SLOs, SLIs, and error budgets for all Atlas5 production services
- Lead incident response for P1-Critical production incidents — own the investigation, RCA, and permanent fix process
- Drive the monitoring gap assessment as a mandatory step in every incident RCA
- Configure auto-escalation for unacknowledged Critical alerts to DevOps Lead and secondary on-call
- Conduct regular alert quality reviews — target false positive rate below 5%
Azure Observability Engineering
Configure and manage Azure Monitor, Log Analytics workspaces, Application Insights, and Diagnostic Settings
- Design and manage KQL queries for alerting, dashboards, and operational reporting
- Optimise monitoring costs — migrate Log Search Alerts to Metric Alerts (30x cheaper) and Activity Log Alerts (free) where applicable
- Configure Prometheus scraping for Azure-native services and manage the Prometheus + Grafana stack
- Instrument applications with Open Telemetry for vendor-neutral, future-proof telemetry
Reliability & Performance Engineering
- Proactively identify reliability risks — throttling, capacity saturation, latency spikes, and dependency failures
- Partner with DevOps and application teams to instrument new services and define monitoring parameters
- Work with development team to define and implement EDA logging and integration layer monitoring parameters
- Contribute to deployment reliability — flag monitoring gaps in pre-deployment validation and change management
Required Experience & Skills
Core Requirements
- 8 – 10 years of experience in SRE, observability engineering, or cloud operations roles
- Proven experience designing and implementing monitoring frameworks from scratch in Azure environments
- Strong Azure Monitor expertise — Log Analytics, Application Insights,
Metric Alerts, Log Search Alerts, Diagnostic Settings
- Hands-on experience with at least 2 of: Prometheus, Grafana, New Relic, Datadog, Dynatrace, Zabbix, SigNoz
- Experience integrating monitoring with incident management tools (Jira, PagerDuty, or equivalent)
Technical Skills
- Azure: Azure Monitor, Log Analytics (KQL), Application Insights, Azure Alerts, Diagnostic Settings, Service Health
- Monitoring tools: Prometheus, Grafana, New Relic, Dynatrace, Zabbix, SigNoz — hands-on with at least 2
- Instrumentation: Open Telemetry (OTel) — logs, metrics, traces; Application Insights SDK
- Alerting: Alert rule design, severity taxonomy, escalation policies, on-call automation
- Scripting: KQL (mandatory), PowerShell or Python for alert automation and tooling
- Azure infrastructure: App Services, Function Apps, Cosmos DB, SQL DB, Microsoft Fabric, AKS — operational monitoring knowledge
- Cost optimisation: experience reducing Azure Monitor costs through Metric Alert migration and Log tier management
Nice to Have
- Experience in fintech or regulated financial services environments
- Familiarity with EDA (Event-Driven Architecture) and integration layer monitoring (Service Bus, Logic Apps, APIM)
- Experience with SigNoz or Open Telemetry-native observability stacks
- Dynatrace certification or hands-on Dynatrace deployment experience
Job Type: Full time
Application Question(s):
- Are you comfortable with shift timings (1 PM to 10 PM)?
Experience:
- Azure: 8 years (Required)
- Monitoring Tools: 5 years (Required)
- Open Telemetry (OTel): 5 years (Required)
- KQL: 5 years (Required)
Location:
- Bangalore City, Bengaluru, Karnataka (Required)
Work Location: In person
📌 Sr SRE Engineer (India)
🏢 Robotico Digital
📍 India