17 Sep
|
Robotico Digital
|
India
17 Sep
Robotico Digital
India
Key Responsibilities
Monitoring Framework Ownership
Own the design, implementation, and ongoing management of the Atlas5 monitoring framework across all 5 layers:
Layer 1: Infrastructure (VMs, compute, networking)
Layer 2: Data Tier (SQL DB, Cosmos DB, Elastic Pools, Microsoft Fabric)
Layer 3: Application (App Services, Function Apps, BRE)
Layer 4: Integration (EDA, Logic Apps, APIM)
Layer 5: File Exchange (FTP/SFTP, file drop, inbound feeds)
Define and configure alert thresholds, severity levels, and escalation paths for all layers
Ensure every alert triggers a Jira incident ticket and Dev Ops channel notification within defined SLA
Enforce mandatory RCA completion for all P1-Critical incidents before ticket closure
Tool Selection & Implementation
Evaluate, recommend, and implement the right monitoring toolset for each layer - Azure Monitor, Prometheus, Grafana, New Relic, Sig Noz, Zabbix, Dynatrace, or equivalent
Build a cost-effective hybrid monitoring stack that balances coverage, operational overhead, and cost vs. native Azure Monitor
Deploy and configure monitoring agents across all Atlas5 production infrastructure
Integrate monitoring tools with Jira Service Management for automated incident creation
Build and maintain unified monitoring dashboards for real-time 5-layer status visibility
SRE Practice & Incident Management
Define and maintain SLOs, SLIs, and error budgets for all Atlas5 production services
Lead incident response for P1-Critical production incidents - own the investigation, RCA, and permanent fix process
Drive the monitoring gap assessment as a mandatory step in every incident RCA
Configure auto-escalation for unacknowledged Critical alerts to Dev Ops Lead and secondary on-call
Conduct regular alert quality reviews - target false positive rate below 5%
Azure Observability Engineering
Configure and manage Azure Monitor, Log Analytics workspaces, Application Insights, and Diagnostic Settings
Design and manage KQL queries for alerting, dashboards, and operational reporting
Optimise monitoring costs - migrate Log Search Alerts to Metric Alerts (30x cheaper) and Activity Log Alerts (free) where applicable
Configure Prometheus scraping for Azure-native services and manage the Prometheus + Grafana stack
Instrument applications with Open Telemetry for vendor-neutral, future-proof telemetry
Reliability & Performance Engineering
Proactively identify reliability risks - throttling, capacity saturation, latency spikes, and dependency failures
Partner with Dev Ops and application teams to instrument new services and define monitoring parameters
Work with development team to define and implement EDA logging and integration layer monitoring parameters
Contribute to deployment reliability - flag monitoring gaps in pre-deployment validation and change management
Required Experience & Skills
Core Requirements
8 10 years of experience in SRE,
observability engineering, or cloud operations roles
Proven experience designing and implementing monitoring frameworks from scratch in Azure environments
Strong Azure Monitor expertise - Log Analytics, Application Insights, Metric Alerts, Log Search Alerts, Diagnostic Settings
Hands-on experience with at least 2 of: Prometheus, Grafana, New Relic, Datadog, Dynatrace, Zabbix, Sig Noz
Experience integrating monitoring with incident management tools (Jira, Pager Duty, or equivalent)
Technical Skills
Azure: Azure Monitor, Log Analytics (KQL), Application Insights, Azure Alerts, Diagnostic Settings, Service Health
Monitoring tools: Prometheus, Grafana, Current Relic, Dynatrace, Zabbix, Sig Noz - hands-on with at least 2
Instrumentation: Open Telemetry (OTel) - logs, metrics, traces; Application Insights SDK
Alerting: Alert rule design, severity taxonomy, escalation policies, on-call automation
Scripting: KQL (mandatory), Power Shell or Python for alert automation and tooling
Azure infrastructure: App Services, Function Apps, Cosmos DB, SQL DB, Microsoft Fabric, AKS - operational monitoring knowledge
Cost optimisation: experience reducing Azure Monitor costs through Metric Alert migration and Log tier management
Nice to Have
Experience in fintech or regulated financial services environments
Familiarity with EDA (Event-Driven Architecture) and integration layer monitoring (Service Bus, Logic Apps, APIM)
Experience with Sig Noz or Open Telemetry-native observability stacks
Dynatrace certification or hands-on Dynatrace deployment experience
📌 Senior SRE / Observability Engineer (India)
🏢 Robotico Digital
📍 India