About the Company:
Eton Solutions is a hypergrowth fintech transforming the Family Office segment of the Wealth Management industry. Eton’s AtlasFive® is a comprehensive enterprise management platform specifically designed to allow today’s modern Family Office meet the unique and varied challenges of Ultra High Net Worth families.
For more details visit: https://eton-solutions.com/
Senior SRE / Observability Engineer
Role Summary
We are seeking a Senior Site Reliability Engineer (SRE) with deep expertise in observability, monitoring framework design, and Azure platform operations. This is a hands-on technical leadership role responsible for owning and driving the end-to-end monitoring framework for all Atlas5 production environments. The ideal candidate has a strong background in monitoring tool evaluation and implementation, alert engineering, incident response, and SRE best practices, and can operate independently to design, deploy, and continuously improve the observability stack across a complex, multi-layer cloud architecture.
Key Responsibilities
Monitoring Framework Ownership
- Own the design, implementation, and ongoing management of the Atlas5 monitoring framework across all 5 layers:
- Layer 1: Infrastructure (VMs, compute, networking)
- Layer 2: Data Tier (SQL DB, Cosmos DB, Elastic Pools, Microsoft Fabric)
- Layer 3: Application (App Services, Function Apps, BRE)
- Layer 4: Integration (EDA, Logic Apps, APIM)
- Layer 5: File Exchange (FTP/SFTP, file drop, inbound feeds)
- Define and configure alert thresholds, severity levels, and escalation paths for all layers
- Ensure every alert triggers a Jira incident ticket and DevOps channel notification within defined SLA
- Enforce mandatory RCA completion for all P1-Critical incidents before ticket closure
Tool Selection & Implementation
- Evaluate, recommend, and implement the right monitoring toolset for each layer — Azure Monitor, Prometheus, Grafana, New Relic, SigNoz, Zabbix, Dynatrace, or equivalent
- Build a cost-effective hybrid monitoring stack that balances coverage, operational overhead, and cost vs. native Azure Monitor
- Deploy and configure monitoring agents across all Atlas5 production infrastructure
- Integrate monitoring tools with Jira Service Management for automated incident creation
- Build and maintain unified monitoring dashboards for real-time 5-layer status visibility
SRE Practice & Incident Management
- Define and maintain SLOs, SLIs, and error budgets for all Atlas5 production services
- Lead incident response for P1-Critical production incidents — own the investigation, RCA, and permanent fix process
- Drive the monitoring gap assessment as a mandatory step in every incident RCA
- Configure auto-escalation for unacknowledged Critical alerts to DevOps Lead and secondary on-call
- Conduct regular alert quality reviews — target false positive rate below 5%
Azure Observability Engineering
- Configure and manage Azure Monitor, Log Analytics workspaces, Application Insights, and Diagnostic Settings
- Design and manage KQL queries for alerting, dashboards, and operational reporting
- Optimise monitoring costs — migrate Log Search Alerts to Metric Alerts (30x cheaper) and Activity Log Alerts (free) where applicable
- Configure Prometheus scraping for Azure-native services and manage the Prometheus + Grafana stack
- Instrument applications with Open Telemetry for vendor-neutral, future-proof telemetry
Reliability & Performance Engineering
- Proactively identify reliability risks — throttling, capacity saturation, latency spikes, and dependency failures
- Partner with DevOps and application teams to instrument new services and define monitoring parameters
- Work with development team to define and implement EDA logging and integration layer monitoring parameters
- Contribute to deployment reliability — flag monitoring gaps in pre-deployment validation and change management
Required Experience & Skills
Core Requirements
- 8 – 10 years of experience in SRE, observability engineering, or cloud operations roles
- Proven experience designing and implementing monitoring frameworks from scratch in Azure environments
- Strong Azure Monitor expertise — Log Analytics, Application Insights, Metric Alerts, Log Search Alerts, Diagnostic Settings
- Hands-on experience with at least 2 of: Prometheus, Grafana, Recent Relic, Datadog, Dynatrace, Zabbix, SigNoz
- Experience integrating monitoring with incident management tools (Jira, PagerDuty, or equivalent)
Technical Skills
- Azure: Azure Monitor, Log Analytics (KQL), Application Insights, Azure Alerts, Diagnostic Settings, Service Health
- Monitoring tools: Prometheus, Grafana, New Relic, Dynatrace, Zabbix, SigNoz — hands-on with at least 2
- Instrumentation: Open Telemetry (OTel) — logs, metrics, traces; Application Insights SDK
- Alerting: Alert rule design, severity taxonomy, escalation policies, on-call automation
- Scripting: KQL (mandatory), PowerShell or Python for alert automation and tooling
- Azure infrastructure: App Services, Function Apps, Cosmos DB, SQL DB, Microsoft Fabric, AKS — operational monitoring knowledge
- Cost optimisation: experience reducing Azure Monitor costs through Metric Alert migration and Log tier management
Nice to Have
- Experience in fintech or regulated financial services environments
- Familiarity with EDA (Event-Driven Architecture) and integration layer monitoring (Service Bus, Logic Apps, APIM)
- Experience with SigNoz or Open Telemetry-native observability stacks
- Dynatrace certification or hands-on Dynatrace deployment experience
if interested, share your updated resume to
[email protected]
📌 Site Reliability Engineer (Bengaluru)
🏢 Eton Solutions
📍 Bengaluru