Site Reliability Engineer (Bengaluru)

Site Reliability Engineer (Bengaluru)

18 Sep
|
Eton Solutions
|
Bengaluru

18 Sep

Eton Solutions

Bengaluru

About the Company:

Eton Solutions is a hypergrowth fintech transforming the Family Office segment of the Wealth Management industry. Eton’s AtlasFive® is a comprehensive enterprise management platform specifically designed to allow today’s modern Family Office meet the unique and varied challenges of Ultra High Net Worth families.

For more details visit: https://eton-solutions.com/

Senior SRE / Observability Engineer

Role Summary

We are seeking a Senior Site Reliability Engineer (SRE) with deep expertise in observability, monitoring framework design, and Azure platform operations. This is a hands-on technical leadership role responsible for owning and driving the end-to-end monitoring framework for all Atlas5 production environments. The ideal candidate has a strong background in monitoring tool evaluation and implementation, alert engineering, incident response, and SRE best practices, and can operate independently to design, deploy, and continuously improve the observability stack across a complex, multi-layer cloud architecture.

Key Responsibilities

Monitoring Framework Ownership

- Own the design, implementation, and ongoing management of the Atlas5 monitoring framework across all 5 layers:
- Layer 1: Infrastructure (VMs, compute, networking)
- Layer 2: Data Tier (SQL DB, Cosmos DB, Elastic Pools, Microsoft Fabric)
- Layer 3: Application (App Services, Function Apps, BRE)
- Layer 4: Integration (EDA, Logic Apps, APIM)
- Layer 5: File Exchange (FTP/SFTP, file drop, inbound feeds)
- Define and configure alert thresholds, severity levels, and escalation paths for all layers
- Ensure every alert triggers a Jira incident ticket and DevOps channel notification within defined SLA
- Enforce mandatory RCA completion for all P1-Critical incidents before ticket closure

Tool Selection & Implementation

- Evaluate, recommend, and implement the right monitoring toolset for each layer — Azure Monitor, Prometheus, Grafana, New Relic, SigNoz, Zabbix, Dynatrace, or equivalent




- Build a cost-effective hybrid monitoring stack that balances coverage, operational overhead, and cost vs. native Azure Monitor
- Deploy and configure monitoring agents across all Atlas5 production infrastructure
- Integrate monitoring tools with Jira Service Management for automated incident creation
- Build and maintain unified monitoring dashboards for real-time 5-layer status visibility

SRE Practice & Incident Management

- Define and maintain SLOs, SLIs, and error budgets for all Atlas5 production services
- Lead incident response for P1-Critical production incidents — own the investigation, RCA, and permanent fix process
- Drive the monitoring gap assessment as a mandatory step in every incident RCA
- Configure auto-escalation for unacknowledged Critical alerts to DevOps Lead and secondary on-call
- Conduct regular alert quality reviews — target false positive rate below 5%

Azure Observability Engineering

- Configure and manage Azure Monitor, Log Analytics workspaces, Application Insights, and Diagnostic Settings
- Design and manage KQL queries for alerting, dashboards, and operational reporting
- Optimise monitoring costs — migrate Log Search Alerts to Metric Alerts (30x cheaper) and Activity Log Alerts (free) where applicable
- Configure Prometheus scraping for Azure-native services and manage the Prometheus + Grafana stack
- Instrument applications with Open Telemetry for vendor-neutral, future-proof telemetry

Reliability & Performance Engineering

- Proactively identify reliability risks — throttling, capacity saturation, latency spikes, and dependency failures
- Partner with DevOps and application teams to instrument new services and define monitoring parameters




- Work with development team to define and implement EDA logging and integration layer monitoring parameters
- Contribute to deployment reliability — flag monitoring gaps in pre-deployment validation and change management

Required Experience & Skills

Core Requirements

- 8 – 10 years of experience in SRE, observability engineering, or cloud operations roles
- Proven experience designing and implementing monitoring frameworks from scratch in Azure environments
- Strong Azure Monitor expertise — Log Analytics, Application Insights, Metric Alerts, Log Search Alerts, Diagnostic Settings
- Hands-on experience with at least 2 of: Prometheus, Grafana, Recent Relic, Datadog, Dynatrace, Zabbix, SigNoz
- Experience integrating monitoring with incident management tools (Jira, PagerDuty, or equivalent)

Technical Skills

- Azure: Azure Monitor, Log Analytics (KQL), Application Insights, Azure Alerts, Diagnostic Settings, Service Health
- Monitoring tools: Prometheus, Grafana, New Relic, Dynatrace, Zabbix, SigNoz — hands-on with at least 2
- Instrumentation: Open Telemetry (OTel) — logs, metrics, traces; Application Insights SDK
- Alerting: Alert rule design, severity taxonomy, escalation policies, on-call automation
- Scripting: KQL (mandatory), PowerShell or Python for alert automation and tooling
- Azure infrastructure: App Services, Function Apps, Cosmos DB, SQL DB, Microsoft Fabric, AKS — operational monitoring knowledge
- Cost optimisation: experience reducing Azure Monitor costs through Metric Alert migration and Log tier management

Nice to Have

- Experience in fintech or regulated financial services environments
- Familiarity with EDA (Event-Driven Architecture) and integration layer monitoring (Service Bus, Logic Apps, APIM)
- Experience with SigNoz or Open Telemetry-native observability stacks
- Dynatrace certification or hands-on Dynatrace deployment experience

if interested, share your updated resume to [email protected]

📌 Site Reliability Engineer (Bengaluru)
🏢 Eton Solutions
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (bengaluru) / bengaluru