09 Sep
|
Moodys Analytics
|
Bengaluru
09 Sep
Moodys Analytics
Bengaluru
Job purpose
We are seeking a Site Reliability Engineer to safeguard the reliability, observability and resilience of critical platforms supporting a Global Capital Markets platform-modernization program. The role ensures that Front
Office Trading systems and the emerging target-state services spanning event-driven integration, data-lake /
Databricks platforms, reporting, migration and the front-to-back trade lifecycle for FX Cash, Cleared IRS and future products run reliably across the trading day and its critical batch cycles. The ideal candidate blends strong engineering with an operations mindset, automating toil, engineering for resilience, and improving supportability before services reach production.
Key responsibilities
Define and implement service-health metrics, monitoring, alerting and dashboards to provide end-to-end observability.
Support incident response, problem management and operational readiness, including root-cause analysis and post-incident reviews.
Identify and remediate reliability, performance and resilience risks across critical trading and post trade services.
Monitor batch / EOD cycles and time-critical processes, detecting failures, delays and job / queue issues early.
Engineer automation to reduce operational toil (self-healing, runbooks, scheduling) and support capacity planning and performance tuning and improve MTTR.
Partner with engineering teams to improve supportability, deployment safety and production readiness before release.
Contribute to SLIs / SLOs, error-budget management and continuous improvement of operational resilience and BCP.
Support Kubernetes-based platforms including monitoring cluster health, workload performance and platform reliability.
Key competencies
Bachelors / Masters degree with 8+ years in SRE / production-engineering / DevOps roles supporting business-critical, high-availability systems.
Strong observability skills monitoring, alerting and dashboards using tools such as Prometheus/Grafana, ELK, Splunk, Datadog, Dynatrace, AppDynamics or Cloud-native monitoring platforms.
lines
Incident and problem management experience, including on-call, escalation and RCA; scheduling tools (e.g., Autosys) and alerting (e.g., PagerDuty).
Automation and scripting (Python / Bash / PowerShell), with CI/CD and infrastructure-as-code familiarity.
Cloud experience (AWS and/or Azure), strong Linux fundamentals, and understanding of resilience, performance and capacity engineering.
Solid analytical, communication and cross-team collaboration skills, with attention to detail.
Understanding of disaster recovery (DR), high availability (HA), failover mechanisms and resilience engineering practices.
Preferred qualifications
Experience defining and operating against SLIs / SLOs and error budgets.
Exposure to containers / orchestration (Docker / Kubernetes) and streaming / event-driven platforms (Kafka / MQ).
Prior experience in Capital Markets or financial services, including awareness of settlement and stress-test batch processing.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Delivery Lead-SRE (Bengaluru)
🏢 Moodys Analytics
📍 Bengaluru