09 Sep
|
Acuity Knowledge Partners
|
Bengaluru
09 Sep
Acuity Knowledge Partners
Bengaluru
Job purpose
We are seeking a Site Reliability Engineer to safeguard the reliability, observability and resilience of critical platforms supporting a Global Capital Markets platform-modernization program. The role ensures that Front
Office Trading systems and the emerging target-state services — spanning event-driven integration, data-lake /
Databricks platforms, reporting, migration and the front-to-back trade lifecycle for FX Cash, Cleared IRS and future products — run reliably across the trading day and its critical batch cycles. The ideal candidate blends strong engineering with an operations mindset, automating toil, engineering for resilience, and improving supportability before services reach production.
Key responsibilities
- Define and implement service-health metrics, monitoring, alerting and dashboards to provide end-to-end observability.
- Support incident response, problem management and operational readiness, including root-cause analysis and post-incident reviews.
- Identify and remediate reliability, performance and resilience risks across critical trading and post trade services.
- Monitor batch / EOD cycles and time-critical processes, detecting failures, delays and job / queue issues early.
- Engineer automation to reduce operational toil (self-healing, runbooks, scheduling) and support capacity planning and performance tuning & and improve MTTR.
- Partner with engineering teams to improve supportability, deployment safety and production readiness before release.
- Contribute to SLIs / SLOs, error-budget management and continuous improvement of operational resilience and BCP.
- Support Kubernetes-based platforms including monitoring cluster health, workload performance and platform reliability.
Key competencies
- Bachelor’s / Master’s degree with 8+ years in SRE / production-engineering / DevOps roles supporting business-critical, high-availability systems.
- Solid observability skills — monitoring, alerting and dashboards using tools such as Prometheus/Grafana, ELK, Splunk, Datadog, Dynatrace, AppDynamics or Cloud-native monitoring platforms.
- lines
- Incident and problem management experience, including on-call, escalation and RCA; scheduling tools (e.g., Autosys) and alerting (e.g., PagerDuty).
- Automation and scripting (Python / Bash / PowerShell), with CI/CD and infrastructure-as-code familiarity.
- Cloud experience (AWS and/or Azure), strong Linux fundamentals, and understanding of resilience, performance and capacity engineering.
- Strong analytical, communication and cross-team collaboration skills, with attention to detail.
- Understanding of disaster recovery (DR), high availability (HA), failover mechanisms and resilience engineering practices.
Preferred qualifications
- Experience defining and operating against SLIs / SLOs and error budgets.
- Exposure to containers / orchestration (Docker / Kubernetes) and streaming / event-driven platforms (Kafka / MQ).
- Prior experience in Capital Markets or financial services, including awareness of settlement and stress-test batch processing.
Job Snapshot
Job ID
JOB_8384
Department
Data and Technology Services
Location
ETV -SEZ1, Bengaluru, Karnataka, India
Experience
6 - 9 Years
Employee Type
Permanent
📌 Delivery Lead-SRE (Bengaluru)
🏢 Acuity Knowledge Partners
📍 Bengaluru