Senior Principal Infrastructure Services (SRE Practice) (Karnataka)

Senior Principal Infrastructure Services (SRE Practice) (Karnataka)

02 Aug
|
Northern Trust
|
Karnataka

02 Aug

Northern Trust

Karnataka

This role will play a pivotal part in ensuring the reliability and performance of the company s systems and services. As a Site Reliability DevOps Engineer, you will be responsible for defining and deploying key observability services with a deep focus on architecture, production operations, capacity planning, performance management, deployment, and release engineering. You will work with cross-functional teams to assist with providing efficiency of our services. Your expertise in both software engineering and system operations will enable our partners to drive continuous improvements in our platform s reliability. This role will focus on bringing complete observability across all technologies.

This role will be responsible for a number of key functions that both support and drive improvements to the reliability of Northern Trust s IT Landscape.

What you will do

Reliability Focused System Design Architecture

- Lead the design and evolution of highly reliable, scalable, and performant distributed systems , applying SRE principles across infrastructure and application layers.
- Partner with engineering and architecture teams to influence system design decisions that improve resilience, fault tolerance, and operational simplicity .
- Define and promote reliability patterns, architectural best practices, and non functional requirements aligned with business criticality.

SRE Operations Automation

- Drive an automation first approach by designing and developing tools, scripts, and platforms that reduce manual effort, operational toil, and human error.
- Embed reliability engineering into the software delivery lifecycle through CI/CD integration, safe deployments, and repeatable operational workflows.
- Establish transparent operational metrics and service health indicators to ensure transparency and accountability.

Incident Management Root Cause Analysis

- Participate in and lead incident response for production systems, ensuring timely mitigation and minimal customer or business impact.
- Conduct and drive blameless post incident reviews , focusing on identifying systemic causes rather than individual faults.
- Implement long term corrective actions to prevent recurrence and measurably improve system reliability.

Monitoring, Alerting Observability

- Architect and implement end to end observability across systems using metrics, logs, and traces to enable rapid diagnosis and proactive issue detection.
- Define and manage Service Level Indicators (SLIs),



Service Level Objectives (SLOs), and error budgets to balance reliability with feature velocity.
- Build and maintain actionable dashboards and alerts that provide real time insights into system health, performance, and risk.

Continuous Reliability Improvement

- Identify reliability gaps through data analysis, failure reviews, and resilience testing, driving targeted improvement initiatives.
- Lead efforts such as capacity planning, load testing, chaos engineering, and fault injection to validate system behavior under stress.
- Continuously reduce operational toil, improve mean time to detect (MTTD) and mean time to recover (MTTR), and raise overall service maturity.

Documentation Knowledge Sharing

- Create and maintain clear, accurate, and actionable documentation including system architectures, runbooks, operational standards, and incident playbooks.
- Ensure documentation supports operational readiness, repeatability, and effective knowledge transfer across teams.

Cross Functional Collaboration Influence

- Work closely with product, development, platform, security, and operations teams to embed SRE principles into roadmap planning and delivery.
- Act as a trusted advisor, translating reliability data and operational risk into business relevant insights for technical and non technical stakeholders.
- Advocate for SRE best practices and help build a strong reliability culture across the organization.

Project Initiative Leadership

- Manage and prioritize multiple reliability focused initiatives, balancing short term operational needs with long term system health.
- Drive execution of strategic SRE programs that measurably improve system resilience, scalability, and operational efficiency.

Qualifications Experience

- Bachelor s degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience demonstrating advanced technical and leadership capabilities.

- 15+ years of progressive experience in systems engineering with a strong emphasis on site reliability, large scale systems operations, and software engineering in complex enterprise or cloud environments.





- 7+ years of experience in a technical leadership role (Team Lead or Hands on Technical Manager), with a proven track record of driving cross functional initiatives and delivering complex projects to successful completion.

- Strong proficiency in one or more modern programming languages such as Python, Go, Java, Ruby , or equivalent, with a software engineering mindset applied to operational challenges.

- Demonstrated experience operating and supporting systems across hybrid environments , including both on premises infrastructure and public/private cloud platforms .

- Hands on experience with containerization and container orchestration technologies , enabling scalable, resilient, and repeatable deployments.

- Proven ability to design and implement observability solutions , including metrics, logs, traces, dashboards, and alerts that provide actionable insights into system health and performance.

- Deep understanding of distributed systems, networking fundamentals, failure modes, and modern software architectures , with the ability to reason about complex system behaviors under load or failure conditions.

- Exceptional problem solving skills with the ability to diagnose, mitigate, and permanently resolve complex, high impact technical issues .

- Strong customer and stakeholder orientation, with excellent communication skills and the ability to articulate complex reliability strategies clearly and persuasively to both technical and non technical audiences.

- Prior experience designing and delivering Infrastructure as Code (IaC) through automated CI/CD pipelines , ensuring consistency, scalability, and reliability of infrastructure changes.

- Demonstrated success in mentoring, coaching, and developing high performing technical teams , fostering a culture of engineering excellence, ownership, and continuous improvement.

- Hands on expertise in implementing automated remediation and corrective actions driven by observability signals and reliability metrics.

- Practical experience working within Agile and DevOps environments , collaborating closely with product and engineering teams to balance reliability, velocity, and innovation.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Senior Principal Infrastructure Services (SRE Practice) (Karnataka)
🏢 Northern Trust
📍 Karnataka

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior principal infrastructure services (sre practice) (karnataka) / karnataka

Subscribe to this job alert:

Get the latest job offers by email for: senior principal infrastructure services (sre practice) (karnataka) / karnataka