Senior Site Reliability EngineerProduction Support) (Bengaluru)

Senior Site Reliability EngineerProduction Support) (Bengaluru)

09 Sep
|
CodeXray | Agentic SRE | Full-Stack Observability | API Security
|
Bengaluru

09 Sep

CodeXray | Agentic SRE | Full-Stack Observability | API Security

Bengaluru

Senior Site Reliability Engineer (SRE)

Experience: 4+ Years

Role Type: Full-time

Deployment: Client Location (Bengaluru) / Client Production Setting

Work Schedule: Shift-based, primarily aligned to US time zones

Who We Are

CodeXray is building a modern observability and reliability platform for complex, high-volume production environments.

We work across observability, cloud infrastructure, distributed systems, automation, and AI-assisted SRE, helping enterprises improve production visibility, troubleshooting, incident response, and reliability.

For this role, you will be deployed as part of the CodeXray SRE team at a client environment, working closely with client engineering and operations teams on business-critical production systems.

The Role

We are looking for a hands-on Senior SRE with strong production ownership, troubleshooting depth, and the ability to operate confidently in high-volume AWS environments.

This is not a monitoring or ticket-closure role.

You will be expected to lead critical incidents, run war rooms, troubleshoot across multiple technology layers, guide L1 SREs, and drive issues through resolution, RCA, and preventive action.

What You Will Own

- Lead critical production incidents and war rooms
- Troubleshoot issues across AWS, Kubernetes, Linux, applications, networks, databases, and dependent services
- Own incidents from detection through mitigation, resolution, RCA, and follow-up actions
- Lead and guide a team of L1 SREs / Operations Engineers
- Work closely with client application, infrastructure, cloud, database, and engineering teams
- Correlate metrics, logs, traces, and infrastructure signals to identify root causes
- Improve monitoring, alerting, runbooks, SOPs, escalation paths, and operational readiness




- Ensure clear communication and structured handovers across shifts
- Identify recurring issues and drive corrective and preventive actions

What We Are Looking For
- 4+ years of experience in SRE, Production Engineering, DevOps, Cloud Operations, or similar roles
- Strong hands-on troubleshooting experience in AWS production environments
- Experience supporting high-volume, distributed, business-critical systems
- Robust knowledge of Linux, networking, containers, and Kubernetes
- Good understanding of DNS, TCP/IP, HTTP/HTTPS, load balancers, proxies, and connectivity troubleshooting
- Experience with observability tools such as CloudWatch, Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, or similar
- Strong understanding of incident management, escalation, RCA, and problem management
- Ability to troubleshoot across multiple layers rather than depend only on dashboards or runbooks
- Strong communication skills and confidence working directly with client stakeholders

What Matters to Us

We value engineers who are:

- High on intent, ownership, and accountability
- Highly sensitive to production impact
- Comfortable taking charge during critical incidents
- Able to establish a troubleshooting path when the root cause is unclear
- Evidence-driven and methodical under pressure
- Strong enough technically to guide L1 engineers while remaining hands-on
- Comfortable working in a client-facing environment

What You Will Gain





This role provides exposure to a large-scale enterprise production environment where reliability and operational discipline genuinely matter.

You will get the opportunity to:

- Work hands-on with high-volume AWS production systems
- Build deep expertise in SRE, cloud troubleshooting, Kubernetes, and observability
- Lead real production incidents and complex war rooms
- Work closely with senior client engineering and operations teams
- Gain exposure to enterprise-scale operational processes and reliability practices
- Work with CodeXray’s observability and AI-assisted RCA capabilities
- Grow toward broader SRE leadership, production engineering, or reliability architecture roles

Shift & Client Location Requirement

This is a client-location deployment and candidates must be comfortable working from the client office in Bengaluru as required.

The role is shift-based and will have significant alignment to US working hours.

Candidates must be comfortable with evening/night shifts in India, rotational schedules, and critical incident support as required.

Willingness to work from the client location and primarily in US-aligned shifts is mandatory.

Good to Have

Experience with AWS services such as EC2, EKS, RDS, ALB/NLB, Route 53, IAM, S3, and CloudWatch, along with exposure to Kafka, Redis, PostgreSQL/MySQL, Terraform, CI/CD, or other distributed systems.

How We Think About SRE

Understand impact → establish the troubleshooting path → lead the war room → correlate signals → drive resolution → complete RCA → prevent recurrence.

- If you enjoy being close to production and taking ownership of difficult operational problems, we would like to speak with you.

📌 Senior Site Reliability EngineerProduction Support) (Bengaluru)
🏢 CodeXray | Agentic SRE | Full-Stack Observability | API Security
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineerproduction support) (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineerproduction support) (bengaluru) / bengaluru