09 Sep
|
CodeXray | Agentic SRE | Full-Stack Observability | API Security
|
Bengaluru
09 Sep
CodeXray | Agentic SRE | Full-Stack Observability | API Security
Bengaluru
Senior Site Reliability Engineer (SRE)
Experience: 4+ Years
Role Type: Full-time
Deployment: Client Location (Bengaluru) / Client Production Setting
Work Schedule: Shift-based, primarily aligned to US time zones
Who We Are
CodeXray is building a modern observability and reliability platform for complex, high-volume production environments.
We work across observability, cloud infrastructure, distributed systems, automation, and AI-assisted SRE, helping enterprises improve production visibility, troubleshooting, incident response, and reliability.
For this role, you will be deployed as part of the CodeXray SRE team at a client environment, working closely with client engineering and operations teams on business-critical production systems.
The Role
We are looking for a hands-on Senior SRE with strong production ownership, troubleshooting depth, and the ability to operate confidently in high-volume AWS environments.
This is not a monitoring or ticket-closure role.
You will be expected to lead critical incidents, run war rooms, troubleshoot across multiple technology layers, guide L1 SREs, and drive issues through resolution, RCA, and preventive action.
What You Will Own
- Lead critical production incidents and war rooms
- Troubleshoot issues across AWS, Kubernetes, Linux, applications, networks, databases, and dependent services
- Own incidents from detection through mitigation, resolution, RCA, and follow-up actions
- Lead and guide a team of L1 SREs / Operations Engineers
- Work closely with client application, infrastructure, cloud, database, and engineering teams
- Correlate metrics, logs, traces, and infrastructure signals to identify root causes
- Improve monitoring, alerting, runbooks, SOPs, escalation paths, and operational readiness
- Ensure clear communication and structured handovers across shifts
- Identify recurring issues and drive corrective and preventive actions
What We Are Looking For
- 4+ years of experience in SRE, Production Engineering, DevOps, Cloud Operations, or similar roles
- Strong hands-on troubleshooting experience in AWS production environments
- Experience supporting high-volume, distributed, business-critical systems
- Robust knowledge of Linux, networking, containers, and Kubernetes
- Good understanding of DNS, TCP/IP, HTTP/HTTPS, load balancers, proxies, and connectivity troubleshooting
- Experience with observability tools such as CloudWatch, Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, or similar
- Strong understanding of incident management, escalation, RCA, and problem management
- Ability to troubleshoot across multiple layers rather than depend only on dashboards or runbooks
- Strong communication skills and confidence working directly with client stakeholders
What Matters to Us
We value engineers who are:
- High on intent, ownership, and accountability
- Highly sensitive to production impact
- Comfortable taking charge during critical incidents
- Able to establish a troubleshooting path when the root cause is unclear
- Evidence-driven and methodical under pressure
- Strong enough technically to guide L1 engineers while remaining hands-on
- Comfortable working in a client-facing environment
What You Will Gain
This role provides exposure to a large-scale enterprise production environment where reliability and operational discipline genuinely matter.
You will get the opportunity to:
- Work hands-on with high-volume AWS production systems
- Build deep expertise in SRE, cloud troubleshooting, Kubernetes, and observability
- Lead real production incidents and complex war rooms
- Work closely with senior client engineering and operations teams
- Gain exposure to enterprise-scale operational processes and reliability practices
- Work with CodeXray’s observability and AI-assisted RCA capabilities
- Grow toward broader SRE leadership, production engineering, or reliability architecture roles
Shift & Client Location Requirement
This is a client-location deployment and candidates must be comfortable working from the client office in Bengaluru as required.
The role is shift-based and will have significant alignment to US working hours.
Candidates must be comfortable with evening/night shifts in India, rotational schedules, and critical incident support as required.
Willingness to work from the client location and primarily in US-aligned shifts is mandatory.
Good to Have
Experience with AWS services such as EC2, EKS, RDS, ALB/NLB, Route 53, IAM, S3, and CloudWatch, along with exposure to Kafka, Redis, PostgreSQL/MySQL, Terraform, CI/CD, or other distributed systems.
How We Think About SRE
Understand impact → establish the troubleshooting path → lead the war room → correlate signals → drive resolution → complete RCA → prevent recurrence.
- If you enjoy being close to production and taking ownership of difficult operational problems, we would like to speak with you.
📌 Senior Site Reliability EngineerProduction Support) (Bengaluru)
🏢 CodeXray | Agentic SRE | Full-Stack Observability | API Security
📍 Bengaluru