29 Aug
|
CodeXray | Agentic SRE | Full-Stack Observability | API Security
|
Bengaluru
29 Aug
CodeXray | Agentic SRE | Full-Stack Observability | API Security
Bengaluru
Company Description
CodeXray delivers a modular, integrated suite of Full-Stack Observability, Agentic SRE, and API Security solutions that help engineering, DevOps, and security teams detect, understand, and resolve issues with confidence. Its eBPF-powered platform captures comprehensive MELT + P telemetry across applications, infrastructure, databases, and user experience, and supports modern and legacy stacks through OTEL SDKs, collectors, and custom agents. CodeXray’s agentic SRE capabilities enable automated anomaly investigation and root cause analysis, while its API Security Observability covers OWASP Top 10 risks, PII detection, active and passive testing, and risk scoring.
Built for both SaaS and on-premise deployments, the platform focuses on providing unified context and intelligence rather than a monolithic toolset. Recognized as Best Product in Agentic SRE & Full-Stack Observability Tools at CIO Conclave & Awards 2025, CodeXray is committed to helping teams run reliable, high-performing, and secure systems.
About the Role
We are seeking a highly motivated and resilient Associate Site Reliability Engineer (L1) to join our managed services team. In this role, you will act as the first line of defense, working from our or client location in Bengaluru.
You will combine a startup mindset—agility, fast learning, and ownership—with the operational rigor required to support critical infrastructure in a highly regulated industry. This role is a perfect launchpad for a junior engineer looking to master large-scale cloud operations, infrastructure-as-code, and enterprise-grade observability.
Core Responsibilities
1.
24x7 L1 Production Support & Incident Triage
- Actively monitor production environments in a 24x7 rotational shift setup to ensure maximum uptime and system reliability.
- Acknowledge, validate, and perform initial triage on automated alerts and system anomalies.
- Coordinate incident response, escalate severe bottlenecks to L2/L3 engineering teams with clear data points, and manage stakeholder communications.
2. System Health Checks & Observability
- Perform scheduled, rigorous system health checks across production, staging, and disaster recovery environments.
- Maintain, configure, and optimize automated alerts using enterprise observability platforms.
- Identify recurring noise or false-alarm alerts and work toward refining thresholds.
3. Infrastructure Provisioning & Maintenance
- Assist in provisioning, scaling, and managing cloud-native infrastructure using Infrastructure as Code (IaC) principles.
- Execute standard runbooks for deployment verification, patching, and system maintenance.
4. Blameless Post-Mortem & Documentation
- Participate in incident review meetings, contributing objective timelines and logs.
- Drive blameless post-mortem coordination to identify root causes and prevent incident re-occurrence.
- Standardize operational processes by updating and authoring technical runbooks and internal knowledge base documents.
Required Technical Skills
Operating Systems: Solid foundational command over Linux/Unix administration (navigation, system logs, process management, network configuration).
Scripting & Automation: Proficiency in writing basic automation scripts using Python, Bash, or Shell to eliminate repetitive manual work (Toil).
Cloud Infrastructure: Exposure to major cloud environments (AWS, Azure, or GCP)—understanding of EC2/VMs, VPCs, IAM, and managed services.
Infrastructure as Code: Familiarity with basic Terraform or cloud-native templates (CloudFormation/ARM) for infrastructure modification.
Observability Stack: Hands-on exposure to tools like Prometheus, Grafana, ELK stack (Elasticsearch, Logstash, Kibana), Datadog, or New Relic.
Behavioral & Operational Competencies
Communication: Exceptional verbal and written English communication skills; ability to convey highly technical incident details clearly to US-based teams.
Adaptability: Willingness and physical readiness to work in a continuous 24x7 rotational shift workplace.
Problem Solving: A structured, logical approach to debugging system anomalies under time constraints.
Mindset: An innate curiosity to understand why systems fail and a commitment to continuous improvement.
What We Offer
- Exposure to complex, regulated critical infrastructure parameters.
- Fast-tracked career progression into L2 SRE and core DevOps engineering roles.
- Shift allowances, comprehensive health insurance, and continuous upskilling certifications.
📌 Associate Site Reliability EngineerProduction Support) (Bengaluru)
🏢 CodeXray | Agentic SRE | Full-Stack Observability | API Security
📍 Bengaluru