Senior SRE + AI / GenAI Engineer — Detailed Job Description
Senior SRE + Ai
Experience: 6–10 Years
Location: Pan India
Notice Period: Immediate to 30 Days
Employment Type: Full time
Work Mode: Hybrid / Remote / On-site, depending on business requirements
Position Overview
We are looking for a highly skilled and experienced Senior SRE + AI / GenAI Engineer with 6–10 years of hands-on experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Automation, and modern AI/Generative AI technologies.
The candidate will be responsible for ensuring the availability, scalability, reliability, performance, and security of mission-critical applications and infrastructure, while also leveraging AI/GenAI to automate and improve engineering and operational processes.
The ideal candidate should have a strong background in SRE/DevOps and cloud-native technologies, combined with practical experience in AI/ML, Generative AI, LLMs, RAG, AI-powered automation, or AIOps.
This role requires someone who can work across infrastructure, application, observability, automation, and AI layers and can independently troubleshoot complex production issues.
Key Responsibilities
1.
Site Reliability Engineering
- Own and improve the reliability, availability, scalability, and performance of production applications and platforms.
- Define and implement SLIs, SLOs, SLAs, and error budgets for critical services.
- Establish reliability standards and operational best practices across engineering teams.
- Monitor production environments and proactively identify reliability and performance risks.
- Participate in production support and on-call activities.
- Perform detailed troubleshooting of production incidents.
- Conduct Root Cause Analysis (RCA) for major incidents and drive permanent corrective actions.
- Track recurring incidents and identify opportunities for automation and reliability improvements.
- Reduce operational toil through automation and self-service capabilities.
- Develop and maintain operational runbooks, troubleshooting guides, and knowledge documentation.
- Conduct capacity planning and scalability assessments.
- Identify system bottlenecks and implement performance improvements.
📌 Site Reliability Engineer (Bengaluru)
🏢 .
📍 Bengaluru