DevOps SRE-AI
Experience:
5–19 Years
Locations:
Pune, Mumbai, Chennai, Hyderabad, Bengaluru, Noida, Kolkata
Notice Period:
Immediate Joiners to 30 Days Only
Work Mode:
Hybrid
Employment Type:
Full-time
Job Description
We are looking for an experienced
Site Reliability Engineer (SRE) / Platform Engineer / DevOps Engineer
with solid hands-on experience in SRE, Platform Engineering, DevOps, Cloud, Kubernetes, Infrastructure as Code, and GenAI.
The candidate will be responsible for building, scaling, securing, and operating highly reliable and automated platform services. The role will focus on improving system reliability, reducing operational toil, enabling developer productivity, and contributing to platform engineering initiatives, technology accelerators, internal products, and Centers of Excellence (COEs).
Key Responsibilities
- Build and operate highly available, scalable, secure, and reliable platform services.
- Implement SRE practices including
SLIs, SLOs, SLAs, error budgets, and reliability engineering
.
- Manage, deploy,
and scale Kubernetes-based workloads across cloud environments.
- Develop and maintain CI/CD pipelines for application and infrastructure deployments.
- Design, implement, and maintain Infrastructure as Code using
Terraform
and related tools.
- Own production reliability and participate in on-call activities, incident response, troubleshooting, and Root Cause Analysis (RCA).
- Enhance observability using metrics, logs, traces, monitoring, dashboards, and alerting solutions.
- Identify and eliminate operational toil through automation and engineering best practices.
- Collaborate with development, DevOps, cloud, security, and infrastructure teams to improve platform reliability and availability.
- Develop automation scripts and tools using
Python, Go, Bash
, or similar programming languages.
- Support cloud-native application deployments and platform services across
AWS, Azure, and GCP
environments.
- Contribute to platform enginee
📌 DevOps SRE-AI (India)
🏢 Harp
📍 India