DevOps SRE-AI (India)

DevOps SRE-AI (India)

14 Aug
|
Harp
|
India

14 Aug

Harp

India

DevOps SRE-AI

Experience: 5–19 Years

Locations: Pune, Mumbai, Chennai, Hyderabad, Bengaluru, Noida, Kolkata

Notice Period: Immediate Joiners to 30 Days Only

Work Mode: Hybrid

Employment Type: Full-time

We are looking for an experienced Site Reliability Engineer (SRE) / Platform Engineer / DevOps Engineer with strong hands-on experience in SRE, Platform Engineering, DevOps, Cloud, Kubernetes, Infrastructure as Code, and GenAI.

The candidate will be responsible for building, scaling, securing, and operating highly reliable and automated platform services. The role will focus on improving system reliability, reducing operational toil, enabling developer productivity, and contributing to platform engineering initiatives, technology accelerators, internal products, and Centers of Excellence (COEs).

Key Responsibilities

- Build and operate highly available, scalable, secure, and reliable platform services.
- Implement SRE practices including SLIs, SLOs, SLAs, error budgets, and reliability engineering .
- Manage, deploy, and scale Kubernetes-based workloads across cloud environments.
- Develop and maintain CI/CD pipelines for application and infrastructure deployments.
- Design, implement, and maintain Infrastructure as Code using Terraform and related tools.
- Own production reliability and participate in on-call activities, incident response, troubleshooting, and Root Cause Analysis (RCA).
- Enhance observability using metrics, logs, traces, monitoring, dashboards, and alerting solutions.
- Identify and eliminate operational toil through automation and engineering best practices.
- Collaborate with development, DevOps, cloud, security,



and infrastructure teams to improve platform reliability and availability.
- Develop automation scripts and tools using Python, Go, Bash , or similar programming languages.
- Support cloud-native application deployments and platform services across AWS, Azure, and GCP environments.
- Contribute to platform engineering initiatives, internal COE projects, intellectual properties, and technology accelerators.
- Support development of internal platforms and products based on business and technical requirements.
- Monitor system health, capacity, availability, performance, and reliability of production environments.
- Conduct incident analysis, troubleshooting, RCA, and implement preventive and corrective actions.
- Continuously improve deployment, monitoring, reliability, and operational processes.
- Collaborate with cross-functional teams to define and implement reliability and automation standards.

Mandatory Skills
- 5–16 years of relevant experience in SRE, Platform Engineering, DevOps, Cloud Engineering, or AI Engineering .
- Strong hands-on experience with Kubernetes and Docker .
- Strong knowledge of Linux administration and troubleshooting .
- Robust scripting/programming skills in Python, Go, Bash , or similar languages.




- Hands-on experience with at least one major cloud platform: AWS, Azure, or GCP .
- Strong understanding of SRE principles including SLIs, SLOs, SLAs, error budgets, reliability, and availability .
- Experience developing and implementing CI/CD pipelines .
- Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools .
- Experience with monitoring, observability, logging, metrics, tracing, and alerting tools.
- Experience in production support, incident management, troubleshooting, and Root Cause Analysis.
- Strong understanding of cloud-native architecture, automation, scalability, and high availability.
- Strong analytical, troubleshooting, and problem-solving skills.

Good to Have
- Experience with Internal Developer Platforms (IDP) and platform engineering frameworks.
- Exposure to service mesh technologies such as Istio or Linkerd .
- Knowledge of FinOps, cloud cost optimization, and resource management .
- Experience with Jenkins, Git, GitLab, Azure DevOps , or similar CI/CD tools.
- Exposure to Prometheus, Grafana, ELK/EFK, OpenTelemetry , or similar observability platforms.
- Experience with AWS EKS, Azure AKS, or Google GKE .
- Experience with Helm, ArgoCD, GitOps , or similar cloud-native deployment technologies.
- Exposure to GenAI, LLMs, AI-powered DevOps/SRE tools, or AI-based operational automation .
- Experience developing AI/GenAI-based automation or intelligent SRE solutions.
- Cloud or Kubernetes certifications are an advantage.

Education Bachelor's/Master's degree in Computer Science, Information Technology, Engineering , or a related field.

📌 DevOps SRE-AI (India)
🏢 Harp
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: devops sre-ai (india) / india