Site Reliability Engineer (SRE) (Pune)

Site Reliability Engineer (SRE) (Pune)

07 Aug
|
EmbarkingOnVoyage Digital Solutions
|
Pune

07 Aug

EmbarkingOnVoyage Digital Solutions

Pune

Job Title: Site Reliability Engineer (SRE) – AI & Cloud Infrastructure

Location: Pune (Work From Office)

Experience: 5–8 Years

Employment Type: Full-Time

About the Role

We are looking for an experienced Site Reliability Engineer (SRE) to build and scale AI-powered reliability capabilities from the ground up. In this role, you will drive contemporary observability, automation, and cloud reliability initiatives while leveraging AI/ML for incident management, forecasting, and infrastructure optimization.

You will own the end-to-end reliability strategy across cloud-native AWS environments, enabling high availability, performance, and operational excellence through automation, intelligent monitoring, and proactive engineering.

Key Responsibilities

- Design, implement, and manage highly available, scalable, and secure cloud infrastructure on AWS.
- Build and maintain an end-to-end observability platform using Open Telemetry, Grafana, Datadog, CloudWatch, and related tools.
- Implement AIOps capabilities, including:
- LLM-assisted incident triage
- AI-powered root cause analysis
- ML-driven forecasting and anomaly detection
- Intelligent alert correlation and noise reduction
- Lead production incident management, on-call response, postmortems, and Root Cause Analysis (RCA).
- Automate operational workflows using Infrastructure as Code (Terraform/CloudFormation) and CI/CD pipelines.
- Drive infrastructure rightsizing, capacity planning, utilization analysis, and cloud cost optimization.
- Build dashboards, SLOs, SLIs, and error budgets to improve service reliability.
- Develop automation scripts using Python, Bash, or Go to eliminate manual operational tasks.
- Monitor application and infrastructure health while ensuring high uptime and service performance.
- Collaborate with Development, DevOps, Security, Platform Engineering, and Product teams across multiple time zones.
- Establish operational best practices for monitoring, incident response, disaster recovery, and resilience engineering.




- Maintain Linux-based production systems and troubleshoot OS, networking, storage, and performance issues.

Required Skills & Qualifications

- 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure Engineering.
- Strong experience with AWS services including EC2, ECS/EKS, Lambda, VPC, IAM, CloudWatch, RDS, Route 53, S3, and Auto Scaling.
- Hands-on experience with Open Telemetry, Grafana, Datadog, Prometheus, or similar monitoring platforms.
- Strong knowledge of Linux administration, networking, system performance tuning, and troubleshooting.
- Experience with Infrastructure as Code using Terraform or CloudFormation.
- Proficiency in scripting using Python, Bash, or Go.
- Experience with Kubernetes and containerized workloads.
- Strong understanding of CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.).
- Experience leading incident management, production support, and RCA processes.
- Knowledge of SRE principles including SLIs, SLOs, and Error Budgets.
- Experience implementing monitoring, logging, alerting, and observability frameworks.
- Strong analytical, troubleshooting, and communication skills.

Preferred Qualifications

- Experience building or implementing AIOps solutions.
- Exposure to Large Language Models (LLMs) for operational automation.
- Experience with machine learning-based forecasting or anomaly detection.
- Hands-on experience administering Adobe Experience Manager (AEM).
- Experience managing Cloudflare CDN, WAF, DNS, and caching strategies.
- Knowledge of FinOps, cloud cost optimization, and capacity planning.
- AWS Solutions Architect, DevOps Engineer, or Kubernetes certifications are a plus.

Skills:- AIOps, Large Language Models (LLM), Machine Learning (ML), Amazon Web Services (AWS), Cloud-Native Infrastructure, OpenTelemetry, grafana, Datadog, Responsive Design, Root Cause Analysis (RCA), Infrastructure Design & Management, Infrastructure Rightsizing, Capacity Planning, Cost Anomaly Detection, Linux administration, Systems Administration, Adobe Experience Manager (AEM) Administration, Cloudflare CDN and Cross-functional Collaboration

📌 Site Reliability Engineer (SRE) (Pune)
🏢 EmbarkingOnVoyage Digital Solutions
📍 Pune

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (sre) (pune) / pune

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (sre) (pune) / pune