07 Aug
|
EmbarkingOnVoyage Digital Solutions
|
Pune
07 Aug
EmbarkingOnVoyage Digital Solutions
Pune
Job Title: Site Reliability Engineer (SRE) – AI & Cloud Infrastructure
Location: Pune (Work From Office)
Experience: 5–8 Years
Employment Type: Full-Time
About the Role
We are looking for an experienced Site Reliability Engineer (SRE) to build and scale AI-powered reliability capabilities from the ground up. In this role, you will drive contemporary observability, automation, and cloud reliability initiatives while leveraging AI/ML for incident management, forecasting, and infrastructure optimization.
You will own the end-to-end reliability strategy across cloud-native AWS environments, enabling high availability, performance, and operational excellence through automation, intelligent monitoring, and proactive engineering.
Key Responsibilities
- Design, implement, and manage highly available, scalable, and secure cloud infrastructure on AWS.
- Build and maintain an end-to-end observability platform using Open Telemetry, Grafana, Datadog, CloudWatch, and related tools.
- Implement AIOps capabilities, including:
- LLM-assisted incident triage
- AI-powered root cause analysis
- ML-driven forecasting and anomaly detection
- Intelligent alert correlation and noise reduction
- Lead production incident management, on-call response, postmortems, and Root Cause Analysis (RCA).
- Automate operational workflows using Infrastructure as Code (Terraform/CloudFormation) and CI/CD pipelines.
- Drive infrastructure rightsizing, capacity planning, utilization analysis, and cloud cost optimization.
- Build dashboards, SLOs, SLIs, and error budgets to improve service reliability.
- Develop automation scripts using Python, Bash, or Go to eliminate manual operational tasks.
- Monitor application and infrastructure health while ensuring high uptime and service performance.
- Collaborate with Development, DevOps, Security, Platform Engineering, and Product teams across multiple time zones.
- Establish operational best practices for monitoring, incident response, disaster recovery, and resilience engineering.
- Maintain Linux-based production systems and troubleshoot OS, networking, storage, and performance issues.
Required Skills & Qualifications
- 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure Engineering.
- Strong experience with AWS services including EC2, ECS/EKS, Lambda, VPC, IAM, CloudWatch, RDS, Route 53, S3, and Auto Scaling.
- Hands-on experience with Open Telemetry, Grafana, Datadog, Prometheus, or similar monitoring platforms.
- Strong knowledge of Linux administration, networking, system performance tuning, and troubleshooting.
- Experience with Infrastructure as Code using Terraform or CloudFormation.
- Proficiency in scripting using Python, Bash, or Go.
- Experience with Kubernetes and containerized workloads.
- Strong understanding of CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.).
- Experience leading incident management, production support, and RCA processes.
- Knowledge of SRE principles including SLIs, SLOs, and Error Budgets.
- Experience implementing monitoring, logging, alerting, and observability frameworks.
- Strong analytical, troubleshooting, and communication skills.
Preferred Qualifications
- Experience building or implementing AIOps solutions.
- Exposure to Large Language Models (LLMs) for operational automation.
- Experience with machine learning-based forecasting or anomaly detection.
- Hands-on experience administering Adobe Experience Manager (AEM).
- Experience managing Cloudflare CDN, WAF, DNS, and caching strategies.
- Knowledge of FinOps, cloud cost optimization, and capacity planning.
- AWS Solutions Architect, DevOps Engineer, or Kubernetes certifications are a plus.
Skills:- AIOps, Large Language Models (LLM), Machine Learning (ML), Amazon Web Services (AWS), Cloud-Native Infrastructure, OpenTelemetry, grafana, Datadog, Responsive Design, Root Cause Analysis (RCA), Infrastructure Design & Management, Infrastructure Rightsizing, Capacity Planning, Cost Anomaly Detection, Linux administration, Systems Administration, Adobe Experience Manager (AEM) Administration, Cloudflare CDN and Cross-functional Collaboration
📌 Site Reliability Engineer (SRE) (Pune)
🏢 EmbarkingOnVoyage Digital Solutions
📍 Pune