02 Sep
|
Umanist Staffing
|
Pune
02 Sep
Umanist Staffing
Pune
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Location: Viman Nagar, Pune – Work From Office
Experience Overall(must have): 8 Years
CTC: Up to ₹25 LPA
Notice Period: Immediate Joiners Only within 15d or (if serving max 30days)
Working Hours: 3:00 PM – 12:00 AM, Monday to Friday
On-Call: 24/7 Production Support – On-Call Rotation Required
Employment Type: Full-Time
About the Role
We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.
The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.
Must-Have Skills & Experience1. SRE & Production Operations
- Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering.
- Hands-on experience with 24/7 production support and on-call operations.
- Strong experience in incident management, troubleshooting, RCA, and post-mortems.
- Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering.
- Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies.
- Ability to improve system availability, performance, scalability, and operational reliability.
2. Cloud & Infrastructure
- Strong hands-on experience with Microsoft Azure, AWS, and/or GCP.
- Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.
- Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.
- Experience with:
- Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS
- AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS
- GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring
3. Kubernetes & Containerization
- Strong hands-on experience with Kubernetes and containerized workloads.
- Experience with AKS / EKS / GKE or equivalent Kubernetes environments.
- Hands-on experience with Helm deployments.
- Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.
4. Infrastructure as Code & DevOps
- Hands-on experience with Terraform / Infrastructure as Code (IaC).
- Experience with Git-based workflows using GitHub, GitLab, or Azure Repos.
- Strong DevOps automation and CI/CD understanding.
- Strong scripting skills in Python and/or Bash.
5. Monitoring & Observability
- Strong hands-on experience with OpenTelemetry.
- Experience with monitoring and observability tools such as:
- Prometheus
- Grafana
- Datadog
- Azure Monitor
- AWS CloudWatch
- GCP Cloud Monitoring
- Strong understanding of metrics, logs, distributed tracing, and alerting.
- Experience implementing monitoring based on Golden Signals:
- Latency
- Traffic
- Errors
- Saturation
- Ability to develop symptom-based, user-impact-focused alerting.
6. Linux & Networking
- Strong knowledge of Linux system administration.
- Strong understanding of:
- DNS
- TCP/IP
- Load Balancing
- SSL/TLS
- Networking fundamentals
- Experience supporting highly available production environments.
7. Incident & Reliability Engineering
- Ability to rapidly diagnose and resolve high-severity production incidents.
- Experience driving MTTR reduction.
- Strong debugging and analytical problem-solving skills.
- Ability to identify recurring issues and implement permanent corrective/preventive solutions.
Good-to-Have Skills
- Experience working across Azure + AWS + GCP in a multi-cloud environment.
- Knowledge of Go (Golang).
- Experience with OpenSearch / ELK Stack.
- Experience supporting AI/ML workloads in production.
- Exposure to Azure AI Services and Azure AI Foundry.
- Experience supporting RAG (Retrieval-Augmented Generation) workloads.
- Experience designing infrastructure for AI/ML platforms.
- Experience building enterprise-wide OpenTelemetry observability frameworks.
- Strong understanding of distributed systems architecture.
- Exposure to advanced cloud-native architectures and reliability patterns.
- Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.
Key ResponsibilitiesProduction & Incident Management
- Participate in the 24/7 on-call rotation.
- Diagnose, mitigate, and resolve production incidents.
- Lead RCA and post-incident reviews.
- Implement corrective and preventive actions.
- Continuously improve MTTR and production stability.
Reliability Engineering
- Define and improve SLIs, SLOs, SLAs, and Error Budgets.
- Identify and eliminate operational toil.
- Conduct reliability and capacity reviews.
- Improve redundancy, failover, disaster recovery, and system resilience.
Cloud & Infrastructure
- Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.
- Manage Kubernetes clusters and containerized applications.
- Implement and maintain Infrastructure as Code using Terraform.
- Support CI/CD and Git-based development workflows.
Observability & Performance
- Build and improve monitoring, logging, metrics, and tracing.
- Implement OpenTelemetry and distributed tracing.
- Establish Golden Signals-based monitoring and alerting.
- Identify and resolve infrastructure and application performance bottlenecks.
Security
- Implement cloud security best practices around IAM, network segmentation, and secrets management.
- Support vulnerability remediation and compliance initiatives.
- Collaborate with Development, Security, and Infrastructure teams.
Ideal Candidate
We are looking for someone with:
- Robust SRE mindset and production ownership.
- Excellent troubleshooting and incident-management skills.
- Hands-on expertise in Cloud + Kubernetes + Terraform + Observability.
- Strong understanding of OpenTelemetry and Golden Signals.
- Experience working in highly available, production-critical environments.
- Ability to remain calm and make effective decisions during critical incidents.
- Strong communication and cross-functional collaboration skills.
- Passion for automation, scalability, reliability, and continuous improvement.
Important Hiring Criteria
Must be:
- 7+ years relevant experience
- Immediate joiner
- Willing to work from office in Viman Nagar, Pune
- Comfortable with 3:00 PM – 12:00 AM shift
- Comfortable with 24/7 on-call rotation
- Strong hands-on SRE/DevOps experience
- Strong Cloud + Kubernetes + Observability experience
- Strong production incident management experience
Good to have:
- Multi-cloud: Azure + AWS + GCP
- OpenTelemetry
- AI/ML or RAG production workloads
- Azure AI / AI Foundry
- Go
- OpenSearch / ELK
- Distributed systems
📌 Senior Site Reliability Engineer (SRE) / DevOps Engineer (Pune)
🏢 Umanist Staffing
📍 Pune