06 Aug
|
Umanist NA
|
Pune
⚠️ Strict Screening Note: Please submit profiles only if all mandatory criteria are met. Candidates must have 8+ years of relevant experience(Overall in IT 10yrs+) with the required mandatory skills. Official notice period must be Immediate (already serving notice in the current organization) or up to 30 days only. Candidates with notice periods exceeding 30 days will not be considered.
Hiring: Senior Site Reliability Engineer (SRE) / DevOps Engineer ? Location: Viman Nagar, Pune (Work From Office) ? Experience: 10+ Years ? CTC: Up to ₹25-27 LPA ? Joining: Immediate Joiners Preferred ⏰ Shift Timings: 3:00 PM – 12:00 AM (Monday – Friday) ? On-Call: 24/7 Production Support (Rotation) Role Overview We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to drive the reliability, scalability, security, observability, and performance of mission-critical production systems. The ideal candidate should have strong expertise in Azure Cloud, Kubernetes, DevOps, SRE practices, and modern observability tools, with the ability to balance production support and long-term reliability engineering initiatives. Must-Have Skills Experience 10+ years of overall IT experience.
Minimum 6–7 years of hands-on DevOps/SRE experience.
Experience managing large-scale production environments.
Immediate joiners preferred. Cloud & Infrastructure Microsoft Azure (Mandatory)
Kubernetes (Strong hands-on experience)
Terraform (Infrastructure as Code)
Helm
GitHub / GitLab / Azure Repos DevOps & SRE Site Reliability Engineering (SRE) Practices
Incident Response & 24/7 On-Call Support
Root Cause Analysis (RCA)
SLI / SLO / SLA Management
Error Budgets
Capacity Planning
Toil Reduction
Reliability Engineering Observability & Monitoring OpenTelemetry
Prometheus
Grafana
Azure Monitor
Datadog
Distributed Tracing
Metrics, Logs & Traces
Golden Signals Monitoring
Latency
Traffic
Errors
Saturation
Programming & System Administration Python
Bash Scripting
Linux Administration
Networking Fundamentals
DNS
TCP/IP
Load Balancing
SSL/TLS
Good-to-Have Skills Google Cloud Platform (GCP)
AWS (EC2, S3, RDS, IAM, VPC, CloudWatch)
Go (Golang)
OpenSearch / ELK Stack
Azure AI Services
AI Foundry
AI/ML Infrastructure
RAG (Retrieval-Augmented Generation) Workloads
Experience with Distributed Systems Architecture Key Responsibilities Production Support & Incident Management Participate in 24/7 production on-call rotation.
Troubleshoot and resolve high-severity production incidents.
Lead Root Cause Analysis (RCA) and post-mortem activities.
Reduce MTTR through proactive reliability improvements.
Maintain SLAs, SLOs, and service reliability.
Site Reliability Engineering
Implement SRE best practices across production environments.
Define and improve SLIs, SLOs, SLAs, and Error Budgets.
Reduce operational toil through automation.
Improve service availability, scalability, and disaster recovery readiness.
Perform reliability reviews and capacity planning. Cloud & Infrastructure Design, deploy, and manage Azure infrastructure.
Administer Kubernetes clusters and containerized workloads.
Manage Infrastructure as Code using Terraform.
Deploy applications using Helm and Git-based CI/CD workflows. Monitoring & Observability Build and maintain observability solutions using OpenTelemetry.
Implement monitoring with Prometheus, Grafana, Azure Monitor, and Datadog.
Establish monitoring based on Golden Signals.
Improve logging, tracing, alerting, and performance monitoring across distributed systems. Security & Compliance Implement cloud security best practices.
Manage IAM, network security, and secrets management.
Support vulnerability remediation and compliance initiatives.
What We're Looking For Strong ownership and accountability.
Excellent debugging and analytical skills.
Calm decision-making during production incidents.
Deep understanding of SRE principles and observability.
Passion for automation, scalability, and continuous improvement.
Robust communication and collaboration skills.
Screening
Notes (Mandatory) Immediate joiners only.
Minimum 10+ years of experience.
Strong Azure Cloud experience (Mandatory).
GCP exposure is acceptable as an added advantage.
Minimum 6–7 years of hands-on SRE(Primary)/DevOps experience.
Strong expertise in Kubernetes, OpenTelemetry, Golden Signals, Prometheus, Grafana, Terraform, and SRE Practices is mandatory.
Candidates should have experience supporting high-availability production environments and be comfortable with 24/7 on-call rotation.
Skills: grafana,prometheus,sre practices,python,sre,opentelemetry,kubernetes,terraform,golden signals,24/7 on-call rotation,azure cloud
📌 Senior Site Reliability Engineer (SRE) / DevOps Engineer (Pune)
🏢 Umanist NA
📍 Pune