06 Aug
|
Umanist NA
|
Pune
⚠️ Strict Screening Note: Please submit profiles only if all mandatory criteria are met. Candidates must have 8 years of relevant experience(Overall in IT 10yrs ) with the required mandatory skills. Official notice period must be Immediate (already serving notice in the current organization) or up to 30 days only. Candidates with notice periods exceeding 30 days will not be considered.
Hiring: Senior Site Reliability Engineer (SRE) / DevOps Engineer
? Location: Viman Nagar, Pune (Work From Office)
? Experience: 10 Years
? CTC: Up to ₹25-27 LPA
? Joining: Immediate Joiners Preferred
⏰ Shift Timings: 3:00 PM – 12:00 AM (Monday – Friday)
? On-Call: 24/7 Production Support (Rotation)
Role Overview
We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to drive the reliability, scalability, security, observability, and performance of mission-critical production systems. The ideal candidate should have strong expertise in Azure Cloud, Kubernetes, DevOps, SRE practices, and modern observability tools, with the ability to balance production support and long-term reliability engineering initiatives.
Must-Have Skills Experience
- 10 years of overall IT experience.
- Minimum 6–7 years of hands-on DevOps/SRE experience.
- Experience managing large-scale production environments.
- Immediate joiners preferred.
Cloud &
- Infrastructure
- Microsoft Azure (Mandatory)
- Kubernetes (Strong hands-on experience)
- Terraform (Infrastructure as Code)
- Helm
- GitHub / GitLab / Azure Repos
DevOps &
- SRE
- Site Reliability Engineering (SRE) Practices
- Incident Response & 24/7 On-Call Support
- Root Cause Analysis (RCA)
- SLI / SLO / SLA Management
- Error Budgets
- Capacity Planning
- Toil Reduction
- Reliability Engineering
Observability &
- Monitoring
- OpenTelemetry
- Prometheus
- Grafana
- Azure Monitor
- Datadog
- Distributed Tracing
- Metrics, Logs &
- Traces
- Golden Signals Monitoring:
- Latency
- Traffic
- Errors
- Saturation
Programming &
- System Administration
- Python
- Bash Scripting
- Linux Administration
- Networking Fundamentals:
- DNS
- TCP/IP
- Load Balancing
- SSL/TLS
Positive-to-Have Skills
- Google Cloud Platform (GCP)
- AWS (EC2, S3, RDS, IAM, VPC, CloudWatch)
- Go (Golang)
- OpenSearch / ELK Stack
- Azure AI Services
- AI Foundry
- AI/ML Infrastructure
- RAG (Retrieval-Augmented Generation) Workloads
- Experience with Distributed Systems Architecture
Key Responsibilities Production Support &
- Incident Management
- Participate in 24/7 production on-call rotation.
- Troubleshoot and resolve high-severity production incidents.
- Lead Root Cause Analysis (RCA) and post-mortem activities.
- Reduce MTTR through proactive reliability improvements.
- Maintain SLAs, SLOs, and service reliability.
Site Reliability Engineering
- Implement SRE best practices across production environments.
- Define and improve SLIs, SLOs, SLAs, and Error Budgets.
- Reduce operational toil through automation.
- Improve service availability, scalability, and disaster recovery readiness.
- Perform reliability reviews and capacity planning.
Cloud &
- Infrastructure
- Design, deploy,
and manage Azure infrastructure.
- Administer Kubernetes clusters and containerized workloads.
- Manage Infrastructure as Code using Terraform.
- Deploy applications using Helm and Git-based CI/CD workflows.
Monitoring &
- Observability
- Build and maintain observability solutions using OpenTelemetry.
- Implement monitoring with Prometheus, Grafana, Azure Monitor, and Datadog.
- Establish monitoring based on Golden Signals.
- Improve logging, tracing, alerting, and performance monitoring across distributed systems.
Security &
- Compliance
- Implement cloud security best practices.
- Manage IAM, network security, and secrets management.
- Support vulnerability remediation and compliance initiatives.
What We're Looking For
- Strong ownership and accountability.
- Excellent debugging and analytical skills.
- Calm decision-making during production incidents.
- Deep understanding of SRE principles and observability.
- Passion for automation, scalability, and continuous improvement.
- Strong communication and collaboration skills.
Screening Notes (Mandatory)
- Immediate joiners only.
- Minimum 10 years of experience.
- Strong Azure Cloud experience (Mandatory).
- GCP exposure is acceptable as an added advantage.
- Minimum 6–7 years of hands-on SRE(Primary)/DevOps experience.
- Strong expertise in Kubernetes, OpenTelemetry, Golden Signals, Prometheus, Grafana, Terraform, and SRE Practices is mandatory.
- Candidates should have experience supporting high-availability production environments and be comfortable with 24/7 on-call rotation.
Skills: grafana,prometheus,sre practices,python,sre,opentelemetry,kubernetes,terraform,golden signals,24/7 on-call rotation,azure cloud
📌 Senior Site Reliability Engineer (SRE) / DevOps Engineer (Pune)
🏢 Umanist NA
📍 Pune