Key Responsibilities
Production Support & Reliability
- Manage and support high-availability production systems
- Perform incident management, RCA, and problem resolution
- Define and track SLIs, SLOs, and error budgets
Observability & Monitoring
- Design and implement end-to-end observability solutions
- Build and maintain dashboards using Grafana, Kibana, Tableau
- Implement distributed tracing and metrics collection using OpenTelemetry (OTEL)
- Work with ITRS Geneos, GCO (Google Cloud Operations) for monitoring
- Improve logging, tracing, and alerting frameworks
Automation & Toil Reduction
- Identify repetitive manual tasks and eliminate them via automation (toil reduction)
- Develop scripts/tools using Python, Shell, Java
- Automate health checks, deployments, alerting, and recovery processes
- Drive self-healing systems and auto-remediation solutions
DevOps & CI/CD
- Build and maintain CI/CD pipelines
- Implement DevOps best practices for faster and reliable releases
- Ensure pipeline monitoring and optimization
Required Skills
Core
- Unix/Linux
- SQL
Programming
- Python, Java
DevOps & Scheduling
- CI/CD pipeline knowledge
- AutoSys or similar scheduler
Observability Stack
- Grafana, Grafana Alloy
- Kibana / ELK
- OpenTelemetry (OTEL)
- ITRS Geneos
- GCO or equivalent
Cloud & Containers
- Google Cloud Platform (GCP)
- OpenShift / Kubernetes
Data & Visualization
- Airflow
- Dataflow
- Tableau
If Interested, Please share your profile to
[email protected]
📌 Senior DevOps / Site Reliability Engineer (SRE) (Pune)
🏢 LTM
📍 Pune