06 Oct
|
Virtusa
|
Gurugram
Role Overview:
- Drive reliability transformation, establish an SRE Center of Excellence (CoE), and enable business visibility. This role requires defining SRE frameworks while staying hands-on in Java code bases to refactor logging and instrument telemetry using OpenTelemetry.
- Core Responsibilities:
- SRE CoE & Architecture: Set up SRE frameworks, tooling radar, and governance. Guide transition from internal cloud to AWS. Track MTTD, MTTR, DORA metrics, and SLOs/Error Budgets. Lay ground for future AI/AIOps capabilities.
- Hands-on Java Uplift: Work directly in Java microservices code to refactor logging, capture business errors/exceptions, and increase business-level transaction visibility.
- Observability & OpenTelemetry: Implement OpenTelemetry tracing, APM, RUM, and log management across Splunk and Grafana.
Build actionable dashboards and Twilio alerting.
- CI/CD Automation: Build single-click GitLab CI/CD deployment pipelines with automated release governance and quality gates.
- Mandatory Requirements:
- Java Development: Hands-on experience developing/refactoring Java microservices and upgrading code-level logging/instrumentation.
- Observability Stack: OpenTelemetry, Splunk, Grafana, APM, RUM, Log Management, and Twilio integration.
- SRE Practices: Deep knowledge of SLOs, SLIs, Error Budgets, MTTD, MTTR, and DORA metrics.
- CI/CD: Advanced GitLab CI/CD pipeline automation and release quality gates
📌 Site Reliability Engineer (Gurugram)
🏢 Virtusa
📍 Gurugram