30 Sep
|
Experis
|
Bengaluru
Job Title: SRE Observability
Location: Hyderabad, Chennai, Bangalore
Experience: 8+ Years
Job Type: Permanent
Notice Period: Immediate
Key Requirements
- Strong experience in Site Reliability Engineering, Observability Engineering, Platform Engineering or DevOps within cloud-native and business-critical environments.
- Hands-on experience with Dynatrace, including dashboarding, alerting, telemetry analysis, synthetic monitoring and configuration management.
- Strong understanding of observability principles across metrics, logs, traces, events, service dependencies and user journeys.
- Experience applying observability configuration using self-service or automated platforms, with valuable governance and standardisation.
- Strong working knowledge of Kubernetes platforms, preferably GKE, and service mesh technologies such as Istio.
- Experience using OpenTelemetry data, including standardised ingestion, reporting, dashboards and alert generation.
- Strong working knowledge of Kafka observability, including monitoring brokers, topics, partitions, consumers, lag, throughput, errors and resilience indicators.
- Experience defining and implementing SLI/SLO frameworks,
including burn-rate alerting, error budgets and reliability dashboards.
- Strong experience refining alerts to reduce false positives and improve actionable alerting for genuine service issues.
- Good understanding of incident management, problem management, operational readiness and continuous improvement practices.
Preferred Experience
- Banking or core banking platform experience.
- Experience working with Thought Machine Vault.
- Experience supporting observability for large-scale distributed systems and regulated production platforms.
- Experience integrating observability data with cloud cost data to support FinOps insight and cost optimisation.
- Experience with GCP cost management, FinOps practices, cloud consumption reporting or cost-to-service mapping.
- Experience designing training material, running enablement sessions and building internal knowledge across engineering teams.
- Knowledge of operational resilience frameworks, service health reporting and evidence-based reliability improvement.
📌 SRE Observability (Bengaluru)
🏢 Experis
📍 Bengaluru