Lead Engineer - Reliability & Observability (Bengaluru)

Lead Engineer - Reliability & Observability (Bengaluru)

09 Oct
|
Cybrilla
|
Bengaluru

09 Oct

Cybrilla

Bengaluru

Job Summary

We are looking for a Lead Engineer to own and significantly advance our Reliability Observability platform. We already have an established observability stack, including ELK, Grafana, Loki and monitoring dashboards. As our multi-tenant platform scales across multiple deployment sites and environments, you will drive the next phase of its evolution, making it easier for engineering teams to understand, monitor and improve the reliability of the systems they build.

This is a hands-on engineering leadership role. You will own and evolve our observability stack, establish reliability engineering standards and drive automation across our production environment. You will work closely with product/application engineering, cloud infrastructure and security teams to build resilient, self-operating systems across an increasingly complex distributed infrastructure.

What Youll Own

- Observability Platform: Own, scale and continuously improve our existing metrics, logging, distributed tracing, dashboards and alerting infrastructure to support increasing platform scale, multiple deployment sites and multi-tenant environments.
- Reliability Engineering: Establish SLO frameworks, error-budget reporting, production-readiness standards and reliability engineering best practices.
- Developer Enablement: Build self-service observability capabilities, standardized instrumentation, reusable dashboards and actionable alerting for engineering teams.
- Incident Management: Establish incident-management tooling, escalation processes, runbooks and post-incident review practices. Drive improvements that prevent recurring incidents.
- Automation: Eliminate operational toil through automation, intelligent alerting, automated remediation and resilient system design.
- Technical Leadership: Own the architecture and roadmap for Reliability Observability, mentor engineers and collaborate with other engineering teams on cross-cutting reliability challenges.

Application teams own their services, and cloud infrastructure engineers own the underlying infrastructure.



Your role is to provide the shared capabilities, standards and engineering expertise that help them operate reliably.

What Were Looking For

- 7+ years of software engineering, platform engineering or SRE experience, including operating business-critical production systems at high scale and complexity.
- Strong experience with observability technologies such as Elasticsearch, Logstash, Kibana (ELK), Grafana, Loki, OpenTelemetry, Jaeger or equivalent platforms.
- Hands-on experience with AWS, Kubernetes, Linux, networking and distributed systems.
- Robust programming and scripting skills in Go, Java, Python or similar languages, with an automation-first mindset.
- Experience designing, scaling and improving production observability platforms, effective monitoring, alerting, SLOs and incident-management practices.
- A solid understanding of distributed-system failure modes, scalability, fault tolerance and production debugging.
- Experience with infrastructure as code, CI/CD and cloud-native architectures.
- Ability to drive technical initiatives independently, make sound architectural decisions and influence engineering teams without relying on organizational authority.

Good to Have

- Experience operating high-availability, multi-tenant SaaS platforms, particularly in fintech or other regulated environments.
- Experience designing observability architectures for geographically distributed infrastructure, multiple Kubernetes clusters and isolated tenant environments.
- Experience with large-scale telemetry pipelines, observability cost optimization and high-cardinality metrics.
- Experience with chaos engineering, resilience testing and automated incident remediation.
- Experience building internal developer platforms or self-service engineering tools.

Above all, we are looking for an engineer who enjoys building reliable platforms, not just maintaining monitoring tools.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Lead Engineer - Reliability & Observability (Bengaluru)
🏢 Cybrilla
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead engineer - reliability & observability (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: lead engineer - reliability & observability (bengaluru) / bengaluru