16 Aug
|
synapse Business Systems
|
Bengaluru
16 Aug
synapse Business Systems
Bengaluru
Role Overview
We are looking for an experienced Observability / SRE Engineer to join our Software Factory team in Bangalore. The ideal candidate will have strong hands-on experience in building and managing modern observability platforms, with expertise across metrics, logs, traces, SRE practices, and database monitoring.
The role will focus on designing, implementing, and maintaining observability solutions using SigNoz, Prometheus, OpenTelemetry, and Loki, helping engineering teams improve system reliability, performance, and troubleshooting capabilities.
Key Responsibilities
- Design, implement, and maintain enterprise-grade observability solutions across applications and infrastructure.
- Configure and manage SigNoz for application and infrastructure monitoring.
- Develop and maintain Prometheus based metrics collection, monitoring, and alerting.
- Implement distributed tracing using OpenTelemetry across microservices and applications.
- Configure Loki for centralized log collection, aggregation, and analysis.
- Develop dashboards, alerts, and monitoring strategies for metrics, logs, and traces.
- Define and monitor SLIs, SLOs, and SLAs in line with SRE best practices.
- Analyze system performance, identify bottlenecks, and support proactive incident prevention.
- Work closely with development, DevOps, and infrastructure teams to improve application reliability and observability.
- Monitor database performance, availability, capacity, and potential bottlenecks.
- Participate in incident response, root cause analysis, and post-incident reviews.
- Automate monitoring and operational processes wherever possible.
- Establish observability standards and best practices across the Software Factory environment.
Required Skills
- Strong hands-on experience with SigNoz.
- Good expertise in Prometheus and metrics-based monitoring.
- Strong understanding of OpenTelemetry and distributed tracing.
- Experience with Loki and centralized log management.
- Solid understanding of metrics, logs, and traces and how they work together within an observability stack.
- Strong understanding of SRE principles, including SLIs, SLOs, error budgets, incident management, and reliability engineering.
- Experience with database monitoring and performance troubleshooting.
- Good understanding of Linux, networking, application performance, and cloud-native environments.
- Experience working with contemporary DevOps and CI/CD environments is preferred.
Preferred Qualifications
- 4 to 8 years of experience in Observability, SRE, DevOps, Platform Engineering, or related roles.
- Experience working with microservices and distributed systems.
- Knowledge of Kubernetes and cloud platforms such as AWS, Azure, or GCP is an advantage.
- Experience with Grafana or similar visualization platforms is a plus.
- Strong troubleshooting, analytical, and incident management skills.
- Bachelor's degree in Computer Science, Engineering, or a related field.
What You Will Do You will be responsible for ensuring that applications and infrastructure are observable, reliable, and performant by building a unified monitoring strategy across metrics, logs, and traces. You will work closely with engineering teams to identify reliability gaps and continuously improve system health and operational efficiency.
Location: Bangalore, India
📌 Observability / SRE Engineer (Bengaluru)
🏢 synapse Business Systems
📍 Bengaluru