About Exavalu
Founded in 2018 and headquartered in Newport Beach, California, Exavalu is built by former CIOs, CXOs, and Big 5 consulting leaders. We specialize in delivering enterprise-scale digital transformation solutions across Insurance, Banking & Financial Services, and Healthcare industries.
With delivery centers across India, Canada and the US, Exavalu is recognized for its agile culture, innovation mindset, and employee-first environment. We are committed to building modern, scalable, and business-driven technology solutions for global enterprise clients.
Job Title : Senior Observability Engineer / Platform Engineer
Experience : 6-10 years
Timings : 12 pm to 10 pm
Mode : Remote
:
Role Summary
We are looking for a hands-on Observability Engineer with strong experience in cloud-native platforms, monitoring, logging, alerting, and operational excellence. The ideal candidate will have experience building and managing enterprise observability solutions across Kubernetes and public cloud environments, enabling proactive monitoring, incident response, and reliability engineering practices.
Key Responsibilities
Design, implement, and maintain enterprise observability platforms covering metrics, logs, traces, and events. Build observability solutions using tools such as Prometheus, Grafana , OpenSearch,
Splunk , Elastic, Datadog, Dynatrace, New Relic, or equivalent platforms. Develop dashboards, SLOs, SLIs, alerting rules, and service health monitoring frameworks. Integrate monitoring and observability capabilities within Kubernetes and cloud-native environments
. Enable incident management, root cause analysis, and operational troubleshooting through observability best practices. Automate monitoring configuration and platform onboarding using Infrastructure-as-Code and CI/CD pipelines (Good to have).
Collaborate with engineering, platform, SRE, and operations teams to improve reliability, performance, and availability.
Support observability maturity initiatives including distributed tracing, AIOps, and intelligent alerting.
Required Skills:
- 6+ years of experience in Platform Engineering, SRE, DevOps, Cloud Operations, or Observability Engineering.
- Solid expertise in: Prometheus, Grafana, Splunk / OpenSearch / Elastic Distributed tracing solutions (Jaeger, Tempo, Open Telemetry , etc.)
- Experience in creating dashboard mainly on Splunk and good to know this on Grafana as well.
- Good understanding of Kubernetes, Docker, and containerized workloads.
- Experience with AWS, Azure, or GCP environments. Knowledge of incident management, alert tuning, and troubleshooting production environments.
- Experience with scripting and automation using Python, Shell, or similar languages.
- Familiarity with CI/CD and Infrastructure-as-Code tools such as Terraform, Jenkins, GitHub Actions, or Argo CD.
- Preferred Skills Exposure to SRE practices, error budgets, SLO/SLI frameworks.
- Good to have with AIOps , automated remediation, or intelligent incident response.
- Working experience in enterprise-scale production platforms. Exposure to security and governance considerations in cloud-native environments.
Share your updated profile to
[email protected]
Diversity Inclusion:
At Exavalu, we are committed to building a diverse and inclusive workforce. We welcome applications for employment from all qualified candidates, regardless of race, color, gender, national or ethnic origin, age, disability, religion, sexual orientation, gender identity or any other status protected by applicable law. We nurture a culture that embraces all individuals and promotes diverse perspectives, where you can make an impact and grow your career.
Exavalu also promotes flexibility depending on the needs of employees, customers and the business. It might be part-time work, working outside normal 9-5 business hours or working remotely. We also have a welcome back program to help people get back to the mainstream after a long break due to health or family reasons.
📌 Senior Observability Engineer / Platform Engineer (India)
🏢 Exavalu
📍 India