Job Title: Site Reliability Engineer I (Observability & Automation)
Department: Infrastructure / SRE
Report To: Engineering Manager
About the Role
We are looking for an energetic SRE 1 to join our Infrastructure team! We are on a mission to build a world-class observability platform. If you love digging into data,
automating boring tasks, and making systems "talk" to us through metrics and traces, we want to meet you.
You won't just be watching screens; you will be building the tools that keep our platform reliable and helping our developers sleep better at night.
What Youll Be Doing
- Driving OpenTelemetry Adoption: You will help us migrate from legacy agents (like Filebeat) to the OpenTelemetry (OTel) Collector. Youll tackle cool challenges like standardizing log formats and implementing tail-based sampling.
- Kubernetes Automation: Work on creating standardized base images with pre-configured OTel agents to make life easier for our developers.
- Taming the Data: Help us solve "High Cardinality" issues. Youll refine our metric ingestion strategies to ensure we are storing useful data without exploding our storage costs.
- Tooling & Dashboards: Build actionable dashboards and alerts in New Relic, Observe,
and Grafana. Youll help ensure our alerts actually mean something (reducing noise!).
- Incident Response: Participate in the on-call rotation using tools like Pagerduty to detect and resolve issues fast.
- Coding & Scripting: Use Python, Go, or Bash to automate manual operational tasks
What We Are Looking For
- Experience: 13 years of experience in SRE, DevOps, or Software Engineering.
- Kubernetes Pro: You know your way around K8s clusters, pods, and
deployments.
- Observability Obsessed: Experience with at least one major tool (Current Relic, Prometheus/Grafana, or ELK, Datadog).
- Cloud Native: Familiarity with public clouds (AWS, GCP, or Azure).
- Coding Skills: Comfortable writing scripts in Python, Go, or Java.
- Curious Mindset: You ask "why" when a system breaks and "how" we can prevent it next time.
Bonus Points (Nice to Have)
- Experience specifically with OpenTelemetry (OTel) instrumentation.
- Knowledge of Terraform or Ansible.
- Understanding of SLOs/SLIs and Golden Signals.
- Experience optimizing observability costs (ingestion volume, data retention
📌 SRE (Bengaluru)
🏢 Tekion
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.