Senior Site Reliability Engineer (Bengaluru)

Senior Site Reliability Engineer (Bengaluru)

07 Aug
|
Chimera Technologies
|
Bengaluru

07 Aug

Chimera Technologies

Bengaluru

The Team & Role

We are building enterprise-grade AI applications — chatbots, autonomous agents, and multi-model workflows — on AWS and Kubernetes, integrating multiple LLM providers, vector databases, and async pipeline components in production at scale.

Our Stack: AWS (EKS, S3, RDS, Secrets Manager), Kubernetes, Prometheus, Grafana, Loki, Jaeger, OpenTelemetry, Python

What You'll Do

Typical activities that would be expected are.

• Own the end-to-end observability stack — metrics, logs, and traces — across infrastructure, Kubernetes, application services, and the AI/LLM request lifecycle.

• Build Grafana dashboards for three audiences — Engineering (traces, agent latency, errors), Operations (SLO burn, queue depth, MTTR), and Management (uptime, LLM cost, error budgets).

• Define SLIs, SLOs, and Error Budgets for all critical services including AI-specific targets like time-to-first-token for streaming responses.

• Design meaningful alert thresholds based on real user impact — LLM latency, token failures, RAG degradation, agent loop failures, and async queue depth.

• Instrument the full AI request lifecycle — orchestration steps, tool executions, RAG retrieval, and LLM calls — covering both streaming and non-streaming patterns.

• Write Python scripts for operational automation — health checks, custom Prometheus exporters, cert expiry monitoring, and runbook execution.

• Lead blameless post-mortems, enforce production readiness checks, and onboard engineering teams to the observability platform.

• Participate in on-call rotation as the observability and reliability escalation point.

• Alert management(create alerts for platform, infra, applications)

Requirements

• Strong hands-on experience with Grafana — dashboards, alerting, variable templating, and provisioning-as-code.

• Proficiency in Prometheus and PromQL — recording rules,



alert rules, and Alertmanager routing.

• Python scripting — mandatory; for automation, custom exporters, health checks, and operational tooling.

• Experience with OpenTelemetry — SDK instrumentation, collector pipelines, and trace-metric-log correlation.

• Distributed tracing with Jaeger or Tempo — reading multi-service traces and correlating with logs and metrics.

• SRE fundamentals — SLI/SLO/Error Budget design, incident management, post-mortems, and production readiness.

• Understanding of AI/LLM application behaviour — streaming vs non-streaming, time-to-first-token, token cost, and failure modes.

• Experience monitoring asynchronous systems — queue depth, worker saturation, retry rates, and job latency.

• Kubernetes and AWS — EKS, EC2, S3, CloudWatch, and pod-level resource monitoring.

• Strong communication skills across engineering and management audiences.

Good to Have

• Familiarity with RAG pipeline monitoring — extraction, chunking, embedding, and vector search latency.

• Experience tracing multi-step agent workflows and detecting orchestration failures.

• Experience monitoring vector databases — query latency, index health, and retrieval performance.

• Loki and LogQL for structured log querying and log-to-metric derivation.

• LLM cost monitoring — per-model token consumption, caching strategies, and spend attribution.

Benefits

- Participate in several organization wide programs, including a compulsory innovation-based Continuous improvement program providing you with platforms to showcase your talent.

- Insurance benefits for the self and the spouse, including maternity benefits.

- Ability to work on multiple products and platforms in a growing technology workplace.

- Fortnightly sessions to understand the direction of each function and interact with leadership.

- Hybrid working model – Remote + Office

📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Chimera Technologies
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru