06 Aug
|
Chimera Technologies
|
Bengaluru
06 Aug
Chimera Technologies
Bengaluru
The Team & Role
We are building enterprise-grade AI applications — chatbots, autonomous agents, and multi-model workflows — on AWS and Kubernetes, integrating multiple LLM providers, vector databases, and async pipeline components in production at scale.
Our Stack: AWS (EKS, S3, RDS, Secrets Manager), Kubernetes, Prometheus, Grafana, Loki, Jaeger, OpenTelemetry, Python
What You'll Do
Typical activities that would be expected are.
- Own the end-to-end observability stack — metrics, logs, and traces — across infrastructure, Kubernetes, application services, and the AI/LLM request lifecycle.
- Build Grafana dashboards for three audiences — Engineering (traces, agent latency, errors), Operations (SLO burn, queue depth, MTTR), and Management (uptime, LLM cost, error budgets).
- Define SLIs, SLOs, and Error Budgets for all critical services including AI-specific targets like time-to-first-token for streaming responses.
- Design meaningful alert thresholds based on real user impact — LLM latency, token failures, RAG degradation, agent loop failures, and async queue depth.
- Instrument the full AI request lifecycle — orchestration steps, tool executions, RAG retrieval, and LLM calls — covering both streaming and non-streaming patterns.
- Write Python scripts for operational automation — health checks, custom Prometheus exporters, cert expiry monitoring, and runbook execution.
- Lead blameless post-mortems, enforce production readiness checks, and onboard engineering teams to the observability platform.
- Participate in on-call rotation as the observability and reliability escalation point.
- Alert management(create alerts for platform, infra, applications)
Requirements
• Solid hands-on experience with Grafana — dashboards, alerting, variable templating, and provisioning-as-code.
- Proficiency in Prometheus and PromQL — recording rules,
alert rules, and Alertmanager routing.
- Python scripting — mandatory; for automation, custom exporters, health checks, and operational tooling.
- Experience with OpenTelemetry — SDK instrumentation, collector pipelines, and trace-metric-log correlation.
- Distributed tracing with Jaeger or Tempo — reading multi-service traces and correlating with logs and metrics.
- SRE fundamentals — SLI/SLO/Error Budget design, incident management, post-mortems, and production readiness.
- Understanding of AI/LLM application behaviour — streaming vs non-streaming, time-to-first-token, token cost, and failure modes.
- Experience monitoring asynchronous systems — queue depth, worker saturation, retry rates, and job latency.
- Kubernetes and AWS — EKS, EC2, S3, CloudWatch, and pod-level resource monitoring.
- Strong communication skills across engineering and management audiences.
Good to Have
- Familiarity with RAG pipeline monitoring — extraction, chunking, embedding, and vector search latency.
- Experience tracing multi-step agent workflows and detecting orchestration failures.
- Experience monitoring vector databases — query latency, index health, and retrieval performance.
- Loki and LogQL for structured log querying and log-to-metric derivation.
- LLM cost monitoring — per-model token consumption, caching strategies, and spend attribution.
Benefits
- Participate in several organization wide programs, including a compulsory innovation-based Continuous improvement program providing you with platforms to showcase your talent.
- Insurance benefits for the self and the spouse, including maternity benefits.
- Ability to work on multiple products and platforms in a growing technology environment.
- Fortnightly sessions to understand the direction of each function and interact with leadership.
- Hybrid working model – Remote + Office
📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Chimera Technologies
📍 Bengaluru