07 Aug
|
Chimera Technologies
|
Bengaluru
07 Aug
Chimera Technologies
Bengaluru
The Team & Role
We are building enterprise-grade AI applications — chatbots, autonomous agents, and multi-model workflows — on AWS and Kubernetes, integrating multiple LLM providers, vector databases, and async pipeline components in production at scale.
Our Stack: AWS (EKS, S3, RDS, Secrets Manager), Kubernetes, Prometheus, Grafana, Loki, Jaeger, OpenTelemetry, Python
What You'll Do
Typical activities that would be expected are.
• Own the end-to-end observability stack — metrics, logs, and traces — across infrastructure, Kubernetes, application services, and the AI/LLM request lifecycle.
• Build Grafana dashboards for three audiences — Engineering (traces, agent latency, errors), Operations (SLO burn, queue depth, MTTR), and Management (uptime, LLM cost, error budgets).
• Define SLIs, SLOs, and Error Budgets for all critical services including AI-specific targets like time-to-first-token for streaming responses.
• Design meaningful alert thresholds based on real user impact — LLM latency, token failures, RAG degradation, agent loop failures, and async queue depth.
• Instrument the full AI request lifecycle — orchestration steps, tool executions, RAG retrieval, and LLM calls — covering both streaming and non-streaming patterns.
• Write Python scripts for operational automation — health checks, custom Prometheus exporters, cert expiry monitoring, and runbook execution.
• Lead blameless post-mortems, enforce production readiness checks, and onboard engineering teams to the observability platform.
• Participate in on-call rotation as the observability and reliability escalation point.
• Alert management(create alerts for platform, infra, applications)
Requirements
• Strong hands-on experience with Grafana — dashboards, alerting, variable templating, and provisioning-as-code.
• Proficiency in Prometheus and PromQL — recording rules,
alert rules, and Alertmanager routing.
• Python scripting — mandatory; for automation, custom exporters, health checks, and operational tooling.
• Experience with OpenTelemetry — SDK instrumentation, collector pipelines, and trace-metric-log correlation.
• Distributed tracing with Jaeger or Tempo — reading multi-service traces and correlating with logs and metrics.
• SRE fundamentals — SLI/SLO/Error Budget design, incident management, post-mortems, and production readiness.
• Understanding of AI/LLM application behaviour — streaming vs non-streaming, time-to-first-token, token cost, and failure modes.
• Experience monitoring asynchronous systems — queue depth, worker saturation, retry rates, and job latency.
• Kubernetes and AWS — EKS, EC2, S3, CloudWatch, and pod-level resource monitoring.
• Strong communication skills across engineering and management audiences.
Good to Have
• Familiarity with RAG pipeline monitoring — extraction, chunking, embedding, and vector search latency.
• Experience tracing multi-step agent workflows and detecting orchestration failures.
• Experience monitoring vector databases — query latency, index health, and retrieval performance.
• Loki and LogQL for structured log querying and log-to-metric derivation.
• LLM cost monitoring — per-model token consumption, caching strategies, and spend attribution.
Benefits
- Participate in several organization wide programs, including a compulsory innovation-based Continuous improvement program providing you with platforms to showcase your talent.
- Insurance benefits for the self and the spouse, including maternity benefits.
- Ability to work on multiple products and platforms in a growing technology workplace.
- Fortnightly sessions to understand the direction of each function and interact with leadership.
- Hybrid working model – Remote + Office
📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Chimera Technologies
📍 Bengaluru