Senior Site Reliability Engineer (Bengaluru)

Senior Site Reliability Engineer (Bengaluru)

05 Aug
|
Chimera Technologies
|
Bengaluru

05 Aug

Chimera Technologies

Bengaluru

The Team & Role We are building enterprise-grade AI applications chatbots, autonomous agents, and multi-model workflows on AWS and Kubernetes, integrating multiple LLM providers, vector databases, and async pipeline components in production at scale.

Our Stack: AWS (EKS, S3, RDS, Secrets Manager), Kubernetes, Prometheus, Grafana, Loki, Jaeger, OpenTelemetry, Python What You'll Do Typical activities that would be expected are.

Own the end-to-end observability stack metrics, logs, and traces across infrastructure, Kubernetes, application services, and the AI/LLM request lifecycle.

Build Grafana dashboards for three audiences Engineering (traces, agent latency, errors), Operations (SLO burn, queue depth, MTTR), and Management (uptime, LLM cost, error budgets).

Define SLIs, SLOs, and Error Budgets for all critical services including AI-specific targets like time-to-first-token for streaming responses.

Design meaningful alert thresholds based on real user impact LLM latency, token failures, RAG degradation, agent loop failures, and async queue depth.

Instrument the full AI request lifecycle orchestration steps, tool executions, RAG retrieval, and LLM calls covering both streaming and non-streaming patterns.

Write Python scripts for operational automation health checks, custom Prometheus exporters, cert expiry monitoring, and runbook execution.

Lead blameless post-mortems, enforce production readiness checks, and onboard engineering teams to the observability platform.

Participate in on-call rotation as the observability and reliability escalation point.

Alert management(create alerts for platform, infra, applications)



Requirements Solid hands-on experience with Grafana dashboards, alerting, variable templating, and provisioning-as-code.

Proficiency in Prometheus and PromQL recording rules, alert rules, and Alertmanager routing.

Python scripting mandatory; for automation, custom exporters, health checks, and operational tooling.

Experience with OpenTelemetry SDK instrumentation, collector pipelines, and trace-metric-log correlation.

Distributed tracing with Jaeger or Tempo reading multi-service traces and correlating with logs and metrics.

SRE fundamentals SLI/SLO/Error Budget design, incident management, post-mortems, and production readiness.

Understanding of AI/LLM application behaviour streaming vs non-streaming, time-to-first-token, token cost, and failure modes.

Experience monitoring asynchronous systems queue depth, worker saturation, retry rates, and job latency.

Kubernetes and AWS EKS, EC2, S3, CloudWatch, and pod-level resource monitoring.

Strong communication skills across engineering and management audiences.

Good to Have Familiarity with RAG pipeline monitoring extraction, chunking, embedding, and vector search latency.

Experience tracing multi-step agent workflows and detecting orchestration failures.

Experience monitoring vector databases query latency, index health, and retrieval performance.





Loki and LogQL for structured log querying and log-to-metric derivation.

LLM cost monitoring per-model token consumption, caching strategies, and spend attribution.

Benefits Participate in several organization wide programs, including a compulsory innovation-based Continuous improvement program providing you with platforms to showcase your talent.

Insurance benefits for the self and the spouse, including maternity benefits.

Ability to work on multiple products and platforms in a growing technology environment.

Fortnightly sessions to understand the direction of each function and interact with leadership.

Hybrid working model Remote + Office

Mandatory/Primary Technical Skills: · 3+ years of experience in designing and implementing Power Platform solutions.

Expertise in Power Apps (Canvas & Model-driven), Power Automate, Power BI, and Dataverse.

Strong Knowledge of Microsoft Dataverse, SQL, and SharePoint for data storage and retrieval.

Experience in AI Builder to extract the data from files Integration with Microsoft 365 services and third-party applications, services Proficiency in custom connectors, API integrations, and security models (OAuth, AD Authentication, etc.).

Experience with CI/CD pipelines and ALM (Application Lifecycle Management) for Power Platform.

Strong understanding of business units, governance, licensing, and security best practices.

Strong communication Skills Working experience in Agile Scrum Process Excellent problem-solving skills and ability to work in an agile environment.

Required Skill Profession

Other General

📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Chimera Technologies
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru