03 Sep
|
Rimini Street
|
Hyderabad
03 Sep
Rimini Street
Hyderabad
Position Summary
We are actively seeking a Sr. Site Reliability Engineer. This role is based in India, Hyderabad.
The Senior Site Reliability Engineer is responsible for the reliability, availability, performance, observability and operational excellence of Rimini Streets Agentic AI ERP Platform.
The role covers production operations, incident management, resilience engineering, capacity planning and disaster recovery across Kubernetes infrastructure, AI Agents, MCP Servers, AI/LLM Gateways, Temporal workflows, RAG services and related enterprise AI platform components.
The SRE works closely with Platform Engineering, AI Engineering, Product Engineering, DevOps and Security teams to ensure that platform services are scalable, secure, supportable and ready for enterprise production use across cloud-hosted, customer-hosted and air-gapped environments.
Key Responsibilities
Reliability and Production Operations
- Own service reliability, availability and operational health across the Agentic AI Platform.
- Define, measure and maintain Service Level Indicators, Service Level Objectives and Error Budgets.
- Conduct production-readiness and reliability reviews for current services and major platform changes.
- Identify recurring operational risks and drive engineering improvements that prevent recurrence.
Observability and Incident Management
- Build and maintain monitoring, logging, distributed tracing, alerting and service-health dashboards.
- Lead or support production incident response, service restoration, Root Cause Analysis and corrective actions.
- Develop actionable alerts, operational runbooks and escalation procedures.
- Track reliability trends, service degradation and opportunities to reduce detection and recovery time.
Kubernetes, Cloud and Resilience
- Operate and support production Kubernetes environments, including deployments, upgrades,
scaling and troubleshooting.
- Improve high availability, fault tolerance, backup, recovery and disaster-recovery capabilities.
- Perform capacity planning, performance analysis and infrastructure optimization.
- Support cloud-hosted, customer-hosted and air-gapped deployment models.
Agentic AI Platform Reliability
- Operate and monitor AI Agents, Agentic Workflows, MCP Servers, AI/LLM Gateways, Temporal workflows and RAG services.
- Improve latency, throughput, scalability, resilience and recovery across AI platform components.
- Monitor gateway routing, authentication, quotas, rate limits, failover and service health.
- Build automation and self-healing capabilities to reduce operational toil and improve platform stability.
Mandatory Skills
- Production Kubernetes operations using Helm, Docker and container orchestration.
- SRE practices, including SLIs, SLOs, Error Budgets, production-readiness reviews and reliability engineering.
- Production incident management, Root Cause Analysis, corrective actions and recurrence prevention.
- Observability using Prometheus, Grafana, OpenTelemetry, ELK, OpenSearch or equivalent tools.
- Infrastructure as Code using Terraform, Pulumi or equivalent tooling.
- Production operations on at least one major cloud platform: Azure, AWS or GCP.
- Operational automation and scripting using Python, Go, Bash or equivalent.
- High availability, scalability, capacity planning, backup and disaster recovery.
- Secrets management, certificates,
access controls and operational security.
- Hands-on experience operating or supporting AI Agents and Agentic Workflows.
- Hands-on experience deploying, operating or supporting MCP Servers and MCP-based integrations.
- Hands-on experience operating or supporting AI/LLM Gateways such as LiteLLM, Azure AI Gateway, Azure API Management, Kong AI Gateway or equivalent.
- Experience supporting GenAI, RAG or LLM-powered applications in production-oriented environments.
- Strong troubleshooting skills across infrastructure, application and AI platform layers.
Preferred Skills
- Temporal or another durable workflow orchestration platform.
- GitOps practices using Argo CD, Flux or equivalent.
- AI observability tools such as Langfuse, LangSmith, OpenLIT or Arize.
- Vector databases such as Qdrant, Pinecone, Weaviate, Chroma or pgvector.
- Service mesh technologies such as Istio or Linkerd.
- PostgreSQL or other production database operations.
- Multi-region, customer-hosted or air-gapped deployment experience.
- Experience supporting enterprise SaaS or other business-critical platforms.
Experience and Qualifications
- 6 to 10 years of experience in Site Reliability Engineering, Platform Operations, DevOps, Cloud Engineering or Infrastructure Engineering.
- Demonstrated experience operating business-critical production environments.
- Experience supporting Kubernetes-based platforms at scale.
- Bachelors degree in Computer Science, Engineering or a related field.
- CKA, CKAD, cloud platform or SRE-related certifications are advantageous.
Disclaimer: This job posting and Location has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Sr. Site Reliability Engineer (Hyderabad)
🏢 Rimini Street
📍 Hyderabad