07 Aug
|
Mindpool Technologies
|
Bengaluru
07 Aug
Mindpool Technologies
Bengaluru
Top Skills: Kubernetes, Docker, Python/Bash Scripting, and IAC (Infrastructure as Code),Distributed Tracing, Observability and Monitoring, Troubleshooting in K8s,Cloud,Microservices and Good to Have Generative AI Experience
Job Description:
We are seeking a SRE engineering capabilities to ensure the reliability, scalability, performance, and operational excellence of AI powered and full stack web platforms.
This role blends core SRE fundamentals (cloud infrastructure, Kubernetes, CI/CD, Linux, networking, incident management) with hands on GenAI development and platform support. You will not only operate AI systems in production, but also build, integrate, debug, and harden GenAI workflows in close collaboration with Machine Learning Engineers and Full Stack Engineers.
Key Responsibilities
Core SRE & Platform Reliability
• Engineer and operate highly available, fault tolerant production platforms.
• Own SLIs, SLOs, and error budgets for critical services.
• Design and implement resilience, disaster recovery, and capacity planning strategies.
• Identify reliability risks through monitoring, load testing, and failure mode analysis.
• Influence architecture to embed reliability, operability, and scalability by design.
Incident Management & Operational Excellence
• Act as a senior escalation point for complex production incidents.
• Lead incident response across infrastructure, application, and GenAI layers.
• Conduct blameless post incident reviews with strong corrective and preventive actions.
• Reduce MTTR, alert fatigue, and repeat incidents through engineering improvements.
• Own and continuously improve runbooks, playbooks, and on call practices.
GenAI Platform Engineering & Support
• Actively develop, integrate, and support GenAI workflows used in production systems.
• Build and maintain GenAI components using:
o LangChain (chains, tools, agents)
o LangGraph (agent workflows)
o LLM APIs (OpenAI, Gemini)
• Partner with MLEs and application teams to:
o Integrate LLMs into backend services and user workflows
o Debug and optimise prompts, agent execution, and tool usage
• Implement guardrails for protected and predictable LLM behaviour, including:
o Inference parameter governance (temperature, top k, top p)
o Rate limiting, retries, fallbacks, and circuit breakers
• Operationalise GenAI constraints around latency, concurrency, token limits, and cost.
• Add reliability controls directly into GenAI code paths.
Infrastructure, Cloud & Kubernetes (GCP)
• Design, operate, and scale Google Cloud Platform (GCP) infrastructure (mandatory).
• Manage Kubernetes based workloads, including:
o Resource tuning and autoscaling
o Rolling and zero downtime deployments
• Build and manage Infrastructure as Code using Terraform or equivalent tools.
• Partner with Platform teams to design secure, scalable AI platform architectures on GCP.
Full Stack & API Reliability
• Ensure reliability and scalability of:
o Python FastAPI backend services
o React / Next.js frontend applications (SSR & CSR)
• Design and enforce API reliability patterns:
o Timeouts, retries, rate limiting, and back pressure
o Graceful degradation and dependency isolation
• Support GenAI powered APIs consumed by frontend and downstream systems.
Observability & Monitoring
• Build end to end observability across infrastructure, applications, and GenAI systems.
• Monitor and troubleshoot:
o API latency and error rates
o Kubernetes and infrastructure health
o LLM success rates, token usage, and cost
o Frontend availability and performance
• Create dashboards and alerts that are actionable and low noise.
CI/CD, Automation & Cost Efficiency
• Design and operate reliable CI/CD pipelines for backend, frontend, and GenAI services.
• Implement canary deployments, automated rollbacks, and health checks.
• Develop automation tools and scripts (primarily in Python).
• Drive cost optimisation, especially for LLM inference and cloud resource usage.
Security & Responsible AI Operations
• Embed security, access control, and compliance into platform and GenAI designs.
• Ensure responsible and production safe AI usage, including secure handling of data, prompts, and responses.
• Partner with security teams on vulnerability remediation and audits.
Required Skills & Experience
• 6+ years of experience in SRE, Platform, or DevOps Engineering roles.
• Hands on experience with Google Cloud Platform (GCP) in production.
• Strong foundation in:
o Linux systems
o Networking (TCP/IP, DNS, load balancing)
o Distributed systems and failure modes
• Hands on experience with:
o Kubernetes and containerized workloads
o Terraform or similar IaC tools
o CI/CD pipelines and release automation
• Strong Python skills for:
o Automation and internal tooling
o Backend and GenAI system development
• Experience building or supporting GenAI / LLM based systems in production.
• Willingness to work in rotational shifts and participate in on call support.
Nice to Have
• Production experience with LangChain or LangGraph
• Familiarity with Vertex AI
• Experience with load testing, chaos engineering, or performance tuning
What Youll Own
• Production reliability of GenAI powered and full stack platforms
• GenAI workflows that are stable, operable, and cost controlled
• Incident response quality and operational maturity
• Platform automation, observability, and efficiency improvements Role & responsibilities
Preferred candidate profile
Perks and benefits
📌 SRE (Bengaluru)
🏢 Mindpool Technologies
📍 Bengaluru