10 Oct
|
ClearRoute
|
Pune
About Us:
ClearRoute is an engineering consultancy bridging Quality Engineering, Cloud Platforms and Developer Experience. We help enterprises reliably bring high-impact digital products to market faster, cheaper, and safer, working with technology leaders facing complex business challenges.
We take as much pride in our people, culture and work-life balance as we do in making better software. Were not just making better software.
Were making the making of software better. Collaborative, entrepreneurial and dedicated to problem solving, we bring the step change our customer need to sustain innovation. Our values challenge us to do the best we can for ClearRoute, our customers and most importantly our team. This is an opportunity for you to build the organisation from the ground up, use your voice to drive change and help transform organisations and problem domains.
Role:
- Monitor the health, latency, availability and error rates of AI services, model endpoints and supporting infrastructure.
- Triage, investigate and resolve incidents, from failed API calls and rate limiting to degraded model output quality.
- Run root-cause analysis and write clear, blameless post-incident reviews with follow-up actions.
- Take part in a shared on-call rota for production AI services. Platform reliability and observability
- Build and maintain dashboards, alerts and logging for LLM workloads, including token usage, latency, cost and evaluation metrics.
- Maintain and improve deployment pipelines for models, prompts, agents and RAG components.
- Manage vector databases, embedding pipelines and data ingestion jobs.
- Automate repetitive support tasks and runbooks.
Key Responsibilities:
- Act as the escalation point for AI platform issues raised by delivery teams and client users.
- Write and maintain runbooks,
knowledge-base articles and user guidance.
- Feed recurring issues and improvement ideas back to engineering and product owners.
- Help onboard new teams and clients onto the platform.
Required Experience:
- 3+ years in platform support, site reliability engineering, DevOps or production support for cloud-hosted services.
- Hands-on experience supporting applications built on LLM APIs (for example OpenAI, Anthropic, Azure OpenAI, AWS Bedrock or Google Vertex AI).
- Strong working knowledge of at least one major cloud (AWS, Azure or GCP).
- Comfortable with containers and orchestration (Docker, Kubernetes). • Scripting and automation in Python plus Bash or PowerShell.
- Experience with observability tooling such as Datadog, Grafana, Prometheus, CloudWatch or Azure Monitor.
- CI/CD and infrastructure-as-code experience (for example GitHub Actions, GitLab CI, Terraform).
- Solid grasp of REST APIs, authentication, networking basics and log analysis.
- Proven incident management experience, including writing post-incident reviews.
- Clear written and verbal communication with both technical and non-technical aud
Desirable Experience:
- Experience with RAG architectures and vector databases (for example Pinecone, Weaviate, pgvector, Azure AI Search).
- Familiarity with agent and orchestration frameworks such as LangChain, LlamaIndex or the Model Context Protocol (MCP).
- LLM observability and evaluation tools such as LangSmith, Langfuse, Arize or Weights & Biases.
- Experience hosting or serving open-weight models (for example vLLM, Hugging Face, SageMaker).
- Awareness of AI governance and regulation, including the EU AI Act, ISO/IEC 42001 and GDPR.
- ITIL, Kubernetes (CKA/CKAD) or cloud certifications. •
- Experience in a consultancy or client-facing delivery setting, ideally in regulated sectors such as energy or financial services.
📌 AI Platform Support Engineer (Pune)
🏢 ClearRoute
📍 Pune