25 Aug
|
cirruslabs
|
Bengaluru
25 Aug
cirruslabs
Bengaluru
Job Title :Backend AI Engineer
Location : All over India
Notice period : Immediate joiner
Role Overview
- Job Title: Senior Backend AI Engineer (Agentic AI, RAG, LLM Engineering)
- Experience Required: 7+ Years
- Core Focus: Production-ready Agentic AI platforms, Retrieval-Augmented Generation (RAG), and full-lifecycle LLM engineering (data prep, fine-tuning, inference, and cloud deployment).
Key Responsibilities
- Agentic Platforms: Build backend platforms for autonomous workflows, orchestration, tool-calling, state/context management, and safety guardrails.
- RAG Systems: Design and optimize end-to-end RAG pipelines (chunking, embeddings, vector search, metadata filtering, reranking, and retrieval evaluation).
- LLM Fine-Tuning: Develop training pipelines, covering dataset curation, preprocessing, labeling strategies, hyperparameter tuning, and rollback mechanisms.
- Inference & Serving: Deploy high-availability, low-latency inference services with monitoring for hallucinations, retrieval precision/recall, and tool-call accuracy.
- Microservices & APIs: Build containerized AI APIs and asynchronous, event-driven microservices integrated with vector stores and model-serving stacks.
- Infrastructure & Kubernetes: Architect K8s deployment topologies, autoscaling policies (HPA), and disaster recovery plans for AI workloads.
- CI/CD & MLOps: Own end-to-end CI/CD and MLOps practices, including blue/green or canary releases, model versioning, prompt regression tests, and benchmarking.
- Collaboration & Mentorship:
Bridge research prototypes to production alongside product and infra teams; mentor engineers via design reviews and coding standards.
Required Skills & Technologies
- Experience: 7+ years of backend engineering with proven production LLM/AI deployment experience.
- AI/LLM Core: Deep knowledge of Agentic AI frameworks, multi-step orchestration, and fine-tuning methods (PEFT, LoRA, QLoRA, instruction tuning).
- Vector Databases: Hands-on experience with FAISS, Pinecone, Weaviate, pgvector, or Milvus.
- Programming & APIs: Advanced Python, asynchronous programming, and API frameworks (FastAPI, Flask).
- Containers & Orchestration: Docker, Kubernetes (Helm, ingress, secrets, HPA, resource quotas).
- Datastores & Caching: PostgreSQL, Redis, object storage (S3/GCS), and NoSQL databases.
- DevOps & Cloud: CI/CD tools (GitHub Actions, GitLab CI, Jenkins), IaC basics, and core cloud infrastructure (AWS, GCP, or Azure).
- Observability: Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, and alerting pipelines.
Preferred Qualifications
- Serving Engines: Triton Inference Server, vLLM, TGI, TorchServe, or BentoML.
- ML Lifecycle Tools: MLflow, Weights & Biases, DVC, or equivalent experiment-tracking tools.
- GitOps: ArgoCD, Terraform, and declarative infrastructure workflows.
- AI Security: Knowledge of prompt-injection defense, data privacy, guardrail policies, and AI auditability.
Core Expectations
- Deliver scalable, highly reliable, and cost-effective AI backend architectures.
- Effectively balance trade-offs between model accuracy, latency, and compute cost.
- Drive technical roadmap decisions and communicate clearly across cross-functional teams.
📌 Backend AI Engineer (Bengaluru)
🏢 cirruslabs
📍 Bengaluru