04 Sep
|
cirruslabs
|
Bengaluru
04 Sep
cirruslabs
Bengaluru
Senior Backend AI Engineer (7+ Years)
- Agentic AI, RAG, LLM Fine-Tuning Job Title: Senior Backend AI Engineer (Agentic AI, RAG, LLM Engineering)
Experience: 7+ years Role Summary We are seeking a Senior Backend AI Engineer with 7+ years of experience to design, build, and scale production-ready AI systems.
The role has a strong focus on Agentic AI platforms, Retrieval-Augmented Generation (RAG), and full LLM lifecycle ownership including training data preparation, fine-tuning workflows, and deployment/serving on cloud-native infrastructure. Key Responsibilities
- Own and build backend platforms for Agentic AI products, including autonomous workflows, orchestration, tool-calling, state and context handling, and safety guardrails.
- Design, implement, and optimize Retrieval-Augmented Generation (RAG) systems (chunking, embeddings, vector search, reranking, metadata filtering, freshness, and ranking quality evaluation).
- Develop and maintain LLM fine-tuning and training pipelines (dataset curation, preprocessing, labeling strategy, hyperparameter tuning, experiment tracking, and rollback strategy).
- Build scalable inference services for LLM and RAG workloads with high availability, low latency, and strict quality monitoring (hallucination, retrieval precision/recall, tool-call accuracy).
- Develop and operationalize AI APIs and microservices in Dockerized environments, including model serving stacks, vector store integration, and asynchronous event-driven processing.
- Design and manage Kubernetes deployment topologies, autoscaling policies, and disaster recovery for AI services.
- Create and own CI/CD workflows for model and application delivery (training jobs, image builds, infra changes, blue/green or canary releases, and post-deploy verification).
- Implement MLOps practices: experiment tracking, model versioning, canary validation, prompt and policy regression tests, observability, and performance benchmarking.
- Partner with product, data, and infrastructure teams to transform research prototypes into production solutions and drive roadmap decisions.
- Mentor engineers through code reviews, architecture sessions, and AI engineering best practices.
Required Skills and Technologies
- 7+ years of software engineering experience with strong backend ownership.
- Hands-on AI/ML engineering experience with LLM systems in production.
- Deep practical knowledge of Agentic AI frameworks and multi-step workflow engines.
- Robust experience with RAG architecture and vector search technologies (FAISS, Pinecone, Weaviate, pgvector, Milvus).
- Hands-on experience in LLM training and fine-tuning workflows (instruction tuning, domain adaptation, PEFT/LoRA/QLoRA, and parameter-efficient methods).
- Python (primary)
with API frameworks such as FastAPI/Flask and asynchronous programming patterns.
- Containerization and orchestration with Docker and Kubernetes (Helm, ingress, secrets, HPA, resource quotas, monitoring).
- CI/CD ownership using GitHub Actions/GitLab CI/Jenkins and infrastructure-as-code patterns (Terraform/Ansible/Helm).
- Datastores and caching: PostgreSQL, Redis, and object storage, with exposure to NoSQL where needed.
- Cloud fundamentals on AWS/GCP/Azure (compute, container registries, IAM, networking, managed databases).
- Experience with observability stack (Prometheus, Grafana, OpenTelemetry, ELK/Opensearch, alerting).
Preferred / Plus
- Experience with MLflow, DVC, Weights & Biases, or equivalent experiment and dataset lifecycle tooling.
- Experience with serving stacks: Triton, vLLM, TorchServe, Text Generation Inference, BentoML, or equivalent.
- Experience with agent memory stores, retrieval quality benchmarking, and policy/safety layers for autonomous agents.
- Familiarity with Terraform, ArgoCD, or GitOps workflows for AI platform delivery.
- Security/compliance and governance awareness for AI systems (data privacy, prompt-injection controls, and auditability).
What We Expect
- Own complex AI backend features from design to production with measurable impact.
- Design systems for reliability, observability, and cost-aware scaling.
- Drive trade-off decisions across model quality, latency, and infrastructure costs.
- Communicate clearly with engineering, product, and leadership, and mentor other team members.
📌 Backend AI Engineer (Bengaluru)
🏢 cirruslabs
📍 Bengaluru