11 Sep
|
Infinityquest It Services
|
Hyderabad
11 Sep
Infinityquest It Services
Hyderabad
AI Engineer
Role summary
Were hiring an AI Engineer to design and build production-grade GenAI and ML services on GCP. You’ll own core platform components for hybrid RAG, vector + graph data stores, multi-agent orchestration, and secure tool connectivity (including MCP-style patterns). This is a purely technical role with strong ownership from design through running.
Key responsibilities
Build backend GenAI/ML services (GCP-first)
- Design and implement scalable API-first services for LLM applications (chat/search/assistant capabilities).
- Build low-latency, high-throughput systems with robust caching, batching, retries, idempotency, and fallbacks.
- Integrate with GCP services (e.g., Cloud Run/GKE, Pub/Sub, Cloud Storage, Secret Manager, Cloud Logging/Monitoring).
Multi-agent orchestration (ADK-style frameworks)
- Implement agent orchestration patterns: planner–executor, tool-routing, specialist agents, critic/reviewer loops.
- Build state management, memory, and workflow execution with clear traceability and deterministic behavior where needed.
- Create reusable agent/tool SDK components for other engineering teams.
Hybrid RAG pipeline (dense + sparse)
- Build a hybrid retrieval layer (embeddings + keyword/BM25), reranking, metadata filters, query rewriting, and context compression.
- Implement ingestion pipelines: chunking strategies, deduplication, metadata enrichment, and incremental re-indexing.
- Ensure response grounding with citations, provenance, and strict access controls.
Vector DB + Graph DB platform components
- Own vector indexing strategy, tenant isolation, lifecycle/retention policies, and performance tuning.
- Build graph-backed retrieval (GraphRAG / relationship-aware search) using entity linking and multi-hop traversal.
- Design data models and services that combine vector similarity + graph context for better precision/recall.
MCP-style tool connectivity & enterprise integration
- Implement secure, standardised tool connectors (MCP-style) to internal APIs/data sources.
- Enforce authentication/authorisation, rate limits, audit logs, and policy checks per tool invocation.
- Provide a registry/catalogue of tools with versioned schemas and compatibility guarantees.
LLMOps/MLOps: reliability, evaluation, observability
- Build automated evaluation harnesses (golden sets, regression tests, red-teaming, retrieval metrics).
- Implement production observability: tracing across agent steps, prompt/version tracking, cost and latency dashboards.
- Harden services with SLOs, incident runbooks, and protected rollout patterns (canary/blue green).
Responsible AI & secure engineering
- Implement mechanism against prompt injection, data exfiltration, and unsafe tool usage.
- Apply privacy-by-design: PII handling, data minimization, encryption, and auditability.
- Produce clear design docs (HLD/LLD), run technical reviews, and own operational readiness.
Required experience & skills
- 8–10 years building backend systems (microservices, APIs, distributed systems) and/or ML engineering in production.
- Strong Python (preferred for LLM/RAG) and solid backend engineering fundamentals (async, concurrency, profiling).
- Experience deploying on GCP (Cloud Run or GKE), CI/CD, Terraform/IaC, containerization (Docker).
- Hands-on with LLM integration: embeddings, tool/function calling, structured outputs, prompt/version management.
- Proven delivery of RAG systems, ideally hybrid (dense + sparse) with reranking.
- Experience with vector databases (indexing, performance tuning, multi-tenancy) and embedding pipelines.
- Experience with graph databases and modelling relationships for retrieval/reasoning use-cases.
- Experience implementing agentic/multi-agent workflows using an orchestration framework (ADK-style or equivalent).
- Strong security mindset: secrets management, IAM, auditing, least privilege, threat modelling.
📌 Artificial Intelligence Engineer(GCP) (Hyderabad)
🏢 Infinityquest It Services
📍 Hyderabad