Key Responsibilities Data Engineering & Pipeline Development Snowflake and AWS services
Build, test, and maintain scalable batch and real-time data pipelines using Snowflake, PySpark, Delta Lake, and Kafka.
Implement data quality checks, schema validation, monitoring, and alerting.
Modernize legacy ETL/DWH platforms and migrate them to cloud-native AWS architectures.
Manage CI/CD pipelines, automated deployments, rollback strategies, and Infrastructure as Code (Terraform, GitHub Actions).
RAG & Retrieval Infrastructure
Develop end-to-end retrieval systems including document ingestion, embedding pipelines, vector databases, and hybrid search solutions.
Optimize chunking, metadata filtering, reranking, precision, recall, and latency.
Ensure data freshness, index consistency, and retrieval quality monitoring.
Semantic & Knowledge Infrastructure
Build and maintain semantic models, business ontologies, entity mappings, and knowledge graphs (Neo4j).
Manage feature stores, metadata, data lineage, and semantic data contracts for AI/ML applications.
ML/LLMOps
Support ML and LLM workflows including feature engineering, dataset versioning, MLflow tracking, and model evaluation.
Build automated evaluation pipelines for hallucination detection, factual accuracy, and regression testing.
Monitor production systems, pipelines, and model performance.
Agentic AI Infrastructure
Develop APIs,
tool schemas, memory stores, and context services for AI agents.
Implement observability for agent interactions, tool usage, retrieved context, and reasoning traces.
Support text-to-SQL and conversational AI capabilities.
Governance & Security
Implement RBAC, PII masking, audit logging, data classification, and access controls.
Establish data quality monitoring, schema governance, and compliance-ready audit trails.
Required Qualifications
7+ years of Data Engineering experience on AWS/Azure cloud platforms.
2+ years of hands-on experience building AI/ML or LLM-era data infrastructure.
Proven expertise in large-scale batch and streaming data pipelines.
Solid understanding of data governance, security, compliance, and data quality frameworks.
Experience working with RAG architectures, vector databases, and retrieval systems in production environments.
Technical Skills
Primary Skills
Python, SQL, PySpark
Kafka, Snowflake, Databricks, Delta Lake
AWS (S3, Glue, Kinesis, EKS, Redshift)
Docker, Kubernetes
GitHub Actions
Gen AI, LangChain, LangGraph, RAG
Secondary Skills
LangChain, LlamaIndex
OpenAI, Bedrock, Claude, Hugging Face APIs
Pinecone, FAISS, ChromaDB, OpenSearch
MLflow, FastAPI, Neo4j
Prompt Engineering
RLHF Dataset Preparation
LLM Fine-Tuning Workflows
📌 Senior Data Engineer -AI/LLM (Delhi)
🏢 Premium
📍 Delhi