Job Description:
We're looking for an AI Data Engineer to design and build robust data pipelines that power AI/ML and GenAI applications. This role sits at the intersection of traditional data engineering and modern AI infrastructure — you'll be responsible for making data AI-ready, scalable, and production-grade.
Key Roles and Responsibilities
• Design, build, and optimize ETL/ELT pipelines using Python, PySpark, and SQL
• Develop and maintain scalable data pipelines on Databricks (Delta Lake, notebooks, workflows, cluster optimization)
• Build data infrastructure to support AI/ML and GenAI use cases — including feature engineering, embeddings, and vector data pipelines
• Prepare, clean, and structure data (structured & unstructured) for LLM/RAG-based applications
• Collaborate with Data Scientists and ML Engineers to operationalize models and AI pipelines
• Ensure data quality, governance, and performance across pipelines
• Work with cloud platforms (Azure/AWS/GCP) for data storage, compute, and orchestration
• Optimize Spark jobs for performance and cost efficiency
Required Skills
• 3–6 years of experience in Data Engineering
• Strong hands-on expertise in SQL and Python
• Proficiency in PySpark for large-scale data processing
• Working experience with Databricks (Delta Lake, Unity Catalog, notebooks)
• Exposure to AI/ML data pipelines — vector databases, embeddings, or RAG architecture is a plus
• Experience with cloud data platforms (Azure Data Factory, AWS Glue, or equivalent)
• Understanding of data modeling, warehousing, and pipeline orchestration (Airflow/ADF)
• Familiarity with LLM ecosystems (LangChain, LlamaIndex) is an added advantage
Good to Have
• Experience in BFSI or analytics-heavy domains
• Exposure to MLOps tools and CI/CD for data pipelines
• Knowledge of NoSQL/vector databases (Pinecone, FAISS, Chroma)
What We Offer
• Opportunity to work on cutting-edge AI/data infrastructure projects
• Team-oriented, fast-paced environment
• Competitive c
📌 AI Data Engineer (Gurugram)
🏢 EXL
📍 Gurugram