Job Description
Job Description:
n
nWe're looking for an AI Data Engineer to design and build robust data pipelines that power AI/ML and GenAI applications. This role sits at the intersection of traditional data engineering and modern AI infrastructure — you'll be responsible for making data AI-ready, scalable, and production-grade.
n
nKey Roles and Responsibilities
n
- n
- Design, build, and optimize ETL/ELT pipelines using Python, PySpark, and SQLn
- Develop and maintain scalable data pipelines on Databricks (Delta Lake, notebooks, workflows, cluster optimization)n
- Build data infrastructure to support AI/ML and GenAI use cases — including feature engineering, embeddings, and vector data pipelinesn
- Prepare, clean, and structure data (structured & unstructured) for LLM/RAG-based applicationsn
- Collaborate with Data Scientists and ML Engineers to operationalize models and AI pipelinesn
- Ensure data quality, governance, and performance across pipelinesn
- Work with cloud platforms (Azure/AWS/GCP) for data storage, compute, and orchestrationn
- Optimize Spark jobs for performance and cost efficiencyn
nRequired Skills
n
- n
- 3–6 years of experience in Data Engineeringn
- Solid hands-on expertise in SQL and Pythonn
- Proficiency in PySpark for large-scale data processingn
- Working experience with Databricks (Delta Lake, Unity Catalog, notebooks)n
- Exposure to AI/ML data pipelines — vector databases, embeddings, or RAG architecture is a plusn
- Experience with cloud data platforms (Azure Data Factory, AWS Glue, or equivalent)n
- Understanding of data modeling, warehousing, and pipeline orchestration (Airflow/ADF)n
- Familiarity with LLM ecosystems (LangChain, LlamaIndex) is an added advantagen
nGood to Have
n
- n
- Experience in BFSI or analytics-heavy domainsn
- Exposure to MLOps tools and CI/CD for data pipelinesn
- Knowledge of NoSQL/vector databases (Pinecone, FAISS, Chroma)n
nWhat We Offer
n
- n
- Opportunity to work o
📌 AI Data Engineer (Gurugram)
🏢 EXL
📍 Gurugram