Experience building enterprise GenAI or AI Agent solutions. Exposure to DBRX foundation models and Databricks AI platform. Hands-on experience with vector search and semantic retrieval systems. Knowledge of DataOps, MLOps, and LLMOps best practices. Experience with multi-cloud data architectures. Understanding of Responsible AI and model governance frameworks.
Data Engineering & AI Platform Development Design and develop scalable batch and real-time data pipelines using Databricks, Spark, and cloud-native services. Build and maintain enterprise-grade data lakes, lakehouses, and data warehouses. Create ingestion frameworks for structured, semi-structured, and unstructured data. Develop AI-ready data pipelines supporting LLM, RAG, Agentic AI, and predictive analytics use cases. Implement metadata management, data governance, lineage, and data quality controls. AI & Generative AI Enablement Prepare and transform data for LLM training, fine-tuning, and inference workloads. Build Retrieval-Augmented Generation (RAG)
pipelines integrating vector databases and enterprise knowledge sources. Develop data workflows supporting DBRX and other foundation models. Optimize AI data pipelines for performance, scalability, and cost efficiency. Cloud & Data Platform Engineering Design solutions on AWS, Azure, or GCP settings. Leverage cloud-native data services for ingestion, storage, orchestration, and monitoring. Implement CI/CD and Infrastructure as Code (IaC) practices for data platforms. Ensure security, compliance, and governance across cloud data ecosystems.
Bachelor's or Master's degree in Computer Science, Data Engineering, Information Technology, Artificial Intelligence, or related field. Key Skillset - Databricks, DBRX, PySpark, Spark, Python, Snowflake, AWS, Azure, GCP, Delta Lake, Unity Catalog, MLflow, RAG, LangChain, Vector Database, Pinecone, Feature Store, Airflow, CI/CD, Docker, Kubernetes, Data Lakehouse, GenAI, LLM, Data Engineering.
📌 Ai Data Engineers Noida
🏢 Infosys
📍 Noida