Position Description
Data Engineer
Mission:
• Analyze the business and technical needs of data/AI use cases in collaboration with Product Owners and Data Scientists
• Design and implement robust, scalable, and industrialize data pipelines (ingestion, transformation, exposure)
• Ensure data quality, monitoring and traceability
• Implement and optimize data architectures (batch and real time) according to performance and volume constraints
• Industrialize data flows to support analytical, ML and generative AI use cases (RAG, feature pipelines, etc.)
• Manage data storage and access (data lakes, warehouses, vector databases)
• Collaborate closely with Data Science / MLOps teams to facilitate the deployment of models in production
• Ensure data security and regulatory compliance
• Document pipelines and architectures to ensure maintainability and knowledge sharing
• Contribute to the continuous improvement of Data Engineering practices and the standardization of tools and frameworks
Skills:
Must have :
• Minimum 7 years of experience in Data Engineering with large scale data pipeline production
• Solid skills in designing distributed data architectures
• Expertise in: Data Engineering : ETL/ELT, data pipelines, data quality, orchestration Big Data : Spark / PySpark, Azure Data Lake, Databricks Stockage : Data Lake, Data Warehouse (Databricks) Databases: Advanced SQL Streaming: Kafka, Event Hub or similar
• Experience on AI/ML use cases with data exposure for models
• GenAI (data interactions): understanding of RAG architectures, management of vector DBs
• Programming: Python, SQL
• Cloud : Azure (Data Factory, Synapse, Data Lake, Event Hub…) ou équivalent
• DevOps & DataOps : Git, CI/CD, Docker, orchestration (Airflow, AWX)
• Best practice of agile methodologies
• Language: French (C1+) and English (B2+)
Nice to have :
• Experience with vector databases (PGVector, Azure AI Search, Weaviate...)
• Connaissance des outils MLOps (MLflow, feature stores)
• Experienc
📌 Lead Data Engineer (Bengaluru)
🏢 CGI
📍 Bengaluru