Experience: 4+ Years Role Summary: Develop and maintain scalable data engineering solutions on Azure using Python, PySpark, and Databricks, with foundational exposure to GenAI-enabled data platforms.
Key Responsibilities:
- Design, develop, and support data pipelines using Azure Databricks, Apache Spark (PySpark), and Delta Lake.
- Ingest, transform, and integrate structured and semi-structured data into ADLS.
- Implement Python-based data processing logic and optimize Spark jobs for performance and reliability.
- Ensure data quality through validation, reconciliation, and adherence to data governance standards.
- Support analytical and reporting use cases by delivering clean, reliable datasets.
- Assist in building and maintaining vector-ready datasets to support semantic search or GenAI use cases.
- Collaborate with senior engineers to support GenAI-enabled data workflows.
- Document pipelines,
transformations, and operational processes.
Required Skills:
- Strong hands-on experience with Python, PySpark, Apache Spark.
- Experience working with Azure Databricks and ADLS.
- Good understanding of Delta Lake, ETL concepts, and data modeling.
- Familiarity with REST APIs and basic Azure services (Key Vault, security fundamentals).
- Foundational understanding of GenAI concepts, such as embeddings, vector search, or RAG patterns.
Good to Have:
- Exposure to vector databases (Azure AI Search, LanceDB).
- Hands-on or conceptual exposure to Azure OpenAI / embedding models.
- Experience with Databricks Vector Search.
- Basic understanding of microservices and API-based architectures.
- Interest in GenAI strategy and up-to-date data platform evolution.
📌 Hiring For Data Engineer (Bengaluru)
🏢 Cognizant
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.