Must Needed Skills: Apache Airflow, Python, PySpark ,Hive(Trino), SQL and
Snowflake
* Develop and maintain data pipelines, ELT processes, and workflow
orchestration using Apache Airflow, Python, PySpark ,Hive(Trino) and
Snowflake to ensure the effective and reliable delivery of data.
* Design and implement custom connectors to facilitate the ingestion of diverse
data sources into our platform, including structured and unstructured data
from various document formats .
* Collaborate closely with cross-functional teams to gather requirements,
understand data needs, and translate them into technical solutions.
* Design and implement data CI/CD pipelines to enable automated and efficient
data integration, transformation, and deployment processes.
* Monitor and troubleshoot data pipelines, proactively identifying and
resolving issues related to data ingestion, transformation, and loading.
* Conduct data validation and testing to ensure the accuracy, consistency, and
compliance of data.
* Stay up-to-date with emerging technologies and best practices in data
engineering.
* Document data workflows, processes, and technical specifications to
facilitate knowledge sharing and ensure data governance.
📌 Data Engineer (India)
🏢 Photon
📍 India