17 Aug
|
Tecnoprism
|
India
Job Description: Data Engineer (Databricks / PySpark / SQL)
Experience
26 years
Role Overview
We are looking for a skilled Data Engineer with solid expertise in Databricks, PySpark, and SQL to design, develop, and optimize scalable data pipelines. The ideal candidate will work closely with data analysts, data scientists, and business stakeholders to deliver high-quality data solutions and insights.
Key Responsibilities
Design, develop, and maintain scalable data pipelines using PySpark on Databricks
Build and optimize large-scale data processing workflows for structured and unstructured data
Develop and maintain SQL queries for data transformation, extraction, and reporting
Work with Delta Lake and Databricks ecosystem for effective data storage and processing
Collaborate with cross-functional teams to gather and understand data requirements
Ensure data quality, integrity, and performance optimization of pipelines
Troubleshoot and debug data-related issues in production settings
Implement best practices for data engineering, including version control,
CI/CD, and testing
Optimize queries and Spark jobs for performance and cost efficiency
Document technical designs, processes, and workflows
Required Skills & Qualifications
Strong experience with Databricks platform (workflows, notebooks, clusters)
Proficiency in PySpark for large-scale data processing
Advanced knowledge of SQL (joins, window functions, query optimization)
Experience working with big data technologies and distributed computing
Hands-on experience with ETL/ELT pipeline development
Understanding of data warehousing concepts and data modeling (Star/Snowflake schema)
Experience with Delta Lake / Lakehouse architecture
Familiarity with cloud platforms (Azure, AWS, or GCP preferred)
Knowledge of data integration tools and scheduling (Airflow, ADF, etc.) is a plus
Preferred Skills
Experience with Python programming beyond PySpark
Exposure to streaming fram
📌 Databricks Engineer Pune (India)
🏢 Tecnoprism
📍 India