Lead Data Engineer (Databricks & PySpark) (Pune)

Lead Data Engineer (Databricks & PySpark) (Pune)

31 Jul
|
YO IT CONSULTING
|
Pune

31 Jul

YO IT CONSULTING

Pune

Location- Pune Experience - 8 to 14 years Mandatory Requirements 8-14 years of Data Engineering experience.

4-8 years of hands-on Databricks experience.

Experience modernizing legacy code into PySpark.

Production-scale data pipeline development experience.

Experience working in Agile environments.

Mandatory

Certification- Must Possess At Least One Of The Following Databricks Certified Data Engineer Associate

Databricks Certified Data Engineer Professional Job Summary: We are looking for an experienced Lead Data Engineer to design, develop, and optimize scalable data pipelines on the Databricks platform. The ideal candidate will have strong expertise in PySpark, Databricks, Delta Lake, and modern data engineering practices while leading the modernization of legacy ETL pipelines into cloud-native data solutions.

Essential Technical Skills Data Engineering: Strong foundation in data engineering principles, ETL/ELT processes, and data pipeline design patterns

PySpark: Proven hands-on experience developing data pipelines using PySpark, including DataFrames API, Spark SQL, and performance optimization

Databricks Platform: Practical experience with Databricks workspace, cluster management, notebooks, and job orchestration

Workspace AI Agent: Knowledge of Databricks Workspace AI Agent capabilities and integration

Data Modelling: Experience implementing data models including dimensional modeling, data vault, or lakehouse architectures

Delta Lake: Understanding of Delta Lake features including ACID transactions, schema evolution, and optimization techniques

Python: Solid Python programming skills for data processing and automation





]:pointer-events-auto R6Vx5W_threadScrollVars scroll-mb-[calc(var(--scroll-root-safe-area-inset-bottom,0px)+var(--thread-response-height))] scroll-mt-[calc(var(--header-height)+min(200px,max(70px,20svh)))]" dir="auto" data-turn-id="request-WEB:c31b1b9c-8b90-4850-a24f-ffb6013969b0-0" data-turn-id-container="request-WEB:c31b1b9c-8b90-4850-a24f-ffb6013969b0-0" data-testid="conversation-turn-2" data-turn="assistant"> Roles And Responsibilities- Data Pipeline Development & Operations- Design, build, and operate scalable and reliable data pipelines on the Databricks platform Develop end-to-end data workflows from ingestion through transformation to consumption

Implement robust error handling, monitoring, and alerting mechanisms

Ensure data pipeline reliability, performance, and maintainability

Optimize pipeline performance through efficient Spark job design and cluster configuration

Manage and orchestrate complex data workflows using Databricks Jobs and workflows Legacy Code Modernization Refactor legacy code and data pipelines to PySpark for improved performance and scalability

Migrate traditional ETL processes to modern ELT patterns on Databricks

Assess existing codebases and identify opportunities for optimization and modernization

Ensure backward compatibility and data integrity during migration processes

Document refactoring approaches and create migration playbooks





Collaborate with stakeholders to minimize disruption during code transitions Data Engineering Excellence Implement data quality checks and validation frameworks

Design and maintain Delta Lake tables with appropriate optimization strategies

Develop reusable code libraries and frameworks for common data engineering tasks

Follow software engineering best practices including version control, testing, and CI/CD

Participate in code reviews and provide constructive feedback to team members

Troubleshoot and resolve data pipeline issues in production environments Collaboration & Knowledge Sharing Work closely with data architects, analysts, and business stakeholders

Collaborate with Infrastructure (Infra), Applications (Apps), and Cyber teams

Share knowledge and best practices with Team NCS

Mentor junior data engineers on PySpark and Databricks technologies

Document technical solutions and maintain comprehensive documentation Additional Technical Skills SQL proficiency for data querying and transformation

Experience with cloud platforms (Azure, AWS, or GCP)

Understanding of data governance and security best practices

Knowledge of streaming data processing (Structured Streaming)

Familiarity with DevOps practices and CI/CD pipelines

Experience with version control systems (Git)

Understanding of data quality frameworks and testing methodologies Additional Certifications (Preferred) Databricks Certified Associate Developer for Apache Spark

Cloud platform certifications (Azure Data Engineer Associate, AWS Certified Data Analytics, or Google Cloud Professional Data Engineer)

Relevant data engineering or big data certifications

📌 Lead Data Engineer (Databricks & PySpark) (Pune)
🏢 YO IT CONSULTING
📍 Pune

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead data engineer (databricks & pyspark) (pune) / pune

Subscribe to this job alert:

Get the latest job offers by email for: lead data engineer (databricks & pyspark) (pune) / pune