14 Aug
|
YO IT CONSULTING
|
Pune
14 Aug
YO IT CONSULTING
Pune
Location- Pune Experience - 8 to 14 years
Mandatory Requirements
- 8-14 years of Data Engineering experience.
- 4-8 years of hands-on Databricks experience.
- Experience modernizing legacy code into PySpark.
- Production-scale data pipeline development experience.
- Experience working in Agile environments.
Mandatory Certification- Must Possess At Least One Of The Following
- Databricks Certified Data Engineer Associate
- Databricks Certified Data Engineer Professional
Job Summary: We are looking for an experienced Lead Data Engineer to design, develop, and optimize scalable data pipelines on the Databricks platform. The ideal candidate will have strong expertise in PySpark , Databricks , Delta Lake , and modern data engineering practices while leading the modernization of legacy ETL pipelines into cloud-native data solutions.
Essential Technical Skills
- Data Engineering: Solid foundation in data engineering principles, ETL/ELT processes, and data pipeline design patterns
- PySpark: Proven hands-on experience developing data pipelines using PySpark, including DataFrames API, Spark SQL, and performance optimization
- Databricks Platform: Practical experience with Databricks workspace, cluster management, notebooks, and job orchestration
- Workspace AI Agent: Knowledge of Databricks Workspace AI Agent capabilities and integration
- Data Modelling: Experience implementing data models including dimensional modeling, data vault, or lakehouse architectures
- Delta Lake: Understanding of Delta Lake features including ACID transactions, schema evolution, and optimization techniques
- Python: Strong Python programming skills for data processing and automation
- ]:pointer-events-auto R6Vx5W_threadScrollVars scroll-mb-[calc(var(--scroll-root-safe-area-inset-bottom,0px)+var(--thread-response-height))] scroll-mt-[calc(var(--header-height)+min(200px,max(70px,20svh)))]" dir="auto" data-turn-id="request-WEB:c31b1b9c-8b90-4850-a24f-ffb6013969b0-0" data-turn-id-container="request-WEB:c31b1b9c-8b90-4850-a24f-ffb6013969b0-0" data-testid="conversation-turn-2" data-turn="assistant">
Roles And Responsibilities- Data Pipeline Development & Operations- Design, build, and operate scalable and reliable data pipelines on the Databricks platform
- Develop end-to-end data workflows from ingestion through transformation to consumption
- Implement robust error handling, monitoring, and alerting mechanisms
- Ensure data pipeline reliability, performance, and maintainability
- Optimize pipeline performance through efficient Spark job design and cluster configuration
- Manage and orchestrate complex data workflows using Databricks Jobs and workflows
Legacy Code Modernization
- Refactor legacy code and data pipelines to PySpark for improved performance and scalability
- Migrate traditional ETL processes to modern ELT patterns on Databricks
- Assess existing codebases and identify opportunities for optimization and modernization
- Ensure backward compatibility and data integrity during migration processes
- Document refactoring approaches and create migration playbooks
- Collaborate with stakeholders to minimize disruption during code transitions
Data Engineering Excellence
- Implement data quality checks and validation frameworks
- Design and maintain Delta Lake tables with appropriate optimization strategies
- Develop reusable code libraries and frameworks for common data engineering tasks
- Follow software engineering best practices including version control, testing, and CI/CD
- Participate in code reviews and provide constructive feedback to team members
- Troubleshoot and resolve data pipeline issues in production environments
Collaboration & Knowledge Sharing
- Work closely with data architects, analysts, and business stakeholders
- Collaborate with Infrastructure (Infra), Applications (Apps), and Cyber teams
- Share knowledge and best practices with Team NCS
- Mentor junior data engineers on PySpark and Databricks technologies
- Document technical solutions and maintain comprehensive documentation
Additional Technical Skills
- SQL proficiency for data querying and transformation
- Experience with cloud platforms (Azure, AWS, or GCP)
- Understanding of data governance and security best practices
- Knowledge of streaming data processing (Structured Streaming)
- Familiarity with DevOps practices and CI/CD pipelines
- Experience with version control systems (Git)
- Understanding of data quality frameworks and testing methodologies
Additional Certifications (Preferred)
- Databricks Certified Associate Developer for Apache Spark
- Cloud platform certifications (Azure Data Engineer Associate, AWS Certified Data Analytics, or Google Cloud Professional Data Engineer)
- Relevant data engineering or big data certifications
📌 Lead Data Engineer (Databricks & PySpark) (Pune)
🏢 YO IT CONSULTING
📍 Pune