16 Aug
|
Insight Advisors
|
India
16 Aug
Insight Advisors
India
Designation : Data Engineer
Location : Pune, Maharashtra, India
Qualification : B.E./B.Tech, M.E./M.Tech, MCA, or equivalent in Computer Science/IT or a related field.
Work Mode : Hybrid
Primary Skill Set : Databricks, Python, Spark, SQL
Data Engineer:
Data Pipeline Development & Operations:
- Design, build, and operate scalable and reliable data pipelines on the Databricks platform.
- Develop end-to-end data workflows from ingestion through transformation to consumption.
- Implement error handling, monitoring, and alerting mechanisms.
- Ensure pipeline reliability, performance, and maintainability.
- Optimize Spark jobs and cluster configurations for improved performance.
- Manage complex workflows using Databricks Jobs and Workflows.
Legacy Code Modernization:
- Refactor legacy code and ETL pipelines to PySpark.
- Migrate traditional ETL processes to modern ELT patterns on Databricks.
- Identify opportunities for code and pipeline optimization.
- Ensure data integrity and backward compatibility during migrations.
- Document modernization approaches and migration playbooks.
Data Engineering Excellence:
- Implement data quality checks and validation frameworks.
- Design and maintain Delta Lake tables with appropriate optimization strategies.
- Develop reusable data engineering frameworks and libraries.
- Follow best practices for version control, testing and CI/CD.
- Participate in code reviews and troubleshoot production issues.
Collaboration & Knowledge Sharing:
- Work closely with data architects, analysts and business stakeholders.
- Collaborate with Infrastructure, Applications and Cyber teams.
- Share technical knowledge and best practices.
- Mentor junior Data Engineers on PySpark and Databricks.
- Maintain comprehensive technical documentation.
Essential Technical Skills:
- Robust Data Engineering, ETL/ELT and data pipeline design knowledge.
- Hands-on PySpark,
including DataFrames, Spark SQL and performance optimization.
- Practical experience with Databricks Workspace, Cluster Management, Notebooks and Job Orchestration.
- Knowledge of Databricks Workspace AI Agent capabilities and integration.
- Data Modelling experience Dimensional Modeling, Data Vault or Lakehouse Architecture.
- Strong understanding of Delta Lake, including ACID Transactions, Schema Evolution and optimization.
- Strong Python skills for data processing and automation.
Additional Technical Skills:
- Strong SQL proficiency.
- Cloud experience Azure, AWS or GCP.
- Data Governance and Security understanding.
- Knowledge of Structured Streaming.
- DevOps and CI/CD practices.
- Version control using Git.
- Data Quality and Testing methodologies.
Professional Experience:
- Minimum 5 years relevant Data Engineering experience.
- At least 23 years hands-on Databricks experience.
- Proven experience in legacy ETL/code modernization to modern frameworks.
- Experience building and maintaining production-scale data pipelines.
- Experience working with multiple data sources and formats.
- Experience in Agile development environments.
Required Certification Mandatory :
At least one of the following:
- Databricks Certified Data Engineer Associate
- Databricks Certified Data Engineer Professional
Preferred Certifications :
- Databricks Certified Associate Developer for Apache Spark.
- Azure Data Engineer Associate.
- AWS Certified Data Analytics.
- Google Cloud Professional Data Engineer.
- Other relevant Data Engineering / Big Data certifications.
Soft Skills :
- Strong problem-solving and analytical skills.
- Excellent communication and technical explanation skills.
- Strong collaboration and stakeholder management.
- Self-motivated with attention to detail.
- Adaptable to changing priorities and technologies.
- Client-focused approach with commitment to quality delivery.
📌 Data Engineer - ETL/PySpark (India)
🏢 Insight Advisors
📍 India