06 Sep
|
TalentOla
|
Pune
Data Engineer – Databricks, Spark, Python & ETL
Experience: 5–8 Years
Location: [Location]
Employment Type: Full-Time
Job Summary
We are looking for a skilled Data Engineer with strong expertise in Databricks, Apache Spark, Python, SQL, and ETL development. The ideal candidate should have experience building scalable data pipelines, optimizing big data processing workflows, and working with orchestration and CI/CD tools in cloud-based environments.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Databricks and Apache Spark.
- Develop data processing applications using Python and PySpark.
- Build and optimize complex SQL queries, stored procedures, and transformations.
- Work with structured and unstructured datasets for large-scale data processing.
- Implement data integration workflows and orchestration using tools such as Airflow, Azure Data Factory (ADF), or Autosys.
- Develop reusable frameworks and automate data engineering workflows.
- Monitor and troubleshoot production data pipelines and resolve performance bottlenecks.
- Collaborate with business analysts, data scientists, and cross-functional teams for data requirements.
- Implement CI/CD pipelines for deployment automation and version control.
- Ensure data quality, governance, security, and compliance standards are followed.
- Participate in code reviews, testing, and documentation activities.
Required Skills
- Robust hands-on experience with Databricks and Apache Spark
- Expertise in Python / PySpark
- Solid SQL and ETL development experience
- Experience with orchestration tools like:
- Apache Airflow
- Azure Data Factory (ADF)
- Autosys
- Experience with CI/CD tools and deployment pipelines
- Knowledge of data warehousing and big data concepts
- Experience with Git/version control systems
- Strong debugging and performance optimization skills
- Good understanding of distributed data processing systems
Good to Have
- Experience with Azure, AWS, or GCP cloud platforms
- Knowledge of Delta Lake, Kafka, or streaming technologies
- Familiarity with DevOps and Infrastructure as Code (IaC)
- Exposure to Snowflake or other cloud data warehouses
- Experience working in Agile/Scrum environments
Preferred Qualifications
- Bachelor's degree in Computer Science, Information Technology, or related field
- Strong analytical and problem-solving abilities
- Excellent communication and stakeholder management skills
Technology Stack
- Databricks
- Apache Spark / PySpark
- Python
- SQL
- ETL/ELT
- Airflow / ADF / Autosys
- CI/CD
- Git
-
📌 Data Engineer (Pune)
🏢 TalentOla
📍 Pune