01 Oct
|
Important Group
|
Hyderabad
01 Oct
Important Group
Hyderabad
Job Summary
We are seeking a skilled Databricks Developer to design, develop, and optimize scalable data pipelines and analytics solutions using Apache Spark, Python (PySpark), and SQL within up-to-date cloud environments. The ideal candidate will have hands -on experience working with Databricks, Delta Lake, and cloud data platforms, and will be responsible for building high -performance data processing workflows that support data -driven decision -making.
Key Responsibilities
Pipeline Development
- Design, develop, and maintain ETL/ELT data pipelines using PySpark and SQL in Databricks notebooks.
- Process large -scale datasets and ensure reliable and efficient data transformations.
Architecture Design
- Implement data lake and data warehouse architectures using Databricks Delta Lake and Delta Live Tables.
- Build and manage Medallion Architecture (Bronze, Silver, Gold layers) for structured data processing.
Performance Optimization
- Optimize Spark jobs and queries for performance, scalability, and cost -efficiency.
- Manage cluster configurations, partitioning strategies, and caching mechanisms.
Cloud Integration
- Integrate Databricks solutions with cloud services such as:
- Azure Data Factory (ADF)
- Azure Data Lake Storage (ADLS) Gen2
- AWS S3 or other cloud storage platforms
Data Governance & Quality
- Implement data quality checks, validation frameworks, and monitoring processes.
- Ensure data security, encryption, masking, and lineage tracking.
Workflow Automation
- Build and manage automated workflows and scheduling using Databricks Jobs, Airflow, or CI/CD pipelines.
- Integrate with DevOps tools such as Azure DevOps or Jenkins for continuous integration and deployment.
Collaboration
- Work closely with data engineers, analysts, and business stakeholders to translate business requirements into scalable data solutions.
- Participate in Agile development processes including sprint planning and technical discussions.
Required Skills & Qualifications
Core Technologies
- 3–5+ years of experience with Databricks, Apache Spark, and Python (PySpark).
- Strong experience building scalable ETL/ELT pipelines.
SQL Expertise
- Advanced knowledge of Spark SQL or Databricks SQL for data transformation and analysis.
Cloud Platforms
- Hands -on experience with Azure Databricks, AWS, or Google Cloud Platform (GCP).
Data Modeling
- Experience with Delta Lake and Medallion Architecture (Bronze/Silver/Gold layers).
- Strong understanding of data modeling and data warehouse concepts.
Version Control & CI/CD
- Proficiency with Git and CI/CD tools such as Azure DevOps or Jenkins.
Preferred Qualifications
- Experience with Data Build Tool (dbt).
- Knowledge of workflow orchestration tools such as Apache Airflow.
- Experience with Machine Learning workflows using MLflow.
- Familiarity with data governance and enterprise data platforms.
Education
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Data Engineering, or a related field.
Key Competencies
- Strong problem -solving and analytical skills
- Ability to handle large -scale data processing environments
- Good communication and collaboration skills in Agile teams