06 Sep
|
Flexton
|
Bengaluru Urban
06 Sep
Flexton
Bengaluru Urban
KEY RESPONSIBILITIES
- Design, build, and maintain scalable data ingestion pipelines from diverse source systems (databases, APIs, files, streaming) into the Databricks lakehouse.
- Implement and optimize the medallion architecture (bronze/silver/gold layers), ensuring explicit data quality and transformation logic at each stage.
- Develop data transformation and cleansing logic using PySpark, Spark SQL, and Delta Lake to produce curated, analytics-ready datasets.
- Build and orchestrate ETL/ELT workflows using Databricks Workflows, Delta Live Tables, and/or orchestration tools (Airflow, ADF).
- Implement data quality checks, validation rules, and monitoring/alerting across pipelines to ensure trustworthy data.
- Manage schema evolution, partitioning, and performance tuning for large-scale Delta tables.
- Collaborate with data governance teams to apply cataloging, access controls, and lineage tracking (Unity Catalog).
- Partner with BI analysts, data scientists, and AI engineers to understand downstream data requirements and ensure gold-layer tables meet consumption needs.
- Document data pipelines, transformation logic, and data models for maintainability and knowledge sharing.
- Troubleshoot and resolve data pipeline failures, latency issues, and data quality incidents.
REQUIRED SKILLS AND EXPERIENCE
- 4+ years of experience in data engineering, with at least 2 years working hands-on with Databricks.
- Strong hands-on experience implementing medallion architecture (bronze, silver, gold layers) for data ingestion and transformation.
- Proficiency in PySpark and Spark SQL for large-scale data processing.
- Solid experience with Delta Lake, including ACID transactions, schema evolution, time travel, and optimization (Z-ordering, compaction).
- Experience building both batch and streaming ingestion pipelines.
- Strong SQL skills and experience with data modeling (dimensional modeling, star schema) for analytics consumption.
- Experience with orchestration tools such as Databricks Workflows, Delta Live Tables, Apache Airflow, or Azure Data Factory.
- Familiarity with cloud data platforms (Azure Databricks, AWS, or GCP) and associated storage services (ADLS, S3).
- Understanding of data governance concepts: cataloging, access control, and data lineage (Unity Catalog or equivalent).
PREFERRED / NICE-TO-HAVE SKILLS
- Experience with Python for pipeline automation and testing.
- Familiarity with CI/CD for data pipelines (Databricks Asset Bundles, DevOps pipelines).
- Exposure to data quality frameworks (Great Expectations, Deequ).
- Experience supporting downstream BI tools (Power BI, Tableau) or AI/ML use cases.
- Databricks certifications (Data Engineer Associate/Professional) a plus.
📌 Data Engineer (Bengaluru Urban)
🏢 Flexton
📍 Bengaluru Urban