10 Sep
|
Flexton
|
Bengaluru
Job Description
KEY RESPONSIBILITIES
n
n
- Design, build, and maintain scalable data ingestion pipelines from diverse source systems (databases, APIs, files, streaming) into the Databricks lakehouse.
n
- Implement and optimize the medallion architecture (bronze/silver/gold layers), ensuring clear data quality and transformation logic at each stage.
n
- Develop data transformation and cleansing logic using PySpark, Spark SQL, and Delta Lake to produce curated, analytics-ready datasets.
n
- Build and orchestrate ETL/ELT workflows using Databricks Workflows, Delta Live Tables, and/or orchestration tools (Airflow, ADF).
n
- Implement data quality checks, validation rules, and monitoring/alerting across pipelines to ensure trustworthy data.
n
- Manage schema evolution, partitioning, and performance tuning for large-scale Delta tables.
n
- Collaborate with data governance teams to apply cataloging, access controls, and lineage tracking (Unity Catalog).
n
- Partner with BI analysts, data scientists, and AI engineers to understand downstream data requirements and ensure gold-layer tables meet consumption needs.
n
- Document data pipelines, transformation logic, and data models for maintainability and knowledge sharing.
n
- Troubleshoot and resolve data pipeline failures, latency issues, and data quality incidents.
n
n
REQUIRED SKILLS AND EXPERIENCE
n
n
- 4+ years of experience in data engineering,
with at least 2 years working hands-on with Databricks.
n
- Solid hands-on experience implementing medallion architecture (bronze, silver, gold layers) for data ingestion and transformation.
n
- Proficiency in PySpark and Spark SQL for large-scale data processing.
n
- Solid experience with Delta Lake, including ACID transactions, schema evolution, time travel, and optimization (Z-ordering, compaction).
n
- Experience building both batch and streaming ingestion pipelines.
n
- Strong SQL skills and experience with data modeling (dimensional modeling, star schema) for analytics consumption.
n
- Experience with orchestration tools such as Databricks Workflows, Delta Live Tables, Apache Airflow, or Azure Data Factory.
n
- Familiarity with cloud data platforms (Azure Databricks, AWS, or GCP) and associated storage services (ADLS, S3).
n
- Understanding of data governance concepts: cataloging, access control, and data lineage (Unity Catalog or equivalent).
n
n
PREFERRED / NICE-TO-HAVE SKILLS
n
n
- Experience with Python for pipeline automation and testing.
n
- Familiarity with CI/CD for data pipelines (Databricks Asset Bundles, DevOps pipelines).
n
- Exposure to data quality frameworks (Great Expectations, Deequ).
n
- Experience supporting downstream BI tools (Power BI, Tableau) or AI/ML use cases.
n
- Databricks certifications (Data Engineer Associate/Professional) a plus.
n
📌 Data Engineer (Bengaluru)
🏢 Flexton
📍 Bengaluru