20 Aug
|
Conoscenza Technology Solutions
|
Chennai
20 Aug
Conoscenza Technology Solutions
Chennai
Job Description: AWS Databricks Data Engineer
Role Overview:
We are looking for an experienced AWS Databricks Data Engineer to design, developand optimize scalable Lakehouse data solutions. The ideal candidate should have strong hands-on experience with DatabricksAWS cloud services, PySpark, Spark SQLand Delta Lakewith the ability to build reliableproduction-grade data pipelines and translate business requirements into scalable data engineering solutions.
Experience Level
- Data Engineer: 5-7 years of overall data engineering experience, including at least 3 years of hands-on experience with Databricks, Spark and Lakehouse implementations.
- Senior Data Engineer / Lead Engineer: 7-12 years of experience with solid expertise in Databricks architecture, data pipeline design, performance optimization, production deployments and leading enterprise-scale data engineering solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Databricks, Delta Lake, Auto Loader and Delta Live Tables (DLT).
- Design and implement metadata-driven data ingestion and transformation frameworks to enable scalable, reusable and configuration-based pipeline development across multiple data sources.
- Develop framework components for managing source configurations, ingestion rules, schema evolution, data quality validations, audit logging and pipeline execution monitoring.
- Implement Medallion Architecture (Bronze, Silver, Gold) with appropriate data quality checks, validation frameworks, testing, and monitoring.
- Develop ingestion frameworks using Databricks Auto Loader, including checkpointing, schema evolution and handling structured/semi-structured data.
- Configure and manage Unity Catalog including catalogs, schemas, access controls, audit logging, data lineage and security policies.
- Optimize Spark workloads by tuning Spark jobs, cluster configurations, partitioning strategies, and Delta Lake storage layouts for performance and cost efficiency.
- Build and manage workflows using Databricks Jobs/Lakeflow Jobs for batch and streaming data processing.
- Implement CI/CD practices and integrate with Git-based DevOps processes.
- Integrate Databricks solutions with AWS services such as S3, IAM, Glue, Step Functionsand other AWS data services.
- Enable analytics consumption through BI tools such as Power BI, Tableau or Looker using optimized connectivity patterns.
- Collaborate with data architects, analysts, data scientists and business stakeholders to deliver enterprise data platform solutions.
- Troubleshoot pipeline failures, performance issues and operational challenges while maintaining technical documentation.
Required Skills & Experience
Category
Requirements
Databricks Platform
Strong hands-on experience with Databricks, Delta Lake, Unity Catalog, Workflows/Jobs and Lakehouse architecture
Cloud Platform
AWS experience including S3, IAM, Glue, EMR, Step Functions and related services
Programming
Strong programming skills in Python, PySpark, Spark SQL and SQL
Data Engineering
Experience designing ETL/ELT pipelines, batch and streaming processing and data modeling
Architecture
Understanding of Lakehouse architecture, Medallion framework and scalable data platform design
Data Governance
Experience with security, RBAC, PII handling, lineage and compliance requirements
DevOps
Experience with CI/CD pipelines, Git, Databricks Asset Bundles and Infrastructure as Code concepts
BI Integration
Experience connecting Databricks with Power BI, Tableau or similar analytics platforms
Soft Skills
Strong communication skills, stakeholder management, documentation and ability to mentor junior engineers
Nice to Have
- Experience with Databricks advanced capabilities such as MLflow, Feature Store, Vector Search and GenAI workloads.
- Knowledge of Delta Sharing and Databricks Marketplace.
- Experience migrating workloads from platforms such as Snowflake, Azure Databricks or traditional data warehouses.
- Exposure to enterprise-scale data platform implementations.
📌 Aws Databricks (Chennai)
🏢 Conoscenza Technology Solutions
📍 Chennai