12 Sep
|
EliteSquad.AI
|
India
12 Sep
EliteSquad.AI
India
Databricks Architect
Remote
Experience: 10+ Years
We are looking for an experienced Databricks Architect to lead the design, architecture, modernization, and implementation of enterprise-scale data platforms using Databricks, Delta Lake, Apache Spark, and cloud-native data services.
About the Role The ideal candidate will have strong expertise in designing Lakehouse architectures, enterprise data platforms, data pipelines, data governance, performance optimization, and large-scale distributed data processing. The candidate should also have practical exposure to Snowflake, including understanding its architecture, data warehousing capabilities, workload patterns, performance optimization, security, and integration with Databricks. This role requires someone who can operate at an architecture and technical leadership level, engage with business and technology stakeholders, evaluate technology choices, define target-state architecture, and guide engineering teams in implementing scalable and secure data solutions.
Responsibilities
- Data Platform Architecture
- Design and implement enterprise-grade Data Lakehouse architectures using Databricks.
- Define current-state and target-state data architecture for large-scale data platforms.
- Design scalable architectures covering data ingestion, transformation, storage, serving, analytics, and AI/ML workloads.
- Develop architecture standards, reference architectures, design patterns, and reusable frameworks.
- Define appropriate data storage and processing strategies based on business and technical requirements.
- Design solutions supporting batch, near-real-time, and real-time data processing.
- Evaluate architectural trade-offs between Databricks, Snowflake, cloud-native services, and other data platforms.
- Design highly available, secure, scalable,
and cost-optimized data platforms.
- Lead cloud data platform modernization and migration initiatives.
- Databricks Architecture
- Provide deep technical leadership across the Databricks Data Intelligence / Lakehouse platform.
- Design Databricks workspaces, compute architecture, clusters, SQL warehouses, jobs, workflows, and data access patterns.
- Architect solutions using:
- Apache Spark
- Delta Lake
- Delta Live Tables / Lakeflow Declarative Pipelines
- Databricks Workflows / Jobs
- Databricks SQL
- Unity Catalog
- MLflow
- Structured Streaming
- Databricks notebooks
- Serverless and/or managed compute
- Define appropriate compute strategies for different workloads.
- Optimize Spark applications and Databricks workloads for performance and cost.
- Establish standards for cluster sizing, autoscaling, workload isolation, and job execution.
- Design reusable frameworks for enterprise data ingestion and transformation.
- Implement medallion architecture patterns such as:
- Bronze
- Silver
- Gold
- Design scalable data products and domain-oriented data architectures.
- Delta Lake & Lakehouse
- Design enterprise-grade Delta Lake architectures.
- Define strategies for:
- ACID transactions
- Schema evolution
- Schema enforcement
- Time travel
- Data versioning
- Change Data Capture (CDC)
- Incremental processing
- Data retention
- Compaction
- Partitioning
- Optimization
- Implement appropriate table design and data lifecycle strategies.
- Use techniques such as OPTIMIZE, Z-ORDER, liquid clustering, partitioning, and file management where appropriate.
- Design efficient ingestion and transformation patterns for large datasets.
- Define data quality and reconciliation mechanisms.
- Apache Spark
- Provide architecture and technical guidance for large-scale Apache Spark workloads.
- Design Spark applications using PySpark and/or Scala.
- Diagnose and resolve Spark performance bottlenecks.
- Optimize:
- Joins
- Shuffles
- Partitioning
- Broadcast joins
- Caching
- Serialization
- Data skew
- Small-file problems
- Executor sizing
- Parallelism
- Analyze Spark execution plans and job performance.
- Establish engineering standards for scalable Spark development.
- Snowflake Exposure
- The candidate should have solid practical exposure to Snowflake and be able to participate in architecture decisions involving Snowflake and Databricks. Responsibilities may include:
- Design and review Snowflake data warehouse architectures.
- Understand Snowflake's:
- Virtual warehouses
- Databases and schemas
- Storage architecture
- Micro-partitions
- Clustering
- Time Travel
- Zero-copy cloning
- Streams and Tasks
- Snowpipe
- Agile Tables
- Snowpark
- Secure data sharing
- Compare and evaluate Snowflake vs. Databricks based on workload requirements.
- Design architectures where Snowflake and Databricks coexist.
- Define data movement and integration patterns between Databricks and Snowflake.
- Optimize Snowflake queries and warehouse utilization.
- Define appropriate workload isolation and warehouse sizing.
- Support migration from traditional data warehouses to Snowflake and/or Databricks.
📌 Databricks Architect (India)
🏢 EliteSquad.AI
📍 India