20 Aug
|
SPG Consulting
|
Bengaluru
20 Aug
SPG Consulting
Bengaluru
Job Description – Databricks Architect
Position
Databricks Architect
Experience
6–9 Years
Job Summary
We are looking for an experienced Databricks Architect to design and implement scalable, secure, and high-performance data platforms using Databricks and Apache Spark. The ideal candidate should have strong expertise in data architecture, cloud platforms, data engineering, Lakehouse architecture, and enterprise-scale Databricks implementations.
Key Responsibilities
- Design and implement enterprise-grade Databricks Lakehouse architectures.
- Define data architecture strategies, standards, governance, and best practices.
- Design scalable data ingestion, transformation, processing, and analytics pipelines.
- Develop architecture solutions using Databricks, Apache Spark, Delta Lake, and cloud-native services.
- Define data lake and lakehouse structures using Bronze, Silver, and Gold layers.
- Design batch and real-time data processing solutions.
- Provide technical leadership to Data Engineering and BI teams.
- Review existing data platforms and recommend modernization strategies.
- Design high-performance and cost-optimized Databricks solutions.
- Establish security, access control, encryption, and data governance standards.
- Design and implement Unity Catalog for centralized data governance.
- Define data lineage, discovery, auditing, and access-management strategies.
- Design CI/CD and DevOps processes for Databricks notebooks, jobs, workflows, and code.
- Integrate Databricks with enterprise data sources, APIs, warehouses, and cloud storage.
- Lead migration of legacy data platforms to Databricks where applicable.
- Conduct architecture reviews, proof-of-concepts, and technology evaluations.
- Collaborate with Data Engineers, Data Scientists, BI Developers, Cloud Architects, and business stakeholders.
- Provide technical guidance, mentoring, and architectural documentation.
Required Skills
Databricks & Spark
- Strong hands-on experience with Databricks.
- Expert-level knowledge of Apache Spark.
- Strong experience with PySpark and/or Scala.
- Experience with Delta Lake and Delta tables.
- Strong knowledge of Databricks Workflows, Jobs, Clusters, Notebooks, and SQL Warehouses.
- Experience designing and implementing Lakehouse architecture.
- Knowledge of Spark performance tuning and optimization.
Data Architecture
- Strong understanding of:
- Data Lake / Data Warehouse / Lakehouse
- Medallion Architecture
- Data Modeling
- ETL/ELT
- Batch and Streaming
- Data Governance
- Data Quality
- Metadata Management
- Experience designing enterprise data platforms and integration architectures.
Cloud
Robust experience with at least one major cloud platform:
- Microsoft Azure
- Amazon Web Services (AWS)
- Google Cloud Platform (GCP)
Preferred Azure technologies include:
- Azure Data Lake Storage Gen2
- Azure Data Factory
- Azure Synapse
- Azure Key Vault
- Azure Event Hubs
Unity Catalog & Governance
- Strong knowledge of Databricks Unity Catalog.
- Design and implement catalogs, schemas, external locations, and storage credentials.
- Implement role-based access control and data security.
- Establish data lineage and auditing.
- Define enterprise data governance and compliance practices.
DevOps & CI/CD
- Experience implementing CI/CD for Databricks solutions.
- Knowledge of Git/GitHub, Azure DevOps, GitHub Actions,
or Jenkins.
- Experience with Infrastructure as Code such as Terraform.
- Familiarity with automated testing, deployment, and release management.
Performance & Cost Optimization
- Optimize Spark workloads, SQL queries, clusters, and Delta tables.
- Experience with partitioning, caching, file optimization, and indexing strategies.
- Knowledge of Delta Lake OPTIMIZE, VACUUM, Z-Ordering, and liquid clustering.
- Implement appropriate cluster sizing and autoscaling strategies.
- Monitor and optimize Databricks platform costs.
Good to Have
- Experience with Databricks SQL and BI integration.
- Knowledge of MLflow and Machine Learning workloads.
- Experience with streaming technologies such as Kafka or Azure Event Hubs.
- Knowledge of Terraform and Infrastructure as Code.
- Experience with Generative AI, RAG, or Vector Search.
- Familiarity with Microsoft Fabric or Snowflake.
- Experience with large-scale cloud migration projects.
- Knowledge of enterprise security and regulatory requirements.
Key Technologies
Databricks | Apache Spark | PySpark | Scala | Delta Lake | Unity Catalog | Databricks SQL | Azure/AWS/GCP | ADLS | ADF | Kafka | Terraform | Git | CI/CD | MLflow
Education
Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, Data Science, or a related discipline.
Preferred Candidate Profile
The ideal candidate should have strong experience designing enterprise Databricks Lakehouse platforms and leading complex data engineering initiatives. The candidate should combine hands-on Databricks expertise with strong knowledge of cloud architecture, Spark, Delta Lake, Unity Catalog, security, governance, performance optimization, and DevOps.
Experience leading large-scale Databricks migration and modernization programs will be highly valued.
📌 Databricks Architect (Bengaluru)
🏢 SPG Consulting
📍 Bengaluru