01 Sep
|
Sourcebae
|
Bengaluru
01 Sep
Sourcebae
Bengaluru
Solution Architect – Databricks
Experience – 10+ Years
Duration – Contract 6 – 12 months
Note - Initially have to go onsite for 2 weeks in Bangalore then Remote
Experience Required
- 10+ years of overall consulting / technology experience
- 7+ years of experience in Data Engineering, Data Platforms, Big Data, or Analytics
- Strong hands-on Databricks experience
- Minimum 6–8 end-to-end Databricks project implementations
Must Have - We are looking for candidates with approximately:
- 10–15+ years overall experience
- 7+ years Data Engineering / Big Data experience
- Strong Databricks and Spark experience
- 6–8+ Databricks project implementations
- Strong client-facing consulting background
- Databricks Data Engineer Professional certification
- Deep Spark / PySpark and Spark internals knowledge
- Strong performance tuning expertise
- Deep knowledge of one cloud platform and exposure to another
Strong solution architecture and technical leadership capabilities
Role Overview
We are looking for a highly experienced Senior FDE / Resident Solution Architect – Databricks who can work closely with enterprise customers to design, develop, optimize, and support scalable data engineering and analytics solutions on the Databricks platform.
The consultant should have strong hands-on expertise in Databricks, Apache Spark, PySpark, distributed computing, cloud platforms, performance optimization, and solution architecture .
This is a highly technical and client-facing role. The consultant should be able to independently drive architecture discussions,
troubleshoot complex Databricks/Spark issues, provide implementation guidance, and support production deployments.
Key Responsibilities
- Design and implement scalable Databricks Lakehouse solutions.
- Work directly with customers to understand technical and business requirements.
- Define end-to-end data engineering and platform architecture.
- Build and optimize data pipelines using Databricks, Spark, PySpark, SQL, and Delta Lake.
- Design batch and streaming data-processing solutions.
- Provide technical guidance on Databricks architecture, development standards, and best practices.
- Troubleshoot complex Spark and Databricks performance issues.
- Optimize workloads for performance, scalability, reliability, and cost.
- Support enterprise Databricks platform implementation and modernization initiatives.
- Work with Databricks capabilities such as Delta Lake, Unity Catalog, Workflows, Auto Loader, Databricks SQL, Lakeflow/DLT, and Serverless.
- Design and support CI/CD processes for Databricks deployments.
- Work with DevOps and Infrastructure-as-Code tools such as Git, Terraform, Azure DevOps, GitHub, GitLab, or Jenkins.
- Provide guidance on Databricks security, governance, access control,
and Unity Catalog.
- Support customer teams with architecture reviews, code reviews, troubleshooting, and technical mentoring.
- Work as a trusted technical advisor to customer architects, engineering teams, and stakeholders.
Mandatory Skills
Databricks
- Robust hands-on experience in Databricks development and architecture
- Minimum 6–8 Databricks projects delivered
- Delta Lake
- Databricks Lakehouse Architecture
- Unity Catalog
- Databricks Workflows
- Auto Loader
- Databricks SQL
- Batch and streaming workloads
- Cluster / compute configuration
- Performance optimization
Apache Spark / PySpark
- Strong hands-on Spark and PySpark development
- Deep understanding of Spark internals including:
- Driver and Executors
- DAG
- Jobs, Stages, and Tasks
- Partitioning
- Shuffle
- Memory management
- Catalyst Optimizer
- Adaptive Query Execution
- Spark SQL execution plans
- Data skew
- Join optimization
Data Engineering
- Strong ETL / ELT experience
- Data ingestion and transformation
- Data pipelines
- Data modeling
- Batch processing
- Streaming processing
- SQL
- Python / PySpark
- Large-scale distributed data processing
Cloud
Candidate must have
- Deep expertise in at least one cloud platform: AWS / Azure / GCP
- Working knowledge of at least one additional cloud platform
Relevant cloud services may include: AWS: S3, IAM, Glue, Lambda, Kinesis, Redshift
Azure: ADLS, ADF, Key Vault, Entra ID, Synapse, Event Hubs
GCP: GCS, BigQuery, Pub/Sub, Dataflow, IAM
📌 Databricks Architect (Bengaluru)
🏢 Sourcebae
📍 Bengaluru