09 Sep
|
Qloron Technology
|
Nagpur
09 Sep
Qloron Technology
Nagpur
JOB ID:QT-SNB-08-346
Role Overview
We are looking for a highly experienced Solution Architect - Databricks / Senior FDE / Resident Solution Architect to work closely with enterprise customers in designing, developing, optimizing, and supporting scalable data engineering and analytics solutions on the Databricks Lakehouse Platform.
The ideal candidate should have deep hands-on expertise in Databricks, Apache Spark, PySpark, SQL, Delta Lake, distributed computing, cloud platforms, performance optimization, data architecture, and enterprise data platforms.
This is a highly technical and client-facing consulting role. The candidate should be capable of independently leading architecture workshops, solution design, technical discussions, troubleshooting, performance optimization, implementation guidance, and production deployments.
Key Responsibilities Solution Architecture
- Design and implement scalable and secure Databricks Lakehouse architectures for enterprise customers.
- Understand business and technical requirements and translate them into scalable data platform solutions.
- Define end-to-end data engineering, analytics, integration, and platform architectures.
- Lead architecture workshops, technical discussions, solution reviews, and design sessions.
- Provide technical recommendations aligned with Databricks and cloud best practices.
- Define development standards, architecture patterns, security models, and implementation strategies.
- Support enterprise Databricks platform modernization and transformation initiatives.
Databricks & Data Engineering
- Build and optimize large-scale data pipelines using Databricks, Apache Spark, PySpark, SQL, and Delta Lake.
- Design and implement both batch and real-time/streaming data processing solutions.
- Work with Databricks capabilities including:
- Delta Lake
- Unity Catalog
- Databricks Workflows
- Auto Loader
- Databricks SQL
- Lakeflow / Delta Live Tables (DLT)
- Serverless compute
- Design robust ETL/ELT pipelines and data ingestion frameworks.
- Develop scalable data transformation and data modeling solutions.
- Provide implementation guidance to engineering teams.
- Conduct code reviews and recommend improvements in performance, scalability, maintainability, and reliability.
Apache Spark & PySpark
- Demonstrate deep understanding of Apache Spark internals and distributed computing.
- Troubleshoot complex Spark and Databricks issues.
- Analyze Spark execution plans and identify performance bottlenecks.
- Work extensively with:
- Driver and Executors
- DAG
- Jobs, Stages, and Tasks
- Partitioning
- Shuffle
- Memory Management
- Catalyst Optimizer
- Adaptive Query Execution (AQE)
- Spark SQL execution plans
- Data Skew
- Join Optimization
Performance & Cost Optimization
- Perform advanced Spark and Databricks performance tuning.
- Optimize partitioning, shuffles, joins, queries, and data layouts.
- Identify and resolve data skew and inefficient execution patterns.
- Optimize Delta tables using appropriate optimization techniques.
- Provide guidance on Photon, cluster sizing, autoscaling, and compute configuration.
- Optimize workloads for performance, scalability, reliability, and cloud cost.
- Analyze workload patterns and recommend appropriate Databricks compute architectures.
Cloud Architecture
- Demonstrate deep expertise in at least one major cloud platform:
- AWS
- Microsoft Azure
- Google Cloud Platform (GCP)
- Working knowledge of at least one additional cloud platform is preferred.
AWS: S3, IAM, Glue, Lambda, Kinesis, Redshift
Azure: ADLS, ADF, Key Vault, Entra ID, Synapse, Event Hubs
GCP: GCS, BigQuery, Pub/Sub, Dataflow, IAM
- Design secure and scalable cloud-based data architectures.
- Integrate Databricks with cloud-native storage, compute, networking, identity, and data services.
Security, Governance & Unity Catalog
- Provide guidance on Databricks security and governance.
- Design appropriate access control and permission models.
- Implement and support Unity Catalog.
- Define data governance, data access, lineage, and security standards.
- Work with cloud identity and access management services.
- Ensure enterprise security and compliance requirements are incorporated into platform architecture.
CI/CD & DevOps
- Design and support CI/CD processes for Databricks development and deployment.
- Work with:
- Git
- Terraform
- Databricks Asset Bundles
- Azure DevOps
- GitHub
- GitLab
- Jenkins
- Support deployment across Development, Test, UAT, and Production environments.
- Implement infrastructure and deployment automation using Infrastructure-as-Code practices.
- Establish release management and deployment standards for Databricks workloads.
MLOps
- Demonstrate working knowledge of MLflow and MLOps practices.
- Support experiment tracking and model lifecycle management.
- Work with Model Registry and model deployment/serving capabilities.
- Understand the integration of data engineering and machine learning workflows within Databricks.
Client Engagement & Technical Leadership
- Act as a trusted technical advisor to enterprise customers.
- Lead customer-facing technical discussions and architecture workshops.
- Gather and analyze technical and business requirements.
- Present architecture options, recommendations, and technical solutions to stakeholders.
- Support customer teams with troubleshooting, code reviews, and technical mentoring.
- Collaborate with customer architects, data engineers, developers, DevOps teams, and leadership.
- Mentor engineering teams on Databricks, Spark, PySpark, and data engineering best practices.
- Independently drive complex technical issues through investigation and resolution.
Mandatory Skills Databricks
- Strong hands-on Databricks development and architecture experience.
- Minimum 6-8 end-to-end Databricks project implementations.
- Databricks Lakehouse Architecture.
- Delta Lake.
- Unity Catalog.
- Databricks Workflows.
- Auto Loader.
- Databricks SQL.
- Lakeflow / DLT.
- Batch and streaming workloads.
- Cluster and compute configuration.
- Databricks performance optimization.
Apache Spark / PySpark
- Strong hands-on experience with Apache Spark and PySpark.
- Deep understanding of Spark internals.
- Spark performance tuning and troubleshooting.
- Partitioning and shuffle optimization.
- Data skew handling.
- Join optimization.
- Spark SQL and execution plans.
- Catalyst Optimizer and AQE.
Data Engineering
- Strong ETL / ELT experience.
- Data ingestion and transformation.
- Data pipeline development.
- Data modeling.
- Batch processing.
- Streaming processing.
- Advanced SQL.
- Python / PySpark.
- Large-scale distributed data processing.
Cloud
- Deep expertise in AWS, Azure, or GCP.
- Working knowledge of at least one additional cloud platform preferred.
- Strong understanding of cloud data services and integrations.
Performance & Scalability
- Spark performance tuning.
- Query optimization.
- Shuffle optimization.
- Partitioning.
- Data skew handling.
- Join optimization.
- Delta table optimization.
- Photon.
- Cluster sizing.
- Autoscaling.
- Cost optimization.
DevOps / CI/CD
- Git.
- CI/CD pipelines.
- Terraform.
- Databricks Asset Bundles.
- Azure DevOps / GitHub / GitLab / Jenkins.
- Multi-environment deployment processes.
Preferred Skills
- Enterprise Databricks architecture.
- Databricks migration and modernization.
- Hadoop-to-Databricks migration.
- Cloud Data Warehouse-to-Databricks migration.
- Multi-cloud architecture.
- Unity Catalog implementation.
- Terraform and Infrastructure-as-Code.
- Databricks Asset Bundles.
- MLflow / MLOps.
- Data governance and security.
- Advanced streaming architecture.
- Technical leadership and mentoring.
Required Client-Facing Competencies
- Customer-facing technical discussions.
- Architecture workshops.
- Requirement gathering and analysis.
- Solution design.
- Technical presentations.
- Stakeholder management.
- Architecture reviews.
- Technical recommendations.
- Complex troubleshooting.
- Root-cause analysis.
- Problem-solving and decision-making.
Experience & Certification
- 10+ years of overall consulting / technology experience.
- 7+ years of Data Engineering, Big Data, Data Platforms, or Analytics experience.
- Strong hands-on Databricks experience.
- Minimum 6-8 Databricks end-to-end implementations.
- Strong client-facing consulting experience.
- Databricks Certified Data Engineer Professional - Mandatory / Highly Preferred.
- Completion of relevant Databricks training and certification programs is preferred.
Ideal Candidate Profile
The ideal candidate will have:
- 10-15+ years of overall technology/consulting experience.
- 7+ years of Data Engineering / Big Data experience.
- Strong hands-on Databricks and Apache Spark expertise.
- 6-8+ successful Databricks project implementations.
- Deep knowledge of Spark internals and PySpark.
- Strong expertise in Spark and Databricks performance tuning.
- Deep expertise in at least one cloud platform with exposure to another.
- Solid understanding of Lakehouse architecture, Delta Lake, Unity Catalog, streaming, and data governance.
- Strong solution architecture and technical leadership capabilities.
- Excellent customer-facing consulting and communication skills.
- Ability to independently troubleshoot complex technical issues and guide enterprise customers toward scalable solutions.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Solution Architect - Databricks (Nagpur)
🏢 Qloron Technology
📍 Nagpur