03 Sep
|
Celebal Technologies
|
Jaipur
03 Sep
Celebal Technologies
Jaipur
Job Summary
We are looking for an experienced Databricks Solution Architect to join our team. You will design and lead end-to-end big data and analytics solutions on the Databricks Lakehouse Platform, translating complex business requirements into scalable, high-performance architectures. Youll work closely with clients and internal engineering teams to architect data pipelines, transformation frameworks, and governance models that ingest and process petabyte-scale data from diverse batch and streaming sources into cloud data lakes.
Responsibilities
- Next-Gen Architecture Delivery: Design and deploy modern Data Lakehouse solutions using Serverless Compute (for SQL, Jobs, and Notebooks) to eliminate infrastructure overhead and optimize TCO.to understand the business problem
- Production Engineering: Write, optimize, and deploy production-grade code. You are expected to fix what breaks-whether it s a shuffling error in Spark, a CI/CD failure in Databricks Asset Bundles (DABs), or a complex Mosaic AI vector search pipeline.
- Deep Performance Tuning: Diagnose and resolve bottlenecks (Skew, OOM, Spill). You must know when to apply legacy tuning (Z-Order/Partitioning) vs. modern Liquid Clustering to handle changing data patterns automatically
- Governance Federation: Implement strict data governance using Unity Catalog. Configure Lakehouse Federation to query external systems without data movement, and enforce Attribute-Based Access Control (ABAC).
- Customer Obsession: Act as the Expert in the Room. Explain complex concepts (like why Serverless reduces cold starts or how AI/BI Genie handles hallucination) to non-technical stakeholders.
Qualifications
- Core Spark Liquid Clustering: Expert knowledge of the Catalyst Optimizer and AQE. You must understand the architectural shift from static partitioning to Liquid Clustering and how it impacts file layout and query skipping.
- Coding Fluency: Advanced proficiency in Python (PySpark) and SQL. You must be comfortable livecoding complex UDFs and transformation logic without reliance on Google.
- Delta Lakehouse Mastery: Deep knowledge of ACID transactions, Time Travel, and Delta Live Tables (DLT) for declarative pipeline management.
- Unity Catalog Security: Hands-on experience with System Tables for observability, Lakehouse Federation for cross-platform querying, and setting up Volume based access controls.
- Operational Excellence: Experience with Infrastructure-as-Code (Terraform), Databricks Asset Bundles (DABs) for CI/CD, and Databricks Connect v2 for local development.
Desired Skills
- SAP Data Integration (High Value): Experience extracting data from SAP systems (ECC, S/4HANA) into Databricks Delta Lake.
- Familiarity with SAP Datasphere, SAP SLT, or Partner Connectors (e.g., Fivetran/Qlik for SAP) for operational reporting.
- Understanding of common SAP data structures (IDOCs, BAPIs) is a major plus.
Good to Have
- ML Engineering: Hands-on experience with MLflow for experiment tracking and model registry. Familiarity with Feature Store implementation.
- GenAI/Mosaic AI: Experience fine-tuning LLMs (Foundation Models), using Mosaic AI Gateway for governance, or building conversational analytics with AI/BI Genie.
- Libraries: Proficiency with core ML libraries (Scikit-learn, XGBoost, PyTorch) for non-generative use cases.
- Serverless Architectures: Proven experience migrating workloads from Classic Compute to Serverless, including cost analysis.
- Certifications: Databricks Certified Data Engineer Skilled (Strongly Preferred)
Disclaimer: This job posting and Location has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Databricks Solutions Architect (Jaipur)
🏢 Celebal Technologies
📍 Jaipur