26 Aug
|
Qloron Technology
|
Nagpur
26 Aug
Qloron Technology
Nagpur
JOB ID:QT-SUL-08-424
Job Summary
We are looking for a Lakehouse Engineer to build and maintain the data foundation for agentic AI applications. The role involves ingesting structured and unstructured data into the Databricks Lakehouse, enabling document parsing, metadata management, indexing, vector search, governance, and retrieval-ready data pipelines.
This is a hands-on Databricks engineering role requiring solid expertise in Spark, PySpark, SQL, Delta Lake, Unity Catalog, and modern data ingestion and retrieval technologies.
Key Responsibilities
- Build and operate ingestion pipelines for databases, APIs, files, documents, images, and audio.
- Develop AI-based document parsing pipelines covering OCR, layout extraction, table extraction, text extraction, and document chunking.
- Design metadata, tagging, classification, and governance frameworks using Unity Catalog.
- Build embedding pipelines, vector indexes, and hybrid keyword + semantic search capabilities.
- Design and maintain Bronze, Silver, and Gold data layers using Delta Lake.
- Implement data quality, lineage, governance, and fine-grained access controls.
- Optimize Databricks workloads using partitioning, Liquid Clustering, Z-Ordering, caching, and job/cluster tuning.
- Build reliable, retrieval-ready datasets for AI and BI teams.
- Collaborate independently with AI/ML and data engineering teams to deliver solutions within tight timelines.
Required Skills & Experience
- 5+ years of overall Data Engineering experience, with substantial hands-on Databricks experience.
- Strong knowledge of Apache Spark, PySpark, and advanced SQL.
- Hands-on experience with Delta Lake, Delta Live Tables/Lakeflow,
Auto Loader, and Databricks Workflows.
- Strong understanding of Unity Catalog, including governance, lineage, tagging, and permissions.
- Experience building pipelines for unstructured data and document parsing/OCR.
- Hands-on experience with embeddings and vector indexes, preferably Databricks Vector Search or equivalent technologies.
- Working knowledge of LLMs, embeddings, chunking, and retrieval/RAG concepts.
- Experience with Azure, AWS, or GCP.
- Knowledge of CI/CD practices for data engineering pipelines.
- Strong troubleshooting, optimization, and problem-solving skills.
- Ability to work independently in a short-term, delivery-focused engagement.
Good to Have
- Structured Streaming and Kafka experience.
- Databricks Asset Bundles (DABs).
- Terraform / Infrastructure as Code.
- Data observability and quality frameworks.
- Knowledge graphs and ontology-driven metadata models.
Ideal Candidate
The ideal candidate is a hands-on Databricks/Lakehouse Engineer who can independently design and deliver scalable data pipelines, manage structured and unstructured data, and build the retrieval layer required for modern GenAI and agentic AI applications.
Note: The 5+ years requirement refers to overall data engineering experience. Five years of LLM-specific experience is not required; however, a sound understanding of LLMs, embeddings, and retrieval concepts is essential.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Lakehouse Engineer (Nagpur)
🏢 Qloron Technology
📍 Nagpur