17 Sep
|
Datametica
|
Pune
Job Title: Big Data Lead
Experience Level: 8+ Years
Location: Pune, India
Job Summary
We are looking for a Lead Big Data Lead to design, scale, and modernize our high-throughput data platform. You will direct technical strategy, manage data pipelines, and bridge traditional big data architectures with modern Lakehouse patterns and Generative AI applications.
Key Responsibilities
- Platform & Lakehouse Architecture: Drive the architecture and migration from legacy Hadoop/Hive systems to contemporary Apache Iceberg formats for ACID compliance, schema evolution, and time-travel querying.
- Query & Compute Optimization: Build and tune high-performance batch and streaming pipelines using Apache Spark. Deploy Trino for sub-second, federated interactive SQL querying across heterogeneous storage layers.
- Orchestration & DevOps: Manage enterprise-wide DAG workflows using Apache Airflow. Deploy, auto-scale, and manage compute workloads on Kubernetes (K8s)
clusters.
- GenAI Integrations: Architect ingestion and processing pipelines for unstructured data, handling vector embeddings and Retrieval-Augmented Generation (RAG) infrastructure.
- Team Leadership & Governance: Lead engineering teams through code reviews, technical design sessions, platform security enforcement, cost optimization, and SLA management.
Technical Skills & Qualifications
Category Requirements
Big Data Core Apache Spark (PySpark/Scala), Apache Hive, Hadoop/HDFS
Lakehouse & Querying Apache Iceberg, Trino (Presto), Complex SQL
Orchestration & Infra Apache Airflow, Kubernetes (K8s), Docker
GenAI Ecosystem Vector databases (Milvus/Pinecone/Qdrant), RAG ingestion pipelines, LLM data processing
Programming Python, Scala, Java
📌 Big Data Lead (Pune)
🏢 Datametica
📍 Pune