07 Aug
|
eTeam
|
Bengaluru
Lead Data Engineer
Experience: 10+ Years
Location: Bangalore (Hybrid 23 Days/Week)
Shift: 12:00 PM – 10:00 PM IST
Job Description
We are looking for an experienced Lead Data Engineer to design, develop, and deliver scalable, high-performance data platforms using modern cloud technologies. The ideal candidate should have strong hands-on expertise in Databricks, Apache Spark, PySpark, Python, SQL, ETL/ELT, and cloud platforms (AWS/Azure/GCP), along with experience in building enterprise data lakes and lakehouse architectures.
This role involves working with cross-functional teams to build robust data pipelines, optimize large-scale data processing, and enable advanced analytics and GenAI use cases.
Key Responsibilities
- Design and develop scalable batch and streaming data pipelines.
- Build and maintain ETL/ELT workflows using Databricks and Apache Spark.
- Develop and optimize Data Lake/Lakehouse architectures using Delta Lake and Medallion (Bronze/Silver/Gold) architecture.
- Design data models for analytics, reporting, machine learning, and AI workloads.
- Build high-quality datasets to support GenAI/LLM applications.
- Implement workflow orchestration using Apache Airflow or Databricks Workflows.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Lead cloud data migration initiatives from legacy platforms to up-to-date cloud-native architectures.
- Implement CI/CD, monitoring, logging, and data quality frameworks.
- Collaborate with architects, data scientists, and business stakeholders to deliver scalable solutions.
- Mentor junior engineers and drive engineering best practices.
Required Skills
- 10+ years of Data Engineering experience.
- Strong hands-on experience with Databricks Lakehouse Platform.
- Expertise in Apache Spark, PySpark, Spark SQL.
- Advanced proficiency in Python and SQL.
- Strong experience building ETL/ELT pipelines.
- Experience with Delta Lake, Delta Live Tables (DLT), Unity Catalog, and Medallion Architecture.
- Hands-on experience with AWS, Azure, or GCP.
- Strong knowledge of Data Lakes, Data Warehousing, and Data Modeling.
- Experience with Apache Airflow or Databricks Workflows.
- Strong understanding of distributed data processing and Spark performance tuning.
- Experience with Git, CI/CD, Agile, and DevOps practices.
- Excellent communication and stakeholder management skills.
Preferred Skills
- Experience with LLMs, Generative AI, RAG, LangChain, LlamaIndex, Embeddings, and Vector Databases.
- Experience with Kafka or Spark Structured Streaming.
- Knowledge of Terraform, Jenkins, or Azure DevOps.
- Cloud or Databricks certifications are an added advantage.
Notice Period: Immediate
📌 Lead Data Engineer (Bengaluru)
🏢 eTeam
📍 Bengaluru