14 Sep
|
Halcer
|
Bengaluru
Job Title: GCP Lead Data Engineer
Total Experience: 8–13 years
Work Location: Bangalore
Notice Period: Immediate to 30 Days Joiners Only
Role Overview
We are seeking a hands-on Lead Data Engineer with strong technical expertise in Google Cloud Platform (GCP), Scala, PySpark, and Kafka. In this role, you will design, build, and optimize enterprise-grade batch and streaming data processing frameworks. You will lead technical implementations to enable advanced analytics, machine learning, and AI solutions for global clients while adhering to cloud security, performance, and governance best practices.
Key Responsibilities
- Pipeline Engineering: Build, maintain, and optimize robust batch and real-time streaming data pipelines using Scala, PySpark, Apache Spark, and Kafka/PubSub.
- Architecture & Leadership: Drive technical architecture discussions, code reviews, and DevOps standards across the data engineering team.
- Cross-Functional Collaboration: Partner with Data Scientists, Business Analysts, and stakeholders to deliver data platforms that support analytics and machine learning workflows.
- Operations & Quality: Ensure high data quality, pipeline reliability, and optimal performance across dynamic distributed systems.
- Production Support & Optimization: Perform root cause analysis on production issues, optimize resource costs,
and implement strict cloud security and governance standards.
Required Core Skills
- Mandatory Technical Stack: Scala, PySpark, and Apache Kafka integrated with core Data Engineering workflows.
- Data Engineering Expertise: Proven track record in designing high-throughput ETL/ELT pipelines and processing large-scale datasets.
- Distributed Systems: Deep understanding of distributed systems engineering, data modeling, performance tuning, and Apache Beam.
- Cloud & Warehousing: Hands-on experience with SQL, NoSQL systems, and cloud data platform ecosystems (BigQuery, GCP).
- Software Development Practices: Strong experience with CI/CD automation, version control (Git), and Agile delivery methodologies.
Acceptable Alternative Skill Profiles
Candidates matching any of the following technical skill combinations are also eligible:
- GCP + Scala (with or without Kafka)
- GCP + Apache Spark
- Scala + Apache Spark + Kafka
- GCP (BigQuery) + PySpark + SQL (Kafka is a plus)
Preferred (Positive-to-Have) Skills
- Exposure to building real-time streaming architectures using Google Cloud Pub/Sub alongside Kafka.
- Python proficiency combined with Scala background.
- Familiarity with machine learning data pipelines, feature stores, and MLOps integrations.
📌 GCP Lead Data Engineer (Bengaluru)
🏢 Halcer
📍 Bengaluru