Senior Consultant - Lead PySpark Cloud Data Engineer (Chennai)

Senior Consultant - Lead PySpark Cloud Data Engineer (Chennai)

26 Aug
|
enGen Global
|
Chennai

26 Aug

enGen Global

Chennai

Job Summary

Lead PySpark Cloud Data Engineer: This job involves understand the overall requirement of the enterprise Data need and design develop the robust, highly scalable resilient data pipeline using Pyspark, Dataproc, open source table formats and other Google services. Job will involve extensive interfacing co-ordination with other senior tech folks (across India US) and lead the design development of the data ingestion pipelines

Essential Responsibilities

- Design, develop, and implement highly scalable, reliable, and performant data pipelines using PySpark, Dataproc, and other Google cloud-native technologies
- Create pipelines for data lakehouse architectures utilizing Apache Iceberg for open table formats and BigLake Metastore for unified metadata management
- Develop and maintain data transformation logic using Pyspark and / or dbt to create clean, consistent, and production-ready data models
- Work extensively with Google Cloud Platform (GCP) services, including Dataproc, BigQuery, Cloud Storage, and other relevant data services.
- Optimize data pipelines and queries for performance, efficiency, and cost-effectiveness.
- Implement and enforce data quality checks, monitoring, and goverce best practices to ensure data integrity and reliability
- Provide technical guidance and mentorship to junior data engineers, fostering a culture of continuous learning and improvement
- Create and maintain comprehensive technical documentation for data pipelines, data models, and platform architecture
- Provide regular updates on the tasks,



status and risks to project manager

The Experience We Are Looking To Add To Our Team Required
- Bachelors degree or higher from a reputed university
- 8 to 12 years total experience with majority of that experience related to building high performant data pipelines batch and streaming
- Strong Hands on experience / expert-level proficiency in PySpark for data processing, transformation, and analysis
- Extensive experience in implementing large scale data ingestion and curation solutions
- Hands-on experience with cloud-based data platforms, preferably Google Cloud Platform (GCP) and services like Dataproc, BigQuery, and Cloud Storage
- Knowledge in Apache Iceberg for open table formats
- Advanced SQL skills for data querying, manipulation, and optimization
- Proficiency with Git and team-oriented development workflows
- Excellent analytical and problem-solving skills with a keen attention to detail
- Strong communication and interpersonal skills, with the ability to explain complex technical concepts to both technical and non-technical audiences

Nice to have
- Expertise in Google Cloud services
- Proven experience with Apache Iceberg for open table formats and BigLake Metastore for unified metadata management
- Strong experience with dbt (data build tool) for data modeling, transformation, and orchestration
- Experience in Data Goverce and Data quality tools like Atlan, Monte Carlo etc.
- Experience in processing streaming data using Kafka / Pub-Sub
- Healthcare industry experience
- Experience in Agile

📌 Senior Consultant - Lead PySpark Cloud Data Engineer (Chennai)
🏢 enGen Global
📍 Chennai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior consultant - lead pyspark cloud data engineer (chennai) / chennai

Subscribe to this job alert:

Get the latest job offers by email for: senior consultant - lead pyspark cloud data engineer (chennai) / chennai