Job Title: Lead PySpark Cloud Data Engineer
Work Location: Chennai / Hyderabad
Experience: 8 -12 years
Work Model: Hybrid (3 days WHO)
Shift Time: 1PM to 10PM / 3PM to 12PM
Skills Required: Pyspark, GCP, BigQuery, Dataproc
Share your resume to:
[email protected] enGen Global is an emerging global healthcare partner that delivers strategic innovation, expertise, and flexibility to its healthcare partners. Being a US healthcare conglomerate captive, we have direct access to deeper insights that help us accelerate our learning process and keeps us ahead of the curve. Thryve delivers next-generation solutions that enable our healthcare partners to provide positive experiences to their consumers.
Our global collaborative of healthcare, operations, and IT experts creates innovative and sustainable processes for our clients, which keeps the ever-evolving consumers engaged and assists them in managing the future of their healthcare better. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. Thryve is an equal chance employer and places a high value on integrity, diversity, and inclusion in the organization.
We do not discriminate based on any protected attribute. For more information about the organization, please visit www.thryvedigital.com
Role Summary
This job involves understand the overall requirement of the enterprise Data need and design & develop the robust, highly scalable & resilient data pipeline using Pyspark, Dataproc, open source table formats and other Google services.
Job will involve extensive interfacing & co-ordination with other senior tech folks (across India & US) and lead the design & development of the data ingestion pipelines
Essential Responsibilities
- Design, develop, and implement highly scalable, reliable, and performant data pipelines using PySpark, Dataproc, and other Google cloud-native technologies
- Create pipelines for data lakehouse architectures utilizing Apache Iceberg for open table formats and BigLake Metastore for unified metadata management
- Develop and maintain data transformation logic using Pyspark and / or dbt to create clean, consistent, and production-ready data models
- Work extensively with Google Cloud Platform (GCP) services, including Dataproc, BigQuery, Cloud Storage, and other relevant data services.
- Optimize data pipelines and queries for performance, efficiency, and cost-effectiveness.
- Implement and enforce data quality checks, monitoring, and governance best practices to ensure data integrity and reliability
- Provide technical guidance and mentorship to junior data engineers, fostering a culture of continuous learning and improvement
- Create and maintain comprehensive technical documentation for data pipelines, data models, and platform architecture
- Provide regular updates on the tasks,
status and risks to project manager The experience we are looking to add to our team
Required
- Bachelor’s degree or higher from a reputed university
- 8 to 12 years total experience with majority of that experience related to building high performant data pipelines – batch and streaming
- Strong Hands on experience / expert-level proficiency in PySpark for data processing, transformation, and analysis
- Extensive experience in implementing large scale data ingestion and curation solutions
- Hands-on experience with cloud-based data platforms, preferably Google Cloud Platform (GCP) and services like Dataproc, BigQuery, and Cloud Storage
- Knowledge in Apache Iceberg for open table formats
- Advanced SQL skills for data querying, manipulation, and optimization
- Proficiency with Git and collaborative development workflows
- Excellent analytical and problem-solving skills with a keen attention to detail
- Strong communication and interpersonal skills, with the ability to explain complex technical concepts to both technical and non-technical audiences
Good to have
- Expertise in Google Cloud services
- Proven experience with Apache Iceberg for open table formats and BigLake Metastore for unified metadata management
- Strong experience with dbt (data build tool) for data modeling, transformation, and orchestration
- Experience in Data Governance and Data quality tools like Atlan, Monte Carlo etc.
- Experience in processing streaming data using Kafka / Pub-Sub
- Healthcare industry experience
- Experience in Agile
📌 Lead Pyspark Cloud Data Engineer (Hyderabad)
🏢 enGen Global
📍 Hyderabad