PySpark & Dataproc Tech Lead (India)

PySpark & Dataproc Tech Lead (India)

03 Aug
|
Hoonartek
|
India

03 Aug

Hoonartek

India

Hybrid, Bangalore, Pune

About Us

We empower enterprises globally through intelligent, creative, and insightful services for data integration, data analytics and data visualization.

Hoonartek is a leader in enterprise transformation, data engineering and an acknowledged world-class Ab Initio delivery partner.

Using centuries of cumulative experience, research and leadership, we help our clients eliminate the complexities & risk of legacy modernization and safely deliver big data hubs, operational data integration, business intelligence, risk & compliance solutions and traditional data warehouses & marts.

At Hoonartek, we work to ensure that our customers, partners and employees all benefit from our unstinting commitment to delivery, quality and value. Hoonartek is increasingly the choice for customers seeking a trusted partner of vision, value and integrity

How We Work?

Define, Design and Deliver (D3) is our in-house delivery philosophy. It’s culled from agile and rapid methodologies and focused on ‘just enough design’. We embrace this philosophy in everything we do, leading to numerous client success stories and indeed to our own success.

We embrace change, empowering and trusting our people and building long and valuable relationships with our employees, our customers and our partners. We work flexibly, even adopting traditional/waterfall methods where circumstances demand it. At Hoonartek, the focus is always on delivery and value.

Job Description

Roles & Responsibilities

-
Drive solution design and architecture for large-scale data processing platforms
- Define and standardize migration, reconfiguration, and modernization approaches for data workloads
- Lead implementation of platform transitions (e.g., containerized environments to managed services like Dataproc)




- Establish frameworks for performance benchmarking and optimization
- Provide technical guidance for:

Code conversion strategies

Performance tuning

Issue resolution and debugging

- Oversee execution to ensure:

Code quality

Scalability

Reliability

- Define best practices and governance for:

Spark development

Resource utilization

Cost management

- Collaborate with stakeholders to translate requirements into technical solutions
- Identify risks, dependencies, and mitigation strategies
- Mentor engineering teams and ensure adherence to architectural standards
- Drive continuous improvement through optimization and standardization initiatives

Primary Skillsets

-
Strong expertise in Apache Spark / PySpark (development and architecture)
- Hands-on experience with GCP Dataproc and cloud-based data platforms
- Deep understanding of distributed systems and large-scale data processing
- Experience in designing and implementing data platform migrations
- Proven expertise in performance engineering and optimization frameworks
- Experience with Kubernetes / GKE environments
- Solid knowledge of GCP ecosystem (BigQuery, GCS, IAM, networking)
- Familiarity with automation tools and migration framework

Secondary Skillsets

- Exposure to cost governance / FinOps practices
- Strong stakeholder communication and leadership skills
- Experience working in Agile delivery models

Job Requirement





Roles & Responsibilities

- Drive solution design and architecture for large-scale data processing platforms
- Define and standardize migration, reconfiguration, and modernization approaches for data workloads
- Lead implementation of platform transitions (e.g., containerized environments to managed services like Dataproc)
- Establish frameworks for performance benchmarking and optimization
- Provide technical guidance for:

Code conversion strategies

Performance tuning

Issue resolution and debugging

- Oversee execution to ensure:

Code quality

Scalability

Reliability

- Define best practices and governance for:

Spark development

Resource utilization

Cost management

- Collaborate with stakeholders to translate requirements into technical solutions
- Identify risks, dependencies, and mitigation strategies
- Mentor engineering teams and ensure adherence to architectural standards
- Drive continuous improvement through optimization and standardization initiatives

Primary Skillsets

- Strong expertise in Apache Spark / PySpark (development and architecture)
- Hands-on experience with GCP Dataproc and cloud-based data platforms
- Deep understanding of distributed systems and large-scale data processing
- Experience in designing and implementing data platform migrations
- Proven expertise in performance engineering and optimization frameworks
- Experience with Kubernetes / GKE environments
- Strong knowledge of GCP ecosystem (BigQuery, GCS, IAM, networking)
- Familiarity with automation tools and migration framework

Secondary Skillsets

- Exposure to cost governance / FinOps practices
- Strong stakeholder communication and leadership skills
- Experience working in Agile delivery models

📌 PySpark & Dataproc Tech Lead (India)
🏢 Hoonartek
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: pyspark & dataproc tech lead (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: pyspark & dataproc tech lead (india) / india