PySPark /Spark & Dataproc Senior Engineer (India)

PySPark /Spark & Dataproc Senior Engineer (India)

03 Aug
|
Hoonartek
|
India

03 Aug

Hoonartek

India

Hybrid, Bangalore, Pune

About Us

We empower enterprises globally through intelligent, creative, and insightful services for data integration, data analytics and data visualization.

Hoonartek is a leader in enterprise transformation, data engineering and an acknowledged world-class Ab Initio delivery partner.

Using centuries of cumulative experience, research and leadership, we help our clients eliminate the complexities & risk of legacy modernization and safely deliver big data hubs, operational data integration, business intelligence, risk & compliance solutions and traditional data warehouses & marts.

At Hoonartek, we work to ensure that our customers, partners and employees all benefit from our unstinting commitment to delivery, quality and value. Hoonartek is increasingly the choice for customers seeking a trusted partner of vision, value and integrity

How We Work?

Define, Design and Deliver (D3) is our in-house delivery philosophy. It’s culled from agile and rapid methodologies and focused on ‘just enough design’. We embrace this philosophy in everything we do, leading to numerous client success stories and indeed to our own success.

We embrace change, empowering and trusting our people and building long and valuable relationships with our employees, our customers and our partners. We work flexibly, even adopting traditional/waterfall methods where circumstances demand it. At Hoonartek, the focus is always on delivery and value.

Job Description

Roles & Responsibilities

-
Design, develop, and maintain PySpark-based data processing pipelines
- Perform migration and reconfiguration of Spark workloads across different execution environments (e.g., containers to managed platforms like Dataproc)




- Execute and validate automated code conversion or migration utilities
- Analyze and resolve code compatibility and runtime issues during platform transitions
- Perform performance tuning and optimization of Spark jobs for:

Execution efficiency

Resource utilization

Cost effectiveness

- Support benchmarking and performance analysis across environments
- Troubleshoot failures and ensure stable and reliable job execution
- Adhere to best practices for:

Distributed processing

Code modularity and maintainability

- Collaborate with cross-functional teams for integration and deployment
- Maintain documentation for code changes, configurations, and known issues

Primary Skillsets

-
Strong expertise in PySpark / Apache Spark
- Experience with distributed data processing frameworks
- Hands-on with Spark performance tuning and optimization techniques
- Proficiency in:

Spark SQL

DataFrames / RDD operations

- Understanding of Spark internals:

Execution plans, partitioning, shuffle operations

- Experience with GCP services (Dataproc, GCS, BigQuery)
- Exposure to containerized environments (Kubernetes / GKE)
- Familiarity with automation scripting (Python, Shell)

Secondary Skillsets

-
Knowledge of cloud cost optimization principles
- Experience with logging, monitoring, and debugging tools




- Basic understanding of CI/CD pipelines

Job Requirement

Roles & Responsibilities

- Design, develop, and maintain PySpark-based data processing pipelines
- Perform migration and reconfiguration of Spark workloads across different execution environments (e.g., containers to managed platforms like Dataproc)
- Execute and validate automated code conversion or migration utilities
- Analyze and resolve code compatibility and runtime issues during platform transitions
- Perform performance tuning and optimization of Spark jobs for:

Execution efficiency

Resource utilization

Cost effectiveness

- Support benchmarking and performance analysis across environments
- Troubleshoot failures and ensure secure and reliable job execution
- Adhere to best practices for:

Distributed processing

Code modularity and maintainability

- Collaborate with cross-functional teams for integration and deployment
- Maintain documentation for code changes, configurations, and known issues

Primary Skillsets

- Strong expertise in PySpark / Apache Spark
- Experience with distributed data processing frameworks
- Hands-on with Spark performance tuning and optimization techniques
- Proficiency in:

Spark SQL

DataFrames / RDD operations

- Understanding of Spark internals:

Execution plans, partitioning, shuffle operations

- Experience with GCP services (Dataproc, GCS, BigQuery)
- Exposure to containerized environments (Kubernetes / GKE)
- Familiarity with automation scripting (Python, Shell)

Secondary Skillsets

- Knowledge of cloud cost optimization principles
- Experience with logging, monitoring, and debugging tools
- Basic understanding of CI/CD pipelines

📌 PySPark /Spark & Dataproc Senior Engineer (India)
🏢 Hoonartek
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: pyspark /spark & dataproc senior engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: pyspark /spark & dataproc senior engineer (india) / india