Tech Mahindra Hiring For Bigdata With Python and Pyspark (BLR/Chn) (Bengaluru)

Tech Mahindra Hiring For Bigdata With Python and Pyspark (BLR/Chn) (Bengaluru)

01 Oct
|
Glauben Technologies
|
Bengaluru

01 Oct

Glauben Technologies

Bengaluru

Role & responsibilities

Lead PySpark Engineer (Distributed

Processing Platform)

Role Overview

We are seeking a highly experienced Lead PySpark Engineer to design, build, optimize, and lead the development of large-scale distributed processing platform applications using Apache Spark.

This is a hands-on technical leadership role focused on architecting high-performance PySpark solutions capable of processing billions of records with low latency and optimal resource utilization.

This role is not intended for ETL-only developers. We are looking for engineers who deeply understand Spark internals and can build highly optimized distributed processing systems at enterprise scale.

Required Skills

- PySpark
- Python
- Apache Spark internals and distributed systems architecture
- Delta Lake , Parquet , Trino
- Linux , vi commands
- Workflow orchestration (Airflow)
- Minio
- Containerized deployment environments (Kubernetes)
- Git , Azure ADO
- Vscode
- Github Copilot

Preferred / Nice-to-Have
- Trade surveillance domain experience
- Financial services data engineering experience
- Time-series analytics
- Market data processing

Ideal Candidate Profile The ideal candidate is a Spark architecture expert with proven real-world experience in:
- Distributed computing

- Memory and execution optimization

- Performance tuning at scale

- Framework development

- Scalable data platform engineering

You thrive in high-scale environments, combine deep technical execution with leadership, and can deliver robust Spark systems for mission-critical workloads.

Key Responsibilities

1. Platform & Application Engineering

- Design and develop enterprise-scale distributed data processing applications using PySpark.
- Build reusable Spark frameworks and libraries that can be shared across multiple projects.
- Architect scalable batch and near real-time processing pipelines.
- Develop modular, configurable, and highly maintainable Spark applications.




- Establish coding standards and Spark engineering best practices across the team.

1. Performance Engineering (End-to-End Ownership)

Own Spark performance optimization across:
- Memory optimization
- CPU utilization optimization
- Shuffle reduction
- Stage optimization
- DAG optimization
- Data skew handling
- Job parallelism tuning
- Executor sizing
- Garbage Collection (GC) optimization

Be able to identify and resolve bottlenecks using:
- Spark UI
- Explain plans
- Event logs
- Executor/stage-level metrics

1. Distributed Processing Expertise

Demonstrate strong command of:
- Spark execution model

- Driver vs Executor architecture

- Lazy evaluation

- Catalyst optimizer

- Lineage and fault tolerance

- Adaptive Query Execution (AQE)

1. Data Partitioning & Data Layout

Hands-on expertise with:
- Partition strategy design
- Partition pruning
- Static partitioning
- Energetic partitioning

1. Join & Window Optimization

Extensive implementation and optimization experience for:
- Large-to-large dataset joins
- Window joins
- Time-series/time-based joins
- Interval joins

Must be able to select join strategies based on:
- Data volume
- Data skew
- Cluster behavior and resource profile

Strong window processing capabilities:
- Running aggregations
- Sliding windows
- Rolling windows
- Ranking
- Look-back calculations
- Time-based analytics
- Complex business-rule implementation

1. Spark Optimization Techniques

Applied experience with:
- Cache/Persist

- Checkpointing

- Compression

- Predicate pushdown

- Bucketing
- Z-ordering

1. Data Engineering Practices

- Delta Lake design and optimization
- Schema evolution and governance

1. Technical Leadership

- Lead architecture discussions for distributed Spark platforms.
- Conduct Spark code reviews and design reviews.
- Enforce engineering standards and quality practices.
- Mentor junior and senior engineers.
- Drive technical decisions across the data platform.
- Troubleshoot and lead resolution of critical production incidents.

📌 Tech Mahindra Hiring For Bigdata With Python and Pyspark (BLR/Chn) (Bengaluru)
🏢 Glauben Technologies
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: tech mahindra hiring for bigdata with python and pyspark (blr/chn) (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: tech mahindra hiring for bigdata with python and pyspark (blr/chn) (bengaluru) / bengaluru