26 Aug
|
Datagaps
|
Hyderabad
26 Aug
Datagaps
Hyderabad
Role Summary
We are seeking a highly skilled Apache Spark Developer / Spark Performance Engineer with 4+ years of hands-on Spark experience to design, develop, optimize, and maintain large-scale data processing applications. The ideal candidate should possess strong expertise in Spark performance tuning, memory optimization, ETL pipelines, and distributed data processing while working across contemporary data lake and big data platforms.
Key Responsibilities
- Design, develop, and maintain scalable data processing pipelines using Apache Spark.
- Build and optimize batch and real-time data processing solutions utilizing Spark SQL and Spark Streaming.
- Analyze and improve Spark job performance through effective memory management, query optimization, and resource utilization.
- Troubleshoot performance bottlenecks, data skew issues, shuffle inefficiencies, and executor failures.
- Develop robust ETL/ELT workflows integrating data from multiple enterprise systems.
- Work with large-scale datasets stored in Data Lakes, Hadoop, HDFS, and cloud-based storage platforms.
- Collaborate with Data Engineers, Architects, and Business Teams to understand requirements and design optimal data solutions.
- Monitor, maintain, and improve Spark applications using Spark UI, Spark History Server, and cluster monitoring tools.
- Implement best practices for partitioning, caching, persistence, and resource allocation.
- Optimize Spark workloads running on Kubernetes and cloud-based environments.
- Ensure code quality through reviews, testing, documentation, and adherence to development standards.
- Support production deployments and provide performance tuning recommendations for existing Spark workloads.
Required Skills & Experience
Experience
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
- Minimum 4+ years of hands-on experience in Apache Spark Development.
- Experience developing enterprise-grade Big Data and Data Engineering solutions.
Core Apache Spark Expertise
- Strong understanding of Spark Architecture and Distributed Computing concepts.
- Hands-on experience with:
- RDD
- DataFrame API
- Dataset API
- Spark SQL
- Spark Streaming / Structured Streaming
- Spark Session & Spark Context
- DAG Scheduler
- Deep understanding of:
- Partitioning Strategies
- Repartition and Coalesce
- Broadcast Join
- Shuffle Join
- Query Optimization
Spark Performance Tuning & Monitoring (Mandatory)
- Strong expertise in:
- Executor Memory Tuning
- Driver Memory Optimization
- Garbage Collection (GC) Tuning
- Dynamic Resource Allocation
- Shuffle Optimization
- Spill-to-Disk Reduction Techniques
- Hands-on experience using:
- Spark UI
- Spark History Server
- Cluster Monitoring Tools
- Experience handling:
- Data Skew
- Executor Failures
- Performance Bottlenecks
- Memory Leaks
- Expertise in:
- Caching and Persistence Strategies
- MEMORY_ONLY
- MEMORY_AND_DISK
- Storage Optimization
Data Engineering & Storage
- Strong experience building ETL and ELT pipelines.
- Hands-on experience with:
- Data Lakes
- Delta Lake
- Hadoop Ecosystem
- HDFS
- Hive
- Experience working with file formats:
- Parquet
- ORC
- Avro
- Understanding of Data Warehousing concepts and large-scale data processing.
Cloud & Containerization
- Experience running Spark workloads on Kubernetes.
- Understanding of Spark-on-Kubernetes architecture and performance tuning.
- Exposure to cloud platforms such as AWS, Azure, or GCP is preferred.
Nice-to-Have Skills
Databricks
- Hands-on experience with:
- Databricks Platform
- Delta Lake
- Unity Catalog
- Databricks Workflows
- Databricks Jobs
- Databricks Notebooks
- Experience optimizing Spark workloads within Databricks environments.
Additional Skills
- Knowledge of Airflow or workflow orchestration tools.
- Experience with CI/CD pipelines and DevOps practices.
- Familiarity with Kafka and real-time data ingestion frameworks.
- Understanding of data governance and data quality practices.
Preferred Candidate Profile
- Strong analytical and problem-solving skills.
- Ability to diagnose and resolve complex Spark performance issues.
- Excellent communication and stakeholder management skills.
- Experience working in Agile/Scrum environments.
- Self-driven individual capable of working independently and within a team.
📌 Spark Expert/ Spark Developer (Hyderabad)
🏢 Datagaps
📍 Hyderabad