Job Description
Candidates ready to join immediately can share their details via email for quick processing.
n
? CCTC | ECTC | Notice Period | Location Preference
n
[email protected]
n
Act fast for immediate attention! ⏳?
n
___________________________________________________________________________________________________
n
Must-Have Skills
n
n
- 6+ years of overall experience in Data Engineering / Big Data.
n
- Strong understanding of Big Data concepts and architecture.
n
- Strong hands-on experience with Apache Spark.
n
- Expertise in:
n
- Spark Performance Tuning
n
- Spark Optimization
n
- Query/Job Performance Improvement
n
- Troubleshooting Spark workloads
n
- Strong hands-on experience with PySpark and Spark.
n
- Strong programming experience in Python.
n
- Valuable experience working with MySQL / SQL.
n
- Strong experience in designing and developing Data Pipelines.
n
- Hands-on experience with Apache Airflow for data pipeline orchestration and scheduling.
n
- Experience working with at least one Cloud Platform.
n
- GCP experience is preferred.
n
- Good understanding of CI/CD and DevOps concepts.
n
- Experience integrating data engineering workloads with CI/CD pipelines.
n
- Strong debugging, troubleshooting, and problem-solving skills.
n
n
Preferred Skills
n
n
- Hands-on exposure to relevant GCP data services.
n
- Experience handling large-scale and high-volume datasets.
n
- Understanding of distributed data processing and data architecture.
n
- Experience improving the scalability, reliability, and performance of data pipelines.
n
- Exposure to Agile development and DevOps practices.
n
n
Key Responsibilities
n
n
- Design, develop, and maintain scalable Big Data and Data Engineering solutions.
n
- Develop data processing applications using Python, PySpark, and Apache Spark.
n
- Perform Spark performance tuning and optimization for large-scale workloads.
n
- Build, maintain, and monitor robust ETL/ELT data pipelines.
n
- Develop and manage workflow orchestration using Apache Airflow.
n
- Work with MySQL/SQL for data extraction, transformation, and validation.
n
- Deploy and support data engineering solutions in cloud environments, preferably GCP.
n
- Work with DevOps teams to implement and maintain CI/CD pipelines.
n
- Troubleshoot production issues and optimize data processing performance.
n
- Collaborate with engineering and business teams to deliver reliable and scalable data solutions.
n
n
Primary Skill Combination
n
Big Data + Apache Spark + PySpark + Python + Airflow + SQL/MySQL + Cloud (GCP Preferred) + CI/CD
n
Mandatory Focus: Strong hands-on Apache Spark performance tuning and optimization experience.
n
📌 6 + YoE - Data Engineer - Big Data / PySpark - any UST Location - Immediate Joiner (Bengaluru)
🏢 UST
📍 Bengaluru