15 Sep
|
ITC Infotech
|
Bengaluru
15 Sep
ITC Infotech
Bengaluru
Job Description
Senior Big Data Engineer (AWS | PySpark | EMR)
n
Location: Bangalore (Hybrid/Onsite)
n
Experience: 10 to 16Years
n
Relevant Experience: Minimum 8 Years in Data Engineering / Big Data
n
Job Description
n
We are looking for a highly skilled Senior Big Data Engineer with strong hands-on experience in AWS, PySpark/Scala, Apache Spark, EMR, and enterprise-scale data processing.
n
Candidates should have proven experience building and optimizing large-scale distributed data platforms and must be comfortable working with complex ETL ecosystems handling massive data volumes.
n
Mandatory Skills
n
AWS (Must Have)
n
n
- Robust hands-on experience with AWS ecosystem
n
- Expert knowledge of:
n
- AWS EMR
n
- S3
n
- Glue
n
- Redshift
n
- Athena
n
- Lambda
n
- Step Functions
n
- CloudWatch
n
- Experience designing and managing large-scale AWS data platforms
n
n
Big Data Technologies (Must Have)
n
n
- Apache Spark
n
- PySpark and/or Scala
n
- Hadoop Ecosystem
n
- Hive
n
- HDFS
n
- Spark SQL
n
- Distributed Data Processing
n
n
Data Engineering (Must Have)
n
n
- Design and development of scalable ETL/ELT pipelines
n
- Batch and large-scale data processing
n
- Data Lake and Data Warehouse solutions
n
- Performance tuning of Spark applications
n
- Data quality and validation frameworks
n
n
Orchestration (Must Have)
n
n
- Apache Airflow
n
- AWS Step Functions
n
- Workflow orchestration
n
- DAG design and optimization
n
- Job scheduling and dependency management
n
- State management for enterprise data pipelines
n
n
Required Experience
n
n
- 8+ years of Data Engineering / Big Data experience
n
- Strong experience handling TB/PB-scale datasets
n
- Experience working with large Spark clusters
n
- Expertise in Spark optimization techniques:
n
- Partitioning
n
- Caching
n
- Broadcast joins
n
- Shuffle optimization
n
- Memory tuning
n
- Cluster tuning
n
- Experience solving distributed computing challenges
n
n
Preferred Skills
n
n
- Kafka
n
- Spark Streaming
n
- Databricks
n
- Snowflake
n
- CI/CD Pipelines
n
- Terraform
n
- Python
n
- SQL
n
- Cloud Migration Projects
n
n
Responsibilities
n
n
- Build and maintain scalable Big Data platforms on AWS
n
- Develop and optimize PySpark/Scala applications
n
- Design enterprise-grade ETL/ELT frameworks
n
- Create and maintain Airflow DAGs and orchestration workflows
n
- Optimize Spark jobs and EMR cluster performance
n
- Implement monitoring, alerting, and reliability solutions
n
- Work closely with business and analytics teams to deliver data solutions
n
- Troubleshoot production issues and improve platform stability
n
📌 Senior Big Data Engineer (AWS | PySpark | EMR) (Bengaluru)
🏢 ITC Infotech
📍 Bengaluru