14 Sep
|
ITC Infotech
|
Bengaluru
14 Sep
ITC Infotech
Bengaluru
Senior Big Data Engineer (AWS | PySpark | EMR)
Location: Bangalore (Hybrid/Onsite)
Experience: 10 to 16Years
Relevant Experience: Minimum 8 Years in Data Engineering / Big Data
Job Description
We are looking for a highly skilled Senior Big Data Engineer with strong hands-on experience in AWS, PySpark/Scala, Apache Spark, EMR, and enterprise-scale data processing .
Candidates should have proven experience building and optimizing large-scale distributed data platforms and must be comfortable working with complex ETL ecosystems handling massive data volumes.
Mandatory Skills
AWS (Must Have)
Robust hands-on experience with AWS ecosystem
Expert knowledge of:
AWS EMR
S3
Glue
Redshift
Athena
Lambda
Step Functions
CloudWatch
Experience designing and managing large-scale AWS data platforms
Big Data Technologies (Must Have)
Apache Spark
PySpark and/or Scala
Hadoop Ecosystem
Hive
HDFS
Spark SQL
Distributed Data Processing
Data Engineering (Must Have)
Design and development of scalable ETL/ELT pipelines
Batch and large-scale data processing
Data Lake and Data Warehouse solutions
Performance tuning of Spark applications
Data quality and validation frameworks
Orchestration (Must Have)
Apache Airflow
AWS Step Functions
Workflow orchestration
DAG design and optimization
Job scheduling and dependency management
State management for enterprise data pipelines
Required Experience
8+ years of Data Engineering / Big Data experience
Strong experience handling TB/PB-scale datasets
Experience working with large Spark clusters
Expertise in Spark optimization techniques:
Partitioning
Caching
Broadcast joins
Shuffle optimization
Memory tuning
Cluster tuning
Experience solving distributed computing challenges
Preferred Skills
Kafka
Spark Streaming
Databricks
Snowflake
CI/CD Pipelines
Terraform
Python
SQL
Cloud Migration Projects
Responsibilities
Build and maintain scalable Big Data platforms on AWS
Develop and optimize PySpark/Scala applications
Design enterprise-grade ETL/ELT frameworks
Create and maintain Airflow DAGs and orchestration workflows
Optimize Spark jobs and EMR cluster performance
Implement monitoring, alerting, and reliability solutions
Work closely with business and analytics teams to deliver data solutions
Troubleshoot production issues and improve platform stability
📌 Senior Big Data Engineer (AWS | PySpark | EMR) (Bengaluru)
🏢 ITC Infotech
📍 Bengaluru