11 Sep
|
ITC Infotech
|
Bengaluru
11 Sep
ITC Infotech
Bengaluru
Senior Big Data Engineer (AWS | PySpark | EMR)
? Location: Bangalore (Hybrid/Onsite)
? Experience: 10 to 16Years
? Relevant Experience: Minimum 8 Years in Data Engineering / Big Data
We are looking for a highly skilled Senior Big Data Engineer with strong hands-on experience in AWS, PySpark/Scala, Apache Spark, EMR, and enterprise-scale data processing.
Candidates should have proven experience building and optimizing large-scale distributed data platforms and must be comfortable working with complex ETL ecosystems handling massive data volumes.
Mandatory Skills
AWS (Must Have)
- Strong hands-on experience with AWS ecosystem
- Expert knowledge of:
- AWS EMR
- S3
- Glue
- Redshift
- Athena
- Lambda
- Step Functions
- CloudWatch
- Experience designing and managing large-scale AWS data platforms
Big Data Technologies (Must Have)
- Apache Spark
- PySpark and/or Scala
- Hadoop Ecosystem
- Hive
- HDFS
- Spark SQL
- Distributed Data Processing
Data Engineering (Must Have)
- Design and development of scalable ETL/ELT pipelines
- Batch and large-scale data processing
- Data Lake and Data Warehouse solutions
- Performance tuning of Spark applications
- Data quality and validation frameworks
Orchestration (Must Have)
- Apache Airflow
- AWS Step Functions
- Workflow orchestration
- DAG design and optimization
- Job scheduling and dependency management
- State management for enterprise data pipelines
Required Experience
- 8+ years of Data Engineering / Big Data experience
- Solid experience handling TB/PB-scale datasets
- Experience working with large Spark clusters
- Expertise in Spark optimization techniques:
- Partitioning
- Caching
- Broadcast joins
- Shuffle optimization
- Memory tuning
- Cluster tuning
- Experience solving distributed computing challenges
Preferred Skills
- Kafka
- Spark Streaming
- Databricks
- Snowflake
- CI/CD Pipelines
- Terraform
- Python
- SQL
- Cloud Migration Projects
Responsibilities
- Build and maintain scalable Big Data platforms on AWS
- Develop and optimize PySpark/Scala applications
- Design enterprise-grade ETL/ELT frameworks
- Create and maintain Airflow DAGs and orchestration workflows
- Optimize Spark jobs and EMR cluster performance
- Implement monitoring, alerting, and reliability solutions
- Work closely with business and analytics teams to deliver data solutions
- Troubleshoot production issues and improve platform stability
📌 Senior Big Data Engineer (AWS | PySpark | EMR) (Bengaluru)
🏢 ITC Infotech
📍 Bengaluru