- Design and build scalable big data processing applications.
- Strong experience in Apache Spark, Scala, and distributed data processing; work on building high-performance data pipelines for analytics and data engineering use cases.
- Develop and maintain data pipelines using Apache Spark (Scala).
- Process large-scale datasets in distributed environments.
- Implement batch and real-time data processing solutions.
- Write efficient and scalable code using Scala.
- Work extensively with Spark Core, Spark SQL, and Spark Streaming.
- Optimize Spark jobs for performance and resource utilization.
- Ingest data from various sources: Databases (RDBMS, NoSQL), APIs, File systems (HDFS, S3).
- Build ETL/ELT pipelines and data transformation workflows.