1. Design and build scalable data pipelines using Apache Spark & Scala (mandatory skills).
2. Develop and optimize ETL workflows leveraging Spark, Python, and Hive for large-scale data processing.
3. Integrate real-time data streaming solutions using Kafka and related frameworks.
4. Work with relational databases like Oracle DB for data modeling and performance tuning.
5. Develop backend services using Java, Spring Framework, and Microservices architecture.
6. Implement distributed communication using Apache Thrift and service integration patterns.
7. Ensure code quality through unit and integration testing using TestNG (groups, assertions).
8. Collaborate with cross-functional teams to translate business requirements into technical solutions.
9. Optimize data performance, scalability, and reliability across batch and streaming systems.
Follow best practices in CI/CD, version control, and production support for data platforms.