Job DescriptionDear Professionals
nGreetings from Tata consultancy Services,
nJob Title Pyspark Data Engineer
nExperiernce: 6-10 Years
nLocation: Chennai / Kolkata / Hyderabad / Pune
nMode of Work : Work from Office
nJob description
n
- n
- Design scalable PySpark-based test architectures for ETL/data pipelines, including modular frameworks for batch processing.n
- Architect end-to-end data validation systems in Hadoop setting for lineage, schema evolutionn
- Lead system design for Hadoop/Hive test environments, including YARN resource management, dynamic partitioning.n
- Exposure to Zephyr-Jira-ServiceNow integrated test management systems with experience on API-driven automation.n
- Design CI/CD test pipelines for PySpark/Hadoop jobs, incorporating artifact management, parallel execution, and blue-green deployments.n
- Create data quality system designs using PySpark integrated with Hive metadata services.n
- Design testing platforms, test data generatorsn
- Mentor juniors on PySpark testing basics, contribute to testing strategy discussionsn
- Spark session configurations for memory and core allocations for both local and cluster manager settingsn
- Data handling with distributed file systems like HDFS and writing back to hive tablesn
- Implementation of Partitioning, caching techniques in organizing code for transformation pipelinesn
- Performance tuning implementation like salting, minimizing shufflingn