16 Sep
|
Tata Consultancy Services
|
Chennai
16 Sep
Tata Consultancy Services
Chennai
Job DescriptionRole: Pyspark Experience: 5-8 yrs Location: Chennai/Kolkata/Mumbai/Pune/Hyd/Bangalore Responsibility: - Design scalable PySpark-based test architectures for ETL/data pipelines, including modular frameworks for batch processing. - Architect end-to-end data validation systems in Hadoop environment for lineage, schema evolution - Lead system design for Hadoop/Hive test environments, including YARN resource management, agile partitioning. - Exposure to Zephyr-Jira-ServiceNow integrated test management systems with experience on API-driven automation. - Design CI/CD test pipelines for PySpark/Hadoop jobs, incorporating artifact management, parallel execution, and blue-green deployments. - Create data quality system designs using PySpark integrated with Hive metadata services. - Design testing platforms, test data generators - Mentor juniors on PySpark testing basics, contribute to testing strategy discussions - Spark session configurations for memory and core allocations for both local and cluster manager settings - Data handling with distributed file systems like HDFS and writing back to hive tables - Implementation of Partitioning, caching techniques in organizing code for transformation pipelines
📌 Pyspark (Chennai)
🏢 Tata Consultancy Services
📍 Chennai