TCS is hiring for Pyspark Data engineer(Spark & Python)
Desired Experience: 6 years +
Job Location: Chennai / Hyderabad / Pune / Kolkata / Bangalore
Must Have skills:-
6+ Years of experience in Pyspark, Python, Spark.
Design scalable PySpark-based test architectures for ETL/data pipelines, including modular frameworks for batch processing.
Architect end-to-end data validation systems in Hadoop setting for lineage, schema evolution
Lead system design for Hadoop/Hive test environments, including YARN resource management, energetic partitioning.
Exposure to Zephyr-Jira-ServiceNow integrated test management systems with experience on API-driven automation.
Design CI/CD test pipelines for PySpark/Hadoop jobs, incorporating artifact management, parallel execution, and blue-green deployments.
Create data quality system designs using PySpark integrated with Hive metadata services.
Design testing platforms, test data generators
Execute Linux commands in hive
Spark session configurations for memory and core allocations for both local and cluster manager settings
Data handling with distributed file systems like HDFS and writing back to hive tables
Implementation of Partitioning, caching techniques in organizing code for transformation pipelines
Performance tuning implementation like salting, minimizing shuffling