23 Aug
|
Tata Consultancy Services
|
Chennai
23 Aug
Tata Consultancy Services
Chennai
Role- PySpark Developer
Year of Experience- 4 to 15 years
Location -Pune, Hyderabad, Chennai, Bangalore
Technical Skills:
1. Pyspark
2. Python concepts and Framework
3. Spark Architecture
4. Big Data
5. SQL
:
Job Requirements*
- Good work experience on Big Data Platforms like Hadoop, PySpark, Scala, Hive, Impala, SQL, Python
- Good Python, Pyspark, Big Data experience
- Spark UI/Optimization/debugging techniques
- Good python scripting skills
- Intermediate SQL exposure – Subquery, Joins, CTE’s
- Database technologies
Key Responsibilities
- *Design scalable PySpark-based test architectures for ETL/data pipelines, including modular frameworks for batch processing
- .Architect end-to-end data validation systems in Hadoop environment for lineage, schema evolutio
- nLead system design for Hadoop/Hive test environments, including YARN resource management, agile partitioning
- .Exposure to Zephyr-Jira-ServiceNow integrated test management systems with experience on API-driven automation
- .Design CI/CD test pipelines for PySpark/Hadoop jobs, incorporating artifact management, parallel execution, and blue-green deployments
- .Create data quality system designs using PySpark integrated with Hive metadata services
- .Design testing platforms, test data generator
- sMentor juniors on PySpark testing basics, contribute to testing strategy discussion
- sSpark session configurations for memory and core allocations for both local and cluster manager setting
- sData handling with distributed file systems like HDFS and writing back to hive table
- sImplementation of Partitioning, caching techniques in organizing code for transformation pipeline
- sPerformance tuning implementation like salting, minimizing shufflin
g
📌 Pyspark developer (Chennai)
🏢 Tata Consultancy Services
📍 Chennai