22 Aug
|
Tata Consultancy Services
|
Bengaluru
22 Aug
Tata Consultancy Services
Bengaluru
Job Locations : Bengaluru, Chennai, Hyderabad, Pune, Kolkata
Job Requirements:
Experience with robust proficiency in Python and PySpark.
Experience working on Spark SQL, RDD, and DataFrame APIs.
Valuable understanding of Hadoop ecosystem (Hive, HDFS, YARN).
Knowledge of data formats like Parquet, Avro, JSON, etc.
Experience with SQL and writing productive queries.
Familiarity with job orchestration tools (Airflow, Oozie, or similar).
Version control systems like Git.
Exposure to cloud platforms (AWS/GCP/Azure) is a plus.
Key responsibilities:
Design, develop, and maintain robust ETL/ELT pipelines using PySpark and other big data technologies. Optimize Spark jobs for performance and scalability.
Implement data transformations, aggregations, and joins over large datasets.
Perform batch and real-time data processing tasks.
Collaborate with data scientists, analysts, and other engineers to understand requirements and deliver quality solutions. Integrate PySpark solutions with data warehouses (like Hive, Redshift, Snowflake) and other data stores.
Write clean, maintainable, and well-documented code. Follow version control and CI/CD practices using tools like Git, Jenkins, or Azure DevOps.
Troubleshoot data quality and performance issues. Monitor pipeline health and ensure SLAs are met.
Work with tools and platforms like Hadoop, Hive, HDFS, AWS EMR, Databricks, or Azure Synapse as required.
📌 Pyspark Developer Bengaluru
🏢 Tata Consultancy Services
📍 Bengaluru