21 Aug
|
Tata Consultancy Services
|
Bengaluru
21 Aug
Tata Consultancy Services
Bengaluru
Job Locations : Bengaluru, Chennai, Hyderabad, Pune, Kolkata
Job Requirements:
- Experience with strong proficiency in Python and PySpark.
- Experience working on Spark SQL, RDD, and DataFrame APIs.
- Good understanding of Hadoop ecosystem (Hive, HDFS, YARN).
- Knowledge of data formats like Parquet, Avro, JSON, etc.
- Experience with SQL and writing productive queries.
- Familiarity with job orchestration tools (Airflow, Oozie, or similar).
- Version control systems like Git.
- Exposure to cloud platforms (AWS/GCP/Azure) is a plus.
Key responsibilities:
- Design, develop, and maintain robust ETL/ELT pipelines using PySpark and other big data technologies. Optimize Spark jobs for performance and scalability.
- Implement data transformations, aggregations, and joins over large datasets.
Perform batch and real-time data processing tasks.
- Collaborate with data scientists, analysts, and other engineers to understand requirements and deliver quality solutions. Integrate PySpark solutions with data warehouses (like Hive, Redshift, Snowflake) and other data stores.
- Write clean, maintainable, and well-documented code. Follow version control and CI/CD practices using tools like Git, Jenkins, or Azure DevOps.
- Troubleshoot data quality and performance issues. Monitor pipeline health and ensure SLAs are met.
- Work with tools and platforms like Hadoop, Hive, HDFS, AWS EMR, Databricks, or Azure Synapse as required.
📌 Pyspark Developer (Bengaluru)
🏢 Tata Consultancy Services
📍 Bengaluru