21 Aug
|
Tata Consultancy Services
|
Mumbai
21 Aug
Tata Consultancy Services
Mumbai
Role & responsibilities
Data Engineer
Robust hands-on experience with PySpark (Spark SQL, Dataframes and RDD).
Experience with Hive, HDFS and Distributed Storage systems.
Hands-on experience with Hive (HiveQL, table optimization, partition management).
Solid knowledge of MySQL (indexes, query tuning, schema design, stored procedures).
Robust hands-on experience with PySpark and Spark optimization techniques (RDD, DataFrame APIs, Spark SQL, partitioning, caching).
Proficiency in Python (Experience on Pandas and Numpy) •Proficiency in SQL (MySQL, Oracle, SQL Server, PostgreSQL, etc.).
Hands on experience Building end to end Scalable Data Pipelines.
Optimize Spark jobs for performance, memory management and resource utilization.
Implement data validation, quality checks and error handling mechanism.
Experience with Data Warehousing concepts (Star schema, SCD, facts, dimensions).
Knowledge of ETL best practices, data modeling, and metadata management.
Experience working with:
Flat files (CSV, TXT)
Excel/XML/JSON
APIs / REST services
Cloud platforms (AWS, Azure, GCP) optional but preferred Experience with version control (Git, SVN).
Knowledge of Airflow/Oozie or other orchestration tools.
📌 Data Engineer Mumbai
🏢 Tata Consultancy Services
📍 Mumbai