20 Aug
|
Tata Consultancy Services
|
Mumbai
20 Aug
Tata Consultancy Services
Mumbai
Role & responsibilities
Data Engineer
- Strong hands-on experience with PySpark (Spark SQL, Dataframes and RDD).
- Experience with Hive, HDFS and Distributed Storage systems.
- Hands-on experience with Hive (HiveQL, table optimization, partition management).
- Strong knowledge of MySQL (indexes, query tuning, schema design, stored procedures).
- Robust hands-on experience with PySpark and Spark optimization techniques (RDD, DataFrame APIs, Spark SQL, partitioning, caching).
- Proficiency in Python (Experience on Pandas and Numpy) •Proficiency in SQL (MySQL, Oracle, SQL Server, PostgreSQL, etc.).
- Hands on experience Building end to end Scalable Data Pipelines.
- Optimize Spark jobs for performance, memory management and resource utilization.
- Implement data validation, quality checks and error handling mechanism.
- Experience with Data Warehousing concepts (Star schema, SCD, facts, dimensions).
- Knowledge of ETL best practices, data modeling, and metadata management.
- Experience working with:
- Flat files (CSV, TXT)
- Excel/XML/JSON
- APIs / REST services
- Cloud platforms (AWS, Azure, GCP) optional but preferred Experience with version control (Git, SVN).
- Knowledge of Airflow/Oozie or other orchestration tools.
📌 Data Engineer (Mumbai)
🏢 Tata Consultancy Services
📍 Mumbai