14 Aug
|
Tata Consultancy Services
|
Bengaluru
14 Aug
Tata Consultancy Services
Bengaluru
Job Title: Data Engineer (Python, PySpark, SQL, ETL)Role & responsibilities
Preferred candidate profile
We are seeking a skilled and motivated Data Engineer with strong expertise in Python, PySpark, SQL, and ETL development to design, build, and maintain scalable data pipelines and data processing solutions. The ideal candidate will have experience working with large datasets, data warehouses, and cloud-based data platforms, ensuring high-quality, reliable, and effective data delivery for analytics and business intelligence needs.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Python, PySpark, and SQL.
- Extract, transform, and load data from multiple source systems into data lakes and data warehouses.
- Develop and optimize PySpark jobs for large-scale data processing.
- Write complex SQL queries, stored procedures, and performance-tuned database solutions.
- Implement data quality checks, validation rules, and monitoring processes.
- Collaborate with business analysts, data scientists, and stakeholders to understand data requirements.
- Perform root cause analysis and resolve data-related production issues.
- Optimize data models, partitioning strategies, and query performance.
- Maintain documentation of data flows, transformations, and technical designs.
- Support batch and near real-time data processing requirements.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
- 6+ years of experience in Data Engineering or ETL Development.
- Strong programming experience in Python.
- Hands-on experience with PySpark and Apache Spark framework.
- Advanced knowledge of SQL and relational databases.
- Experience designing and implementing ETL workflows.
- Robust understanding of Data Warehousing concepts, including Fact and Dimension modeling.
- Experience working with Parquet, ORC, JSON, and CSV data formats.
- Knowledge of version control systems such as Git.
- Strong analytical and problem-solving skills.
Preferred Qualifications
- Experience with Azure Databricks, Azure Data Factory (ADF), ADLS, Synapse, or AWS data services.
- Knowledge of Delta Lake and Lakehouse architecture.
- Experience with workflow orchestration tools such as Airflow.
- Understanding of CI/CD practices and DevOps concepts.
- Familiarity with streaming technologies such as Kafka or Spark Streaming.
Technical Skills
Must Have:
- Python
- PySpark
- SQL
- ETL Development
- Data Warehousing
- Performance Tuning
Good to Have:
- Azure Databricks
- Azure Data Factory (ADF)
- ADLS
- Apache Airflow
- Delta Lake
- Git
- Kafka
Experience
- 6-10 Years of relevant experience in Data Engineering.
Responsibilities in Daily Operations
- Develop and enhance data pipelines.
- Monitor and troubleshoot ETL jobs.
- Ensure data integrity and consistency.
- Participate in code reviews and design discussions.
- Work closely with cross-functional teams to deliver data solutions.
📌 Data Eng (Python+Pyspark+SQL+ETL) (Bengaluru)
🏢 Tata Consultancy Services
📍 Bengaluru