- Design, develop, and maintain scalable ETL pipelines using PySpark.
- Build robust data ingestion frameworks for processing data from multiple source systems.
- Develop effective batch and near real-time data processing workflows.
- Perform data extraction, transformation, cleansing, validation, and loading activities.
- Optimize PySpark jobs for performance, memory utilization, and execution time.
- Develop reusable data transformation components and utility functions.
- Write complex SQL queries for data analysis, validation, and reconciliation.
- Process large datasets using distributed computing techniques.
- Collaborate with business teams to understand data requirements and translate them into technical solutions.
- Develop data quality checks and validation frameworks to ensure data accuracy and consistency.
- Troubleshoot ETL failures, production issues, and data discrepancies.
- Perform root cause analysis and implement permanent fixes for recurring issues.
- Integrate data from relational databases, cloud storage, APIs, and flat files.
- Support production deployments and monitor ETL job execution.
- Participate in code reviews and ensure adherence to coding standards and best practices.
- Create technical documentation, data mappings, and process flow diagrams.
- Work closely with DevOps and infrastructure teams to improve deployment automation and operational efficiency.
- Contribute to continuous improvement initiatives for data engineering processes.
Preferred Skills :
- Experience with Azure Databricks or similar Spark-based data platforms.
- Understanding of distributed computing and big data processing.
- Familiarity with cloud platforms such as Azure, AWS, or GCP.
- Experience with data pipeline orchestration tools.
- Knowledge of dimensional data modeling and data warehouse architecture.
- Experience working with large-scale enterprise data platforms.
- Exposure to DevOps practices and automated deployment pipelines.