- Design, develop, and maintain scalable data pipelines using AWS services and PySpark.
- Create and maintain optimal data pipeline architecture for efficient data processing.
- Build data ingestion, transformation,
and ETL workflows.
- Develop batch and near real-time data processing solutions.
- Work with large and complex datasets to meet business requirements.
- Manage and optimize data warehouses using AWS Redshift or Hive.
- Ensure data quality, integrity, security, and governance standards.
- Perform SQL tuning and performance optimization.
- Automate data workflows and improve platform scalability.
- Collaborate with Product, Data, Engineering, and Business teams for solution delivery.
- Create technical documentation and support operational readiness.