- Design and implement robust ETL/ELT pipelines using PySpark for large-scale data processing in distributed environments.
- Utilize AWS services such as AWS Glue, EMR, Redshift, and S3 for data storage, transformation, and warehousing.
- Optimize and fine-tune PySpark jobs and SQL queries to enhance data processing speed and reduce costs.
- Design data models and implement quality checks to ensure data consistency and accuracy.
- Work with data architects, data scientists, and analysts to deliver data solutions, participating in code reviews and adhering to best practices.
- Implement streaming solutions using Kafka and automate workflows, often using tools like Airflow.
Required Skills
- Robust proficiency in Python and SQL (window functions, query optimization).
- Expertise in PySpark and Spark architecture for distributed data processing.
- Hands-on experience with AWS cloud environment, specifically Glue, EMR, Redshift, and S3.
- Understanding of Modern Data Architectures, including Snowflake or Databricks.
- 6+ years of relevant experience in data engineering roles.
Preferred Skills
- Knowledge of NoSQL databases (e.g., DynamoDB).
- Familiarity with DevOps tools (CI/CD) and version control (Git).
- Streaming experience (Kafka).
- Relevant AWS or Data Engineering certifications.