We are looking for a skilled AWS Data Engineer with strong expertise in PySpark, Python, and AWS cloud services to design, develop, and optimize scalable data processing solutions. The ideal candidate should have hands-on experience in Big Data technologies and cloud-based data platforms.
Key Responsibilities
- Develop and maintain scalable ETL/data pipelines using PySpark and Python
- Build and optimize Big Data processing solutions on AWS
- Design and implement data ingestion, transformation, and validation workflows
- Work with large-scale datasets and distributed computing frameworks
- Monitor, troubleshoot, and optimize Spark jobs for performance
- Collaborate with business and technical stakeholders to deliver data-driven solutions
- Participate in Agile development processes and project planning
Required Skills
- Strong experience in PySpark, Apache Spark, and Python
- Hands-on experience with AWS services:
- EMR
- S3
- IAM
- Lambda
- SNS
- SQS
- Good understanding of Big Data concepts and distributed processing
- Robust SQL skills including Joins, Subqueries, and CTEs
- Experience with database technologies and data modeling concepts
- Knowledge of Spark performance tuning, debugging, and optimization techniques
Preferred Skills
- Experience with Data Vault, Data Mesh, or Data Fabric architectures
- Exposure to cloud-native and event-driven architectures
- Experience working in Agile environments
Qualifications
- Bachelor's degree in Computer Science, Engineering, IT, or related field
- 4+ years of hands-on development experience in Data Engineering
- Solid coding skills and ability to contribute from Day 1