18 Sep
|
Weekday AI
|
Mumbai
???? ???? ?? ??? ??? ?? ??? ???????'? ???????
?????? ?????: ?? ??????? - ?? ??????? (?? ??? ??-?? ???)
Experience: 3+ yrs
Location: Mumbai, Maharashtra, India
Job Type: Full-time
We are looking for an experienced Data Engineer to design, develop, and maintain scalable data platforms and pipelines using AWS, Apache Spark, Python, SQL, Kafka, and up-to-date data engineering frameworks.
The role will focus on building reliable data solutions for high-volume batch and real-time workloads, including data ingestion, transformation, processing, orchestration, storage, and delivery. The ideal candidate will have strong hands-on experience with AWS data services and distributed data processing, along with a solid understanding of data modelling, streaming architectures, data quality, and pipeline optimisation.
KEY RESPONSIBILITIES
- Design, develop, and maintain end-to-end data pipelines for high-volume data ingestion, transformation, processing, and delivery.
- Build scalable Spark-based ETL/ELT workflows for both batch and real-time data processing.
- Develop data ingestion solutions using Kafka, Amazon Kinesis, and other streaming technologies.
- Build and manage data lakes, warehouses, and lakehouse solutions using AWS S3, Glue, Redshift, Athena, and EMR.
- Develop efficient data models using dimensional modelling, star schemas, partitioning, and other data engineering practices.
- Implement data quality checks, validation rules, monitoring, and error-handling mechanisms.
- Develop automated workflows using Airflow, MWAA, AWS Step Functions, or similar orchestration tools.
- Collaborate with Data Analysts and Data Scientists to deliver clean, structured, and analytics-ready datasets.
- Optimise data pipelines for performance, scalability, reliability, and AWS cost efficiency.
- Integrate data from multiple internal and external systems while maintaining data consistency and reliability.
- Develop Python-based automation and data-processing solutions.
- Monitor production pipelines, troubleshoot failures,
and perform root-cause analysis.
- Follow modern software engineering practices including Git, CI/CD, testing, documentation, and code reviews.
- Contribute to data platform architecture, engineering standards, and continuous improvement initiatives.
- Support data governance, cataloguing, lineage, and metadata management practices where required.
- Work with Linux/Unix environments and efficiently process large datasets.
WHAT MAKES YOU A GREAT FIT
- 3+ years of professional experience in Data Engineering, Big Data, Analytics Engineering, or a related field.
- Strong hands-on programming experience with Python for data processing, automation, and pipeline development.
- Strong expertise in Apache Spark, particularly PySpark and/or Spark SQL.
- Deep working knowledge of the AWS data ecosystem, including S3, Glue, Redshift, Athena, EMR, Kinesis, Lambda, and IAM.
- Hands-on experience with Kafka, Kinesis, Flink, or similar real-time streaming technologies.
- Strong command of SQL and experience with data modelling, dimensional modelling, partitioning, and large-scale data processing.
- Experience with Airflow, MWAA, Step Functions, or comparable workflow orchestration tools.
- Strong understanding of batch and real-time data processing architectures.
- Experience working with high-volume datasets and distributed data processing environments.
- Familiarity with Git, CI/CD, testing, and up-to-date software development practices.
- Comfortable working in Linux/Unix environments.
- Strong troubleshooting, analytical, and problem-solving skills.
- Experience with Delta Lake, Apache Iceberg, Hudi, or other lakehouse technologies is an advantage.
- Knowledge of data governance, cataloguing, metadata, and lineage tools such as Glue Data Catalog, DataHub, or Amundsen is a plus.
- Familiarity with Docker, ECS, or EKS and containerised deployments is desirable.
- Basic understanding of AI/ML data requirements and workflows is an advantage.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical disciplineis preferred.
📌 Data Engineer (Mumbai)
🏢 Weekday AI
📍 Mumbai