02 Aug
|
Diverse Lynx
|
Bengaluru
02 Aug
Diverse Lynx
Bengaluru
Job Title: Data Engineer AWS SageMaker (P1)
Job Summary:
We are looking for an experienced Data Engineer with strong expertise in AWS cloud technologies, PySpark, AWS Glue, Apache Airflow, and Amazon SageMaker to design, develop, and maintain scalable, secure, and high-performance data platforms. The ideal candidate will have experience building end-to-end data pipelines, implementing Medallion Architecture, and supporting analytics and AI/ML workloads in a cloud-native environment.
Key Responsibilities
- Design, develop, and maintain scalable, reliable, and fault-tolerant data pipelines on AWS.
- Build and optimize ETL/ELT workflows using PySpark, AWS Glue, and other AWS services.
- Implement Medallion Architecture (Bronze, Silver, Gold) for structured and semi-structured data processing.
- Develop and orchestrate data workflows using Apache Airflow, including DAG creation, scheduling, dependency management, and monitoring.
- Build data ingestion and transformation pipelines from multiple data sources into enterprise data lakes and data warehouses.
- Work extensively with AWS services including S3, Glue, Athena, Redshift, Lambda, EMR, and related cloud-native technologies.
- Enable AI/ML use cases by building feature engineering pipelines and preparing datasets using Amazon SageMaker.
- Ensure data quality, validation, monitoring, lineage, and observability across the data platform.
- Optimize data storage, partitioning, file formats,
and processing performance for cost-efficient and scalable solutions.
- Implement data security, governance, IAM policies, encryption, and compliance best practices.
- Collaborate with Data Scientists, Data Analysts, Application Developers, and Business teams to deliver high-quality data solutions.
- Participate in architecture discussions, code reviews, and performance optimization initiatives.
- Troubleshoot production issues, monitor pipeline health, and ensure high availability of data platforms.
Required Skills
- Strong experience with Python and PySpark.
- Hands-on experience with AWS Glue, Amazon SageMaker, S3, Athena, Redshift, Lambda, and EMR.
- Experience implementing Medallion Architecture and modern Data Lake/Data Warehouse solutions.
- Strong knowledge of Apache Airflow and DAG development.
- Experience with ETL/ELT design, data modeling, and data integration.
- Positive understanding of SQL, data partitioning, schema evolution, and performance tuning.
- Experience with version control tools such as Git and CI/CD practices.
- Knowledge of data governance, security, IAM, encryption, and cloud best practices.
", Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Data Engineer Sage Maker (Bengaluru)
🏢 Diverse Lynx
📍 Bengaluru