07 Aug
|
TBO.COM
|
Gurugram
Key Responsibilities
- Design, develop, and maintain data pipelines and ETL workflows using Apache Spark and AWS Glue/EMR.
- Design & Develop AI/ML pipelines using pyspark, LLM, Langchain, RAG to solve business use cases.
- Build and optimize data lake solutions leveraging HDFS, Apache Hudi, and AWS S3 for scalable and reliable data storage.
- Develop and manage data ingestion frameworks, transformations, and data models for analytical and machine learning workloads.
- Enable data discovery and querying using AWS Athena and related analytical tools.
- Collaborate with data science teams to prepare, curate, and operationalize datasets for AI/ML model training and deployment.
- Implement and monitor data quality, security, and governance practices across the platform.
- Work closely with DevOps and Cloud (ECS, Lambda) teams to ensure infrastructure is reliable, automated, and cost-efficient.
- Contribute to model deployment, monitoring, and automation pipelines integrating ML components into production workflows.
Required Skills & Experience
- 5–8 years of hands-on experience in Data Engineering with exposure to AI/ML use cases.
- Strong proficiency in Apache Spark (PySpark/Scala) and AWS Glue/EMR for large-scale ETL.
- Deep understanding of HDFS, Data Lake architectures, and distributed data processing.
- Hands-on experience with AWS services such as S3, Athena, Lambda, ECS, IAM, and CloudWatch.
- Strong Python and SQL skills, with experience in data preprocessing, feature engineering, or model integration.
- Familiarity with machine learning concepts (e.g., model lifecycle, data pipelines for ML, basic libraries like scikit-learn or TensorFlow).
- Solid grasp of data modeling, ETL design patterns, and data lifecycle management.
- Knowledge of version control (Git) and CI/CD practices for data and ML pipelines.
- Exposure to end-to-end data science or MLOps pipelines will be a robust plus.
-
📌 Senior Data Engineer (Gurugram)
🏢 TBO.COM
📍 Gurugram