Key Responsibilities
· Design, build, and maintain ETL/ELT pipelines for structured and unstructured
data in AWS.
· Work with AWS services (e.g., S3, Glue, EMR, Lambda, Kinesis, Step Functions,
Redshift, Athena, DynamoDB, RDS) to manage data ingestion, transformation, and
storage.
· Ensure data quality, governance, and security across the full lifecycle.
· Enable real-time and batch data processing to support AI-driven workflows.
· Collaborate with AI/ML teams to prepare datasets for model training, inference,
and fine-tuning.
· Optimize data infrastructure for scalability, cost-efficiency, and performance.
· Implement CI/CD pipelines and best practices for data engineering in a cloudnative environment.
· Monitor and troubleshoot data pipelines to ensure high availability and reliability.
Required Skills & Qualifications
· 3–7 years of experience in data engineering (adjust based on seniority).
· Robust programming skills in Python, SQL, and/or Scala/Java.
· Expertise in AWS cloud ecosystem (S3, Glue, EMR, Redshift, Athena, Kinesis,
Lambda, etc.).
· Experience with data pipeline orchestration tools (Airflow, Step Functions,
Dagster, or similar).
· Proficiency with big data frameworks (Spark, Hadoop, or Flink).
· Familiarity with data modeling, warehousing, and schema design.
· Solid understanding of data governance, lineage, and security (IAM, Lake
Formation, encryption).
· Experience with real-time streaming data (Kafka, Kinesis, or equivalent).
· Knowledge of DevOps practices (Terraform/CloudFormation, CI/CD, Git, Docker)
Pay: ₹509,523.85 - ₹1,868,850.93 per year
Work Location: Remote
📌 Data Engineer (India)
🏢 hirezy
📍 India