Data Pipeline Engineer (Freelancer) (India)

Data Pipeline Engineer (Freelancer) (India)

06 Aug
|
Deccan AI Experts
|
India

06 Aug

Deccan AI Experts

India

About Us

Deccan AI Experts is a pioneering AI company founded by IIT Bombay and IIM Ahmedabad alumni, with a strong founding team from IITs, NITs, and BITS. We specialize in high-quality human-curated data, AI-first operations, and advanced AI evaluation systems. About the Role

We are seeking a Data Pipeline Engineer (Freelancer) to support advanced AI evaluation initiatives focused on data engineering, ETL/ELT workflows, data pipeline development, cloud-based data lakes, and AI-ready data infrastructure.

In this role, you will evaluate AI-generated data pipeline designs, ETL workflows, schema architectures, data validation logic, and monitoring solutions. Your expertise will help improve AI systems designed for scalable data ingestion, transformation, validation, and analytics.

This position is ideal for professionals with experience in Data Engineering, Big Data, Cloud Data Platforms, ETL Development, Analytics Engineering, or AI infrastructure. Responsibilities

Create deliverables addressing real-world data engineering and pipeline automation scenarios.

Annotate and evaluate AI-generated ETL workflows, data pipeline architectures, schema designs, and monitoring strategies.

Assess AI outputs for scalability, reliability, data quality, and production readiness.

Design and evaluate ETL/ELT pipelines for batch, streaming, and incremental data processing across enterprise systems.

Develop and review schema designs, including relational models, dimensional models, partitioning strategies, metadata management, and schema evolution.

Design and validate Amazon S3/data lake ingestion workflows, ensuring efficient ingestion, storage, partitioning, cataloging, and lifecycle management of structured and unstructured data.

Implement and evaluate data validation checks, including schema validation, completeness, consistency, duplicate detection, referential integrity,



and business rule validation.

Review and optimize pipeline monitoring, including job scheduling, logging, alerting, failure recovery, SLA monitoring, and performance optimization.

Identify data quality issues, transformation errors, schema mismatches, ingestion failures, bottlenecks, and operational risks in AI-generated outputs.

Provide structured feedback to improve AI performance in data engineering, pipeline automation, and cloud data infrastructure.

Review peer-developed deliverables to maintain quality and consistency standards. Requirements

Bachelor's degree in Computer Science, Information Technology, Data Engineering, Software Engineering, or a related field.

3+ years of hands-on experience in Data Engineering, ETL Development, Cloud Data Platforms, Big Data, or Analytics Engineering.

Designing and implementing ETL/ELT pipelines using modern data engineering tools and frameworks.

Developing schema designs for transactional systems, data warehouses, and analytics platforms.

Building Amazon S3/data lake ingestion pipelines using cloud-native architectures and scalable storage solutions.

Performing data validation checks, including quality assurance, reconciliation, anomaly detection, and business rule validation.

Implementing pipeline monitoring with logging, alerting, performance metrics, error handling, retry mechanisms, and operational dashboards.

Optimizing pipeline performance, scalability, fault tolerance, and cost efficiency.





Experience with tools such as Apache Airflow, AWS Glue, Apache Spark, dbt, Kafka, Azure Data Factory, Google Cloud Dataflow, or similar data engineering platforms.

Familiarity with cloud services including Amazon S3, AWS Redshift, Snowflake, BigQuery, Azure Synapse, or Databricks.

Experience with Git, CI/CD pipelines, Docker, and cloud-based deployment workflows is preferred.

Solid analytical thinking, debugging skills, and attention to detail.

Excellent written English communication skills.

Ability to critically evaluate AI-generated pipeline architectures for reliability, scalability, and operational excellence.

Ability to work independently in a remote, fast-paced environment.

Preferred

Qualifications

Experience working with large-scale enterprise data platforms or AI/ML data pipelines.

Experience supporting machine learning feature pipelines, analytics platforms, or real-time streaming architectures.

Familiarity with Delta Lake, Apache Iceberg, Apache Hudi, or lakehouse architectures.

Experience implementing data governance, metadata management, lineage tracking, and data observability.

Knowledge of Kubernetes, Terraform, Infrastructure as Code (IaC), and DevOps practices is an advantage.

Professional certifications in AWS, Azure, Google Cloud, Databricks, Snowflake, or Data Engineering are highly desirable.

Why Join

Us

Competitive hourly pay: ₹1,500 - 2,000/hour

Fully remote with adaptable working hours.

Opportunity to contribute to cutting-edge AI and data engineering initiatives.

Exposure to advanced AI systems focused on scalable data infrastructure, analytics, and intelligent automation.

Flexible project-based opportunities with global teams.

Work on next-generation AI solutions supporting enterprise data platforms and AI applications.

📌 Data Pipeline Engineer (Freelancer) (India)
🏢 Deccan AI Experts
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data pipeline engineer (freelancer) (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: data pipeline engineer (freelancer) (india) / india