Data Engineer - ML Training Data Pipeline (Hyderabad)

Data Engineer - ML Training Data Pipeline (Hyderabad)

11 Sep
|
Data Economy
|
Hyderabad

11 Sep

Data Economy

Hyderabad

Job Summary

Job Title: Data Engineer - ML Training Data Pipeline Notice period: 0-30 Days

Experience: 5+ Years

Location: Hyderabad OR Pune

We are looking for Data Engineer - ML Training Data Pipeline who can build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS.

What We Expect

- Build end-to-end data pipelines: raw trace ingestion dedup format conversion quality gating training-ready datasets
- Process large-scale JSONL data on AWS S3 (tens of thousands of traces per batch)
- Convert between chat-completion formats (e.g., OpenAI Llama 3.1 tool-calling format)
- Implement smart deduplication and sampling to balance training distribution
- Design identity-aware train/test splits that measure true generalization
- Build data validation gates to detect schema drift and format anomalies
- Create a continuous pipeline that auto-processes new production traces for retraining

Requirements

- Experience: 6+ years data engineering focused on ML data pipelines
- Python: Strong pandas, pyarrow, JSONL processing at scale
- ML Data Libraries: HuggingFace Datasets, Arrow-based storage
- Data Formats: Multi-turn conversation/chat data structures and tokenizer-specific formatting




- Deduplication: Content hashing, identity-based grouping strategies
- AWS: S3, EC2, batch processing workflows

Preferred

- Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright.

Benefits

- Comprehensive Medical Coverage: Health insurance of Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
- Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
- Retirement Benefits: PF and Gratuity provided as per standard government regulations.
- Flexible Work Options: Enjoy hybrid work arrangements & adaptable working hours.
- Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays.
- Employee Well-being Spaces: Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Data Engineer - ML Training Data Pipeline (Hyderabad)
🏢 Data Economy
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data engineer - ml training data pipeline (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: data engineer - ml training data pipeline (hyderabad) / hyderabad