Job Title:Data Engineer - ML Training Data Pipeline
Notice period: 0-30 Days
Experience : 5+ Years
Location: Hyderabad OR Pune
We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS.
What We Expect:
- Build end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets
- Process large-scale JSONL data on AWS S3 (tens of thousands of traces per batch)
Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright.
Benefits
- Comprehensive Medical Coverage:
Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
- Robust Protection Plans:
Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
- Retirement Benefits:
PF and Gratuity provided as per standard government regulations.
- Flexible Work Options:
Enjoy hybrid work arrangements & versatile working hours.
- Generous Leave Policy:
21 days of annual leave, in addition to 10 company-declared holidays.
- Employee Well-being Spaces:
Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.
📌 Data Engineer — ML Training Data Pipeline (Hyderabad)
🏢 DataEconomy
📍 Hyderabad
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.