Description
Qualifications
Ph.D., Master’s degree, or equivalent practical experience in Computer Science, Artificial Intelligence, Machine Learning, Operations Research, Statistics, or a related technical field.
5+ years of relevant experience with a Master’s degree, or 3+ years with a Ph.D., applying machine learning to real-world problems.
Strong Python programming skills and experience building production-quality ML, GenAI, or data systems.
Hands-on experience with PyTorch and modern deep-learning stacks; experience with Hugging Face, LLMs, VLMs, diffusion models, or multimodal models is strongly preferred.
Experience with data-centric AI or GenAI methods, such as synthetic-data generation, data-quality measurement, dataset curation, weak supervision, model-based labeling, active learning, deduplication, or data augmentation.
Experience designing experiments and interpreting results through statistical analysis, ablation studies, benchmark evaluation, and error analysis.
Robust understanding of model training, inference, evaluation, and production monitoring.
Ability to evaluate research papers, identify practical value, and implement practical techniques in real-world systems.
Experience building scalable data or ML pipelines using distributed compute, cloud storage, batch processing, or workflow orchestration.
Solid written and verbal communication skills, including experience preparing technical proposals, design documents, experiment reports, and stakeholder presentations.
Description
Design and build data-centric Generative AI methods for synthetic data generation, multimodal data curation, augmentation, filtering, deduplication, and data-quality assessment.
Develop and evaluate synthetic-data pipelines for text, speech, vision, and multimodal GenAI use cases, including controllable generation, provenance tracking, safety checks, and domain adaptation.
Build evaluation frameworks that connect data quality with downstream model perfo
📌 Senior Machine Learning Engineer Bengaluru (India)
🏢 Oracle
📍 India