04 Sep
|
Wiingy AI
|
Bengaluru
04 Sep
Wiingy AI
Bengaluru
The role
We are looking for an Applied Research Lead, Music AI & Benchmarking to become Wiingy AI's internal authority on generative music models, music datasets and model evaluation. You should understand how modern text-to-music and generative-audio systems are built, trained, fine-tuned, controlled and evaluated.
Your primary mandate will be to build, launch and continuously advance Wiingy AI's global music benchmark. The benchmark should provide a transparent and technically rigorous way to compare music models across genres, cultures, languages, musical tasks and real-world creative use cases.
You will own the benchmark as a continuing research program. This includes its methodology, prompt and evaluation sets, expert and listener panels, technical harness, release cadence, public leaderboard and research roadmap. You will also guide internal teams and translate customer problems into rigorous evaluation and dataset designs.
Primary mandate: build a global music benchmark
- Benchmark strategy: Define what the benchmark will measure, which users it will serve and how it will remain credible, differentiated and useful as music-generation technology changes.
- Scope and task design: Establish the initial scope for text-to-music generation, including instrumental music and songs with vocals. Create a roadmap for continuation, variation, editing, inpainting, stem generation and other creative tasks.
- Global musical coverage: Design coverage across genres, musical traditions, cultures, languages, instruments, ensemble types, vocal styles and production aesthetics without reducing quality to one universal definition of good music.
- Prompt and test-set design: Build representative prompts and evaluation cases covering simple attributes, complex compositions, references to structure and instrumentation, creative constraints and adversarial conditions.
- Model and version policy: Define how models, model versions, generation settings, seeds, clip lengths and system configurations are selected and disclosed so comparisons remain fair and reproducible.
- Evaluation framework: Combine automated measures with blind human evaluation. Define dimensions for prompt adherence, musical coherence, melody, harmony, rhythm, structure, performance, vocals, genre authenticity, audio quality, artifacts and production readiness.
- Expert and listener evaluation: Determine which questions require musicians, composers, producers, audio engineers, genre specialists or general listeners. Design separate expert and listener tracks where their judgments answer different questions.
- Experimental rigor: Set standards for sample size, randomization, counterbalancing, confidence intervals, significance testing, inter-rater reliability and treatment of failed or truncated generations.
- Evaluation harness: Build or supervise the technical system used to generate outputs, preserve model settings, randomize tracks, collect judgments, calculate results and produce reproducible reports.
- Public release: Publish clear methodology, prompt-set documentation, limitations, model configurations,
results and a versioned leaderboard that external researchers and model builders can scrutinize.
- Continuous advancement: Track current models, add genres, languages, tasks and failure categories, refresh evaluation sets and publish regular updates without losing historical comparability.
- Benchmark integrity: Protect the benchmark from contamination, overfitting, selective reporting, undisclosed methodology changes, rights violations and inappropriate commercial influence.
- Research agenda: Use benchmark findings to identify systematic model failures and publish technical reports, failure analyses and original research on music-model evaluation.
- What you will own
- Technical authority: Serve as the go-to person for music-model, music-data and evaluation questions across research, operations, product and GTM teams.
- Model understanding: Maintain deep knowledge of music-generation architectures, audio and symbolic representations, neural audio codecs, conditioning methods, training and fine-tuning approaches, inference controls and common failure modes.
- Dataset architecture: Design collection and annotation specifications for pre-training, fine-tuning, supervised evaluation and preference data, including audio, lyrics, captions, musical attributes and structured metadata.
- Music-data quality: Define standards for audio formats, clip selection, segmentation, loudness, metadata, text-to-music alignment, genre and instrument labels, musical annotations, provenance, QA and acceptance criteria.
- Human evaluation systems: Create rubrics, calibration sets, gold tasks, rater qualification methods, disagreement-resolution rules and inter-rater reliability standards.
- Expert matching: Determine when a task requires a general listener, trained musician, composer, vocalist, producer, mixing engineer or genre specialist, and convert their knowledge into scalable evaluation criteria.
- Failure analysis: Build taxonomies that explain why a generation fails, including musical, performance, vocal, production, technical and prompt-alignment failures.
- Preference and post-training data: Design pairwise comparisons, rankings, written rationales and other expert data that can support preference modelling or post-training.
- Experiments and tooling: Build or supervise lightweight pipelines for model inference, audio processing, dataset inspection, metric computation, error analysis and benchmark reporting.
- Customer solution design: Translate a customer's model objective into a defensible evaluation plan or dataset specification, identify methodological risks and define pilot success criteria.
- Research communication: Produce benchmark documentation, technical playbooks,
research reports and customer-facing methodology documents in clear language.
- Responsible music-data practice: Incorporate licensing, copyright, performer consent, voice and likeness rights, provenance, cultural context, bias and appropriate dataset use into every research design.
What you should bring
- At least 3 years of relevant experience in music AI, audio ML, machine learning, music information retrieval or a closely related field. Exceptional candidates with fewer years and unusually relevant benchmark work will also be considered.
- Hands-on experience training, fine-tuning or deeply evaluating at least one generative music or generative-audio system. You should be comfortable investigating actual model and data failures, not only reading research papers.
- Strong understanding of music-generation pipelines, including audio preprocessing, representations, tokenization or neural codecs, conditioning, generation, decoding and inference-time controls.
- Practical experience designing music or audio datasets, evaluation sets or model benchmarks and making defensible choices about sampling, coverage, annotation, experimental controls and quality assurance.
- Enough musical understanding to reason about melody, harmony, rhythm, structure, instrumentation, performance and production, and to work effectively with specialist musicians and producers.
- Strong knowledge of human-evaluation methodology and the limitations of automated music and audio metrics.
- Strong experimental-design and statistics fundamentals: sampling, bias, variance, confidence intervals, significance testing, agreement measurement and reproducibility.
- Ability to work in Python and with common ML, audio-processing, model-inference and evaluation tooling, such as PyTorch, Hugging Face, librosa or related libraries.
- Ability to explain technical decisions clearly to researchers, operators, product teams, customers and non-technical stakeholders.
Especially valuable
- Experience designing, launching or maintaining a public benchmark, shared task, model leaderboard or reproducible evaluation harness.
- Experience evaluating several commercial or open-source music-generation models under controlled and comparable settings.
- Education or practical experience as a musician, composer, producer, audio engineer or music technologist.
- Experience in music information retrieval, source separation, music transcription, audio-quality assessment, similarity analysis or text-to-music alignment.
- Experience with multilingual music, non-Western musical traditions, culturally specific genres or evaluation across different listener populations.
- Experience with preference data, RLHF or other post-training methods for generative music or audio models.
- Experience working with large annotation operations, external data vendors or distributed expert workforces.
- Published research, open-source contributions or strong technical writing in music AI, audio ML or music information retrieval.
- A master's degree or PhD in a relevant discipline. Equivalent industry research or engineering experience is equally valuable.
📌 Applied Research Lead — Music AI & Benchmarking (Bengaluru)
🏢 Wiingy AI
📍 Bengaluru