# Yuchen fact-depth questionnaire

Purpose: deepen interview evidence without upgrading plausible details into facts. Answer in Chinese or English. Each answer remains **proposed** until Yuchen marks it verified with a traceable note in Practice Studio or reviews the generated proposal manually.

For every answer, record one status: `verified`, `contains_unknown`, `confidential`, or `hypothetical`. Never include supplier names, contract prices, private dataset names, internal endpoints, exact pretraining totals, salary, or personal identifiers.

## 1. Sourcing and acceptance ownership

1. In real chronological order, how did one data source move from discovery to accepted training data?
2. Which decisions were personally yours? Which belonged to procurement, legal/compliance, model researchers, infrastructure engineers, suppliers, or annotators?
3. What was the most common rejection reason at each stage?
4. Which checks were hard gates, which were scores, and which required human review?
5. What threshold or sampling decision can you safely describe? If none is documented, mark it unknown.
6. Describe one real failure that changed the pipeline. Do not substitute a hypothetical incident.

Minimum evidence before promotion: one end-to-end sequence, a responsibility boundary, one real failure or explicit unknown, and one externally safe validation signal.

## 2. Structured captioning

1. What was wrong with the previous tags, using one real example?
2. How were the ten dimensions defined, and which dimensions were hardest to annotate consistently?
3. How were section boundaries represented when sections overlapped or were ambiguous?
4. What did model pre-annotation do, and what exactly did humans verify or correct?
5. What automation or tooling did you personally build or specify?
6. What evaluation supports the captioner comparison, and what limitations make the comparison non-universal?

Minimum evidence: one before/after annotation example, one ambiguity rule, personal scope, and evaluation design without claiming SongEval causality.

## 3. Lead-sheet annotation

1. Which musical elements were represented: melody, chords, key, meter, sections, repeats, or something else?
2. How did the schema handle modulations, inversions, ambiguous harmony, pickups, repeats, and transcription disagreement?
3. Which labeling-tool decision most improved consistency or throughput?
4. How was label quality checked?
5. What training comparison supports the recognition improvement?
6. Why was the SOTA statement subjective, and what objective benchmark was missing?

Minimum evidence: safe schema detail, one ambiguity policy, tool ownership, and the exact boundary of the subjective SOTA claim.

## 4. Evaluation suite, Suno comparison, and SongEval

1. How were more than ten subjective sets divided by capability, language, genre, and use case?
2. Why use three-rater CMOS for routine releases and ten-rater MOS for competitor comparisons?
3. How were order, loudness, rater fatigue, genre preference, and poor raters detected?
4. How did you treat uncertainty and disagreement?
5. What was personally designed or operated by you versus achieved by the model team?
6. State the Suno v5.5 and SongEval claims in externally safe language, including what you are not claiming.

Minimum evidence: release decision flow, rater design, one QA mechanism, and correct public-benchmark attribution.

## 5. Consumer feedback and 50% user growth

1. Which product behaviors or feedback were considered useful model-quality signals?
2. How did you prevent popularity, exposure, novelty, or recommendation bias from masquerading as quality?
3. How did product evidence enter benchmark or model-improvement priorities?
4. What evidence supports “contributed to 50% user growth” without claiming sole causality?
5. Which parts of the growth analysis are unknown or owned by another function?

Minimum evidence: one feedback-to-evaluation loop, attribution boundary, and explicit unknowns around causal measurement.

## 6. Outcome reward-model labeling

1. What exactly did one preference task ask raters to compare?
2. How were conflicting musical dimensions represented in the rubric?
3. How did you distinguish real taste disagreement from a broken task?
4. What does “nearly 90% human agreement” mean operationally, and what does it not prove?
5. How were difficult examples sampled across iterations?
6. Which changes to the labeling system produced a measurable improvement?

Minimum evidence: task unit, rubric/QA loop, precise agreement definition or an explicit unknown, and personal ownership.

## 7. Automatic evaluation and LLM-as-judge

1. Which text-side dimensions achieved more than 80% agreement with human gold?
2. How were train/tuning examples separated from calibration or holdout examples?
3. Which prompt or judge-design decisions did you personally make?
4. How was disagreement inspected by slice rather than only in aggregate?
5. How was pop bias in musicality/aesthetic scorers detected?
6. What conditions kept humans as the release gold standard?

Minimum evidence: dimension definition, gold-set construction, one calibration choice, one bias slice, and a clear statement that aesthetic scorers were evaluated rather than built by Yuchen.

## 8. Data composition and representation decisions

1. For Latin grooves and blues vocals, what audible failure was observed and what alternatives were considered before blaming composition?
2. What did the sampling-ratio or targeted-sourcing intervention actually change?
3. Which evaluation showed recovery, and what regressions were checked?
4. For the representation decision, what did “higher information density” and “lower quality ceiling” mean in measurable terms?
5. What evidence would have changed your mind?
6. Which actor, timeline, cost, and saved-training-run details are undocumented and must remain unknown?

Minimum evidence: observation, competing hypotheses, intervention, validation, falsifiability, and no invented savings or stakeholder story.

## 9. Collaboration, leadership, and engineering boundary

1. Describe the real approximately ten-person operating structure without naming private individuals.
2. How were instructions, calibration, escalation, retraining, and rater removal handled?
3. Give one documented cross-functional decision with model researchers or product; if no exact story exists, mark it unknown.
4. Which Python, SQL, audio-DSP, Docker, CI, or PyTorch tasks can you independently specify, review, debug, and own today?
5. What did agent-assisted development change in your workflow, and how did you verify generated work?
6. Which infrastructure and distributed-training tasks require stronger engineering support?

Minimum evidence: team operating model, one verified collaboration example or explicit unknown, concrete verification practice, and honest implementation limits.

## Promotion rule

A proposed detail may enter candidate knowledge v5 only after all of the following are true:

1. Yuchen explicitly marks it verified.
2. The evidence note identifies a safe source or precise personal recollection.
3. It does not conflict with the submitted resume or knowledge v4.
4. It passes attribution, confidentiality, numerical-precision, causal-language, and hypothetical-versus-past checks.
5. A human reviews the generated proposal. No script or model directly edits `candidate-kb.md`.
