The entire real-speech claim rests on CS-FLEURS JA–EN read/test: 196 utterances, one speaker (SS), read speech, align-then-swap switch points. Three limitations compound:
- One speaker — results are speaker-specific and we cannot separate "generalises" from "happens to suit this voice"
- Read, not spontaneous — no disfluencies, repairs, or natural prosodic switch marking
- Artificially dense switching — span-swapping produces more switches per utterance than natural speech, inflating absolute error rates
04-evaluation/eval-ft/docs/human-eval-protocol.md already specifies a human gold-set protocol that was never executed. This is the single highest-value thing we could add.
What's needed
- Record 30–60 minutes of natural JA–EN bilingual speech from several speakers (meeting-like, spontaneous)
- Transcribe under the script-policy conventions in the existing protocol
- Add as a third frozen benchmark alongside CS-FLEURS and the FLEURS controls
- Report all systems on it
Even 100 utterances from 5 speakers would materially strengthen every claim in the paper, and would let us test whether the CS-FLEURS ranking holds on spontaneous speech.
Related
Would also let us check whether the synthetic-to-real gap measured in §8 is larger on spontaneous speech than on read speech — a natural follow-up result.
The entire real-speech claim rests on CS-FLEURS JA–EN
read/test: 196 utterances, one speaker (SS), read speech, align-then-swap switch points. Three limitations compound:04-evaluation/eval-ft/docs/human-eval-protocol.mdalready specifies a human gold-set protocol that was never executed. This is the single highest-value thing we could add.What's needed
Even 100 utterances from 5 speakers would materially strengthen every claim in the paper, and would let us test whether the CS-FLEURS ranking holds on spontaneous speech.
Related
Would also let us check whether the synthetic-to-real gap measured in §8 is larger on spontaneous speech than on read speech — a natural follow-up result.