Skip to content

Build a speaker-diverse real-speech evaluation set #6

Description

@Awshesh12

The entire real-speech claim rests on CS-FLEURS JA–EN read/test: 196 utterances, one speaker (SS), read speech, align-then-swap switch points. Three limitations compound:

  • One speaker — results are speaker-specific and we cannot separate "generalises" from "happens to suit this voice"
  • Read, not spontaneous — no disfluencies, repairs, or natural prosodic switch marking
  • Artificially dense switching — span-swapping produces more switches per utterance than natural speech, inflating absolute error rates

04-evaluation/eval-ft/docs/human-eval-protocol.md already specifies a human gold-set protocol that was never executed. This is the single highest-value thing we could add.

What's needed

  • Record 30–60 minutes of natural JA–EN bilingual speech from several speakers (meeting-like, spontaneous)
  • Transcribe under the script-policy conventions in the existing protocol
  • Add as a third frozen benchmark alongside CS-FLEURS and the FLEURS controls
  • Report all systems on it

Even 100 utterances from 5 speakers would materially strengthen every claim in the paper, and would let us test whether the CS-FLEURS ranking holds on spontaneous speech.

Related

Would also let us check whether the synthetic-to-real gap measured in §8 is larger on spontaneous speech than on read speech — a natural follow-up result.

Metadata

Metadata

Assignees

Labels

evaluationEvaluation harness, metrics, benchmarksresearchExploratory direction, new scope

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions