At ~2 dB, noise is a strong stress test. Stability is not correctness: all samples can share the same wrong answer. Four noisy samples give only six pairwise comparisons. Phonetic scores exclude the original and are withheld on incomplete runs. Timestamps are model-predicted segment boundaries, not exact uncertain-word boundaries. Speaker IDs are local to each rollout and can change between variants. Mandarin uses dictionary pinyin (tones ignored by default); English uses eSpeak IPA. Disagreement scores ignore punctuation; probability curves retain it. Phonetic scores use speech text only; curves include generated timestamps and speaker labels. Audio and results are temporary server files, not published as a dataset; download runs you want to keep.