Reducing false captions during silence, music, and transitions
Speech models can infer generic phrases from ambiguous input such as silence, music, applause, noise, or television transitions. VoxRelay's current build filters common false phrases, probable non-speech, low-confidence segments, and repeated overlap, but no model can eliminate every incorrect caption.
Use the latest VoxRelay version, select the cleanest playback source, keep speech at a clear level, and avoid unnecessary audio processing or background music when practical. Include the app version, audio source, and a description of the surrounding sound when reporting a repeatable false-caption problem. Do not post private recordings or transcripts.