Before you begin

1. Requirements

  • A 64-bit computer running Windows 10 or Windows 11.
  • A supported NVIDIA CUDA GPU with a current NVIDIA production driver is optional. Without one, VoxRelay uses CPU processing.
  • CPU mode uses a smaller model and has higher latency and lower translation accuracy than CUDA mode.
  • An enabled Windows playback device, such as speakers, headphones, HDMI audio, or compatible capture hardware.
  • Internet access during the first launch unless the selected speech model is already in the current Windows user's local cache.
  • Several gigabytes of free storage for the speech model, application runtime, and temporary download data.

VoxRelay captures computer playback through Windows WASAPI loopback. It does not use the microphone as its normal audio source.

Installation

2. Install and complete the first run

  1. Install VoxRelay from its Microsoft Store listing.
  2. Open VoxRelay from the Windows Start menu.
  3. Look for the VoxRelay icon in the notification area. Windows may place it under the hidden-icons arrow.
  4. Keep the computer connected to the internet while the default speech model downloads on first use. A model-setup window displays the percentage and transferred size; the download can take several minutes.
  5. When initialization is complete, the floating overlay displays Listening.... Play speech through the selected Windows output.

The selected model is downloaded from Max Power IT's versioned HTTPS model host and cached in the current Windows user profile. CUDA uses the larger Large-v3 model with float16 compute; CPU uses the smaller Small model with int8 compute. VoxRelay verifies every file against a pinned size and SHA-256 hash and can resume an interrupted download. If the primary host is unavailable, VoxRelay automatically uses the original Hugging Face repository as a backup. The progress window identifies the active source and closes after the model is loaded. Model download requests do not contain captured audio or subtitle text.

Playback capture

3. Select the audio source

  1. Right-click the VoxRelay notification-area icon.
  2. Open Select Source.
  3. Choose the Windows playback or loopback source carrying the audio you want to caption.
  4. If a recently connected device is missing, choose Refresh Sources and open the source menu again.
  5. Start playback and allow at least one complete speech segment for the first subtitle.

Changing sources restarts capture. Names come from Windows and the installed audio driver, so the available choices vary by computer.

Select the processing device

  1. Right-click the VoxRelay notification-area icon and open Processing Device.
  2. Choose Auto to prefer NVIDIA CUDA and fall back to CPU automatically.
  3. Choose NVIDIA CUDA to require GPU processing, or CPU to use the Small/int8 model without an NVIDIA GPU.

VoxRelay saves this preference and reloads the appropriate model without restarting. The first switch to a model may require a download.

Floating overlay

4. Position, resize, show, and hide subtitles

  • Move: Drag within the floating subtitle area.
  • Resize: Move the pointer near an invisible edge or corner until the resize pointer appears, then drag.
  • Hide: Press Esc or Q while the overlay has focus, or right-click the tray icon and choose Hide Subtitles.
  • Show: Right-click the tray icon and choose Show Subtitles.
  • Quit: Right-click the tray icon and choose Quit. Hiding the overlay does not stop speech processing.

The subtitle surface is transparent and does not create a normal taskbar button. Only floating white text and its dark outline are visible.

Automatic mode

5. Understand transcription and translation

VoxRelay's standard mode performs language analysis on confident speech and rechecks the spoken language every 30 seconds.

  • English speech is transcribed directly in English.
  • Supported non-English speech is translated into English.
  • Translated captions are prefixed with the detected source language and confidence, such as [Translated from Spanish 96%].
  • A potential language change must be detected confidently more than once before the remembered language changes.
  • Uncertain speech, non-speech audio, repeated overlap, and low-confidence segments may be omitted.

A quiet period does not force a language change. Clear speech after a speaker or language change gives the detector the best chance to update correctly.

TV and home theater audio

6. Use HDMI ARC or eARC audio

A television's ARC or eARC port is not automatically a Windows recording source. VoxRelay can use that audio only when compatible capture hardware and its driver expose the signal to Windows as a selectable playback or loopback source.

  1. Connect the TV or receiver to an ARC/eARC-compatible audio capture device or interface designed to make the return-audio channel available to the PC.
  2. Install the manufacturer's current Windows driver.
  3. Confirm the device appears and carries audio in Windows sound settings.
  4. In VoxRelay, choose Refresh Sources, then select that source.

HDCP-protected sources and hardware that exposes ARC/eARC only as an output may not provide capturable audio. VoxRelay cannot bypass content protection or create a Windows source that the hardware driver does not expose.

Data handling

7. Local processing and model downloads

  • Audio capture, language detection, transcription, and translation run on the local computer.
  • Captured audio is processed in memory and discarded; VoxRelay does not send it to Max Power IT or Hugging Face.
  • VoxRelay has no required account, advertising, analytics, tracking, or telemetry SDK.
  • Max Power IT's model host is contacted only when the selected CUDA or CPU model is missing. Hugging Face is used as a backup or for a separately selected model.
  • Optional SRT output is not enabled by the standard Store launch. When explicitly enabled in another distribution, subtitle text is written to the user-selected file.

Read the complete VoxRelay Privacy Policy.

Problem solving

8. Troubleshooting

No notification-area icon appears

Open the Windows hidden-icons panel. If VoxRelay is not present, start it again from the Start menu. Use Task Manager to end an unresponsive VoxRelay process before relaunching.

The overlay stays on Listening...

  1. Confirm that speech is audible through the selected Windows playback device.
  2. Right-click the tray icon, refresh sources, and select the active output.
  3. Check Windows Volume Mixer to confirm the source application is not muted or routed to another output.
  4. Try continuous, clearly spoken audio. Very quiet input and non-speech content are intentionally ignored.

The model does not load

Keep the computer online during the first model download and confirm that security software permits HTTPS access to maxpowerit.com and huggingface.co. VoxRelay tries maxpowerit.com first and reports when it changes to the Hugging Face backup. An unauthenticated-request warning affects backup-download rate limits; it does not mean audio is uploaded.

A CUDA error appears

Install the latest compatible NVIDIA driver, restart Windows, and confirm the GPU is supported. You can also choose CPU from the notification-area Processing Device menu to run without CUDA.

Captions show generic phrases that were not spoken

Speech models can infer generic phrases from silence, music, applause, or television transitions. The current build filters common phrases, low-confidence segments, and non-speech audio, but no model can eliminate every false caption. Confirm the latest version is installed, use a clean playback source, and reduce background music when practical.

The wrong language remains selected

Play several seconds of clear speech in the new language. VoxRelay rechecks every 30 seconds and requires repeated confident detections before changing languages, which avoids rapid switching during short clips or uncertain audio.

Subtitles lag behind the audio

Latency includes audio collection and model inference. CPU mode deliberately gathers longer speech context and is slower than CUDA. For lower latency, select NVIDIA CUDA when supported, close other GPU-intensive applications, use the latest driver, and select a source without additional Bluetooth, network, or home-theater buffering.

Maintenance

9. Update or uninstall

Update

Microsoft Store builds update through the normal Store update process. Keep Windows and the NVIDIA driver current as well.

Uninstall

  1. Quit VoxRelay from its notification-area menu.
  2. Open Settings > Apps > Installed apps.
  3. Select VoxRelay and choose Uninstall.

Downloaded speech-model files may remain under the current user's VoxRelay model directory or Hugging Face cache. The user may remove those files separately when they are no longer needed.

Still need help?

Check the FAQ or open VoxRelay support.