Skip to content

Live Caption & Subtitles ​

DuRT's Live Caption feature enables ultra-low-latency, real-time speech recognition for live streaming, international video conferences, foreign language classes, and podcasts.


Key Capabilities ​

  • System Audio Capture: Directly transcribes sound from Safari, Chrome, Zoom, Microsoft Teams, YouTube, Spotify, or any application.
  • Microphone Input: Real-time voice-to-text with customizable microphone devices.
  • Acoustic Echo Cancellation (AEC): Filters out speaker feedback when capturing microphone input alongside system audio.
  • Floating Subtitle Window: A customizable translucent window that hovers over other apps.

The Floating Window ​

The floating window provides maximum visibility with minimal distraction:

  • Always on Top: Keeps subtitles visible above full-screen presentations, video players, and code editors.
  • Customizable Appearance: Adjust background transparency, font size, font family, and text color.
  • Bilingual Display: Display source transcription and translated subtitles simultaneously.

Audio Source Comparison ​

Audio SourceTechnologyRecommended Use Case
System AudioScreenCaptureKitOnline meetings (Zoom/Teams), YouTube, podcasts, movies
MicrophoneCoreAudio / AVAudioEnginePersonal speeches, presentations, lectures, interviews
Microphone + AECHardware / DSP AECGroup conference rooms with open speakers

Engine Options for Live Caption ​

  1. Apple Speech Recognition:
    • Zero configuration, built into macOS with on-device offline dictation support.
    • Real-time word streaming with punctuation.
  2. OneASR Gateway:
    • Connect to local or remote OneASR servers for high-accuracy multilingual recognition.
  3. OpenAI / Cloud Realtime APIs:
    • Cloud-accelerated, highly accurate live transcription.

Released under the MIT License.