Live Caption & Subtitles
DuRT's Live Caption feature enables ultra-low-latency, real-time speech recognition for live streaming, international video conferences, foreign language classes, and podcasts.
Key Capabilities
- System Audio Capture: Directly transcribes sound from Safari, Chrome, Zoom, Microsoft Teams, YouTube, Spotify, or any application.
- Microphone Input: Real-time voice-to-text with customizable microphone devices.
- Acoustic Echo Cancellation (AEC): Filters out speaker feedback when capturing microphone input alongside system audio.
- Floating Subtitle Window: A customizable translucent window that hovers over other apps.
The Floating Window
The floating window provides maximum visibility with minimal distraction:
- Always on Top: Keeps subtitles visible above full-screen presentations, video players, and code editors.
- Customizable Appearance: Adjust background transparency, font size, font family, and text color.
- Bilingual Display: Display source transcription and translated subtitles simultaneously.
Audio Source Comparison
| Audio Source | Technology | Recommended Use Case |
|---|---|---|
| System Audio | ScreenCaptureKit | Online meetings (Zoom/Teams), YouTube, podcasts, movies |
| Microphone | CoreAudio / AVAudioEngine | Personal speeches, presentations, lectures, interviews |
| Microphone + AEC | Hardware / DSP AEC | Group conference rooms with open speakers |
Engine Options for Live Caption
- Apple Speech Recognition:
- Zero configuration, built into macOS with on-device offline dictation support.
- Real-time word streaming with punctuation.
- OneASR Gateway:
- Connect to local or remote OneASR servers for high-accuracy multilingual recognition.
- OpenAI / Cloud Realtime APIs:
- Cloud-accelerated, highly accurate live transcription.
