Skip to content

File & Media Transcription ​

Convert long audio recordings, meeting files, interviews, podcasts, and online media streams into structured, synchronized subtitles.


Input Modes ​

1. Local Media Files ​

  • Drag and drop single or multiple audio/video files directly into DuRT.
  • Supported file formats: .mp3, .wav, .m4a, .flac, .aac, .ogg, .mp4, .mov, .mkv, .webm, .avi.
  • Batch processing allows transcribing multiple hours of recordings in the background.
  • Paste links from mainstream video sharing platforms or direct .mp3/.mp4 media URLs.
  • Automatic stream extraction and server-side/client-side audio demuxing.

Advanced Capabilities ​

Word-Level Timestamps ​

  • Outputs millisecond-precise start and end times for every single word.
  • Ideal for karaoke-style subtitle animations, precise video cut points, and word search in video editors.

Speaker Diarization (Multi-Speaker Separation) ​

  • Automatically separates and labels different speakers (Speaker 0, Speaker 1, Speaker 2, etc.).
  • Unsupervised clustering: No pre-recorded voice samples required.
  • Perfect for multi-person interviews, boardroom meetings, and panel podcasts.

Exporting Subtitles ​

Transcribed results can be exported in multiple industry-standard formats:

  • SRT (.srt): Universal format compatible with Premiere Pro, Final Cut Pro, DaVinci Resolve, and VLC.
  • WebVTT (.vtt): Standard format for web HTML5 players.
  • Plain Text (.txt): Clean, formatted text document without timecodes.
  • JSON (.json): Structured data including words, timestamps, and speaker metadata for developers.

Released under the MIT License.