Workflow Steps
DuRT's Workflow Engine lets you chain up to 6 types of steps into a single pipeline — from audio capture or file transcription, through AI processing, all the way to export. Each step is self-contained and configurable; you only include the steps your task requires.
NOTE
Steps are executed in sequence. The output of each step is automatically passed as input to the next one.
Step 1 — RealtimeASR
Purpose: Capture and transcribe live audio in real time, either from your microphone or from system audio (any app playing on your Mac).
How It Works
DuRT uses ScreenCaptureKit to tap into system audio, allowing you to transcribe what you're hearing — a video call, a streaming lecture, a podcast — without any cables or virtual audio routing. Microphone input is also supported for in-person meetings or dictation.
Key Options
| Option | Description |
|---|---|
| Input Source | System Audio (ScreenCaptureKit) or Microphone |
| AEC (Acoustic Echo Cancellation) | Removes microphone feedback when system audio is active; prevents your speakers from being re-transcribed |
| Language | Select the spoken language for best accuracy |
| Engine | Choose your preferred ASR engine (Whisper, etc.) |
Best Practices
TIP
Enable AEC whenever you are using System Audio input and have speakers active. Without it, the ASR engine may pick up its own output and produce duplicate or garbled text.
- For online meetings, use System Audio so all speakers are captured, not just you.
- For dictation or voice notes, use Microphone for the cleanest signal.
- Keep the target language set correctly — a mismatch causes significantly lower accuracy.
Step 2 — MediaTranscription
Purpose: Transcribe pre-recorded audio or video from a local file or a remote URL.
Supported Sources
Local files: MP3, WAV, M4A, FLAC, OGG, MP4, MKV, MOV, AVI, WEBM, and more.
Remote URLs:
- YouTube videos and playlists
- Other video platform links
- Direct audio/video URLs (
.mp3,.mp4,.m3u8, etc.)
Key Options
| Option | Description |
|---|---|
| Source | Local file path or remote URL |
| Word-level Timestamps | Attach a precise timestamp to every word, enabling accurate subtitle splitting |
| Speaker Diarization | Identify and label different speakers (Speaker 1, Speaker 2 …) |
| Language | Set the spoken language, or use auto-detect |
| Engine | Select the ASR engine to use |
Best Practices
TIP
Enable Word-level Timestamps if you plan to export subtitles (SRT/VTT). They make subtitle line-breaking far more precise.
TIP
Use Speaker Diarization for interviews, podcasts, or multi-person meetings so you can tell who said what in the transcript.
- For YouTube links, DuRT downloads and transcribes the audio automatically — no manual download needed.
- If a video has embedded subtitles or an auto-generated transcript available, DuRT can use them as a faster alternative to full ASR transcription.
Step 3 — SubtitleInput
Purpose: Bring existing text or subtitles into the workflow instead of transcribing audio from scratch.
Input Methods
| Method | Description |
|---|---|
| Import SRT file | Load a standard .srt subtitle file from disk |
| Reference a past task | Pull the transcript output from a previous DuRT task |
| Type / paste text | Enter or paste text directly in the editor |
When to Use This Step
Use SubtitleInput when you already have a transcript and only need downstream processing — for example, translating an existing SRT file or running AI proofreading on a transcript you received from someone else.
NOTE
SubtitleInput is typically used as the first step in a workflow, replacing RealtimeASR or MediaTranscription when no new audio capture is needed.
Best Practices
- When referencing a past task, make sure the task has finished and its transcript is available before starting the new workflow.
- Plain text input (no timestamps) works fine for LLMProcess steps, but won't produce timed subtitles on export without timestamps.
Step 4 — LLMProcess
Purpose: Apply AI language model processing to the transcript — proofread, translate, split into sentences, or generate a summary.
Available Modes
| Mode | What It Does |
|---|---|
| Proofreading | Fixes punctuation, grammar, and ASR errors while preserving meaning |
| Translation | Translates the transcript into a target language of your choice |
| Sentence Splitting | Breaks a long block of text into natural, readable sentences with proper line lengths for subtitles |
| Summary | Condenses the full transcript into a concise summary (great for long meetings or lectures) |
Key Options
| Option | Description |
|---|---|
| Mode | One of the four modes above |
| Target Language | For Translation mode — the language to translate into |
| LLM Provider / Model | Choose which AI model to use (OpenAI GPT-4o, etc.) |
| Custom Prompt | Add your own instructions to guide the LLM's behavior |
Best Practices
TIP
Chain LLMProcess steps for complex tasks. For example: first Proofreading, then Translation. Each step refines the output of the previous one.
IMPORTANT
LLMProcess requires a valid API key configured in DuRT's settings for your chosen LLM provider.
- Use Proofreading before Translation for the cleanest result — correcting errors first prevents them from propagating into the translation.
- Sentence Splitting is especially useful after transcription of spoken content, which often lacks sentence boundaries.
- For meeting recordings, combine Proofreading + Summary to get a clean transcript and a quick brief in one workflow.
Step 5 — TextToSpeech
Purpose: Convert the processed transcript into natural-sounding speech audio.
Voice Options
| Option | Description |
|---|---|
| Preset Voices | Choose from a library of built-in voices across multiple languages |
| Voice Cloning | Upload a short audio sample to clone a specific voice |
| Voice Design | Adjust parameters (pitch, style, accent) to create a custom voice |
Key Options
| Option | Description |
|---|---|
| Voice | Select preset, cloned, or designed voice |
| Speed | Adjust speaking rate (e.g., 0.75× for slow, 1.25× for fast) |
| Language / Locale | Ensure the TTS engine matches the text language |
| TTS Provider | Choose the provider (OpenAI TTS, system voices, etc.) |
Best Practices
TIP
When using Voice Cloning, provide a clean recording of 10–30 seconds with minimal background noise for the best clone quality.
TIP
Pair TextToSpeech with a Translation step before it to produce dubbed audio in a different language.
- Match the TTS voice language to the text language — a mismatch produces heavily accented or unintelligible output.
- Use speed adjustment for accessibility: slower speech for language learners, faster for podcasts or narration.
Step 6 — ExportArtifact
Purpose: Save the workflow output to one or more files in your chosen format.
Subtitle Formats
| Format | Extension | Best For |
|---|---|---|
| SRT | .srt | Universal subtitles, video players, most platforms |
| VTT | .vtt | Web video (HTML5), YouTube, streaming |
| TXT | .txt | Plain text transcript, no timestamps |
Audio Formats
| Format | Extension | Best For |
|---|---|---|
| AAC | .aac | High quality, small file size |
| MP3 | .mp3 | Universal compatibility |
Key Options
| Option | Description |
|---|---|
| Format | Select one or more output formats |
| Output Path | Choose the save location on disk |
| Filename | Customize the output filename |
| Split by Speaker | Export separate files per speaker (requires diarization) |
Best Practices
TIP
You can select multiple export formats in one step — for example, export both SRT and TXT simultaneously.
NOTE
Audio export (AAC/MP3) is only available when the workflow includes a TextToSpeech step upstream.
- For video editing workflows, export SRT and drop it directly into your video editor.
- For archiving meeting notes, TXT is the simplest and most portable format.
- If your workflow includes Speaker Diarization, consider enabling Split by Speaker to produce per-participant transcripts.
