Skip to content

Workflow Steps ​

DuRT's Workflow Engine lets you chain up to 6 types of steps into a single pipeline — from audio capture or file transcription, through AI processing, all the way to export. Each step is self-contained and configurable; you only include the steps your task requires.

NOTE

Steps are executed in sequence. The output of each step is automatically passed as input to the next one.


Step 1 — RealtimeASR ​

Purpose: Capture and transcribe live audio in real time, either from your microphone or from system audio (any app playing on your Mac).

How It Works ​

DuRT uses ScreenCaptureKit to tap into system audio, allowing you to transcribe what you're hearing — a video call, a streaming lecture, a podcast — without any cables or virtual audio routing. Microphone input is also supported for in-person meetings or dictation.

Key Options ​

OptionDescription
Input SourceSystem Audio (ScreenCaptureKit) or Microphone
AEC (Acoustic Echo Cancellation)Removes microphone feedback when system audio is active; prevents your speakers from being re-transcribed
LanguageSelect the spoken language for best accuracy
EngineChoose your preferred ASR engine (Whisper, etc.)

Best Practices ​

TIP

Enable AEC whenever you are using System Audio input and have speakers active. Without it, the ASR engine may pick up its own output and produce duplicate or garbled text.

  • For online meetings, use System Audio so all speakers are captured, not just you.
  • For dictation or voice notes, use Microphone for the cleanest signal.
  • Keep the target language set correctly — a mismatch causes significantly lower accuracy.

Step 2 — MediaTranscription ​

Purpose: Transcribe pre-recorded audio or video from a local file or a remote URL.

Supported Sources ​

Local files: MP3, WAV, M4A, FLAC, OGG, MP4, MKV, MOV, AVI, WEBM, and more.

Remote URLs:

  • YouTube videos and playlists
  • Other video platform links
  • Direct audio/video URLs (.mp3, .mp4, .m3u8, etc.)

Key Options ​

OptionDescription
SourceLocal file path or remote URL
Word-level TimestampsAttach a precise timestamp to every word, enabling accurate subtitle splitting
Speaker DiarizationIdentify and label different speakers (Speaker 1, Speaker 2 …)
LanguageSet the spoken language, or use auto-detect
EngineSelect the ASR engine to use

Best Practices ​

TIP

Enable Word-level Timestamps if you plan to export subtitles (SRT/VTT). They make subtitle line-breaking far more precise.

TIP

Use Speaker Diarization for interviews, podcasts, or multi-person meetings so you can tell who said what in the transcript.

  • For YouTube links, DuRT downloads and transcribes the audio automatically — no manual download needed.
  • If a video has embedded subtitles or an auto-generated transcript available, DuRT can use them as a faster alternative to full ASR transcription.

Step 3 — SubtitleInput ​

Purpose: Bring existing text or subtitles into the workflow instead of transcribing audio from scratch.

Input Methods ​

MethodDescription
Import SRT fileLoad a standard .srt subtitle file from disk
Reference a past taskPull the transcript output from a previous DuRT task
Type / paste textEnter or paste text directly in the editor

When to Use This Step ​

Use SubtitleInput when you already have a transcript and only need downstream processing — for example, translating an existing SRT file or running AI proofreading on a transcript you received from someone else.

NOTE

SubtitleInput is typically used as the first step in a workflow, replacing RealtimeASR or MediaTranscription when no new audio capture is needed.

Best Practices ​

  • When referencing a past task, make sure the task has finished and its transcript is available before starting the new workflow.
  • Plain text input (no timestamps) works fine for LLMProcess steps, but won't produce timed subtitles on export without timestamps.

Step 4 — LLMProcess ​

Purpose: Apply AI language model processing to the transcript — proofread, translate, split into sentences, or generate a summary.

Available Modes ​

ModeWhat It Does
ProofreadingFixes punctuation, grammar, and ASR errors while preserving meaning
TranslationTranslates the transcript into a target language of your choice
Sentence SplittingBreaks a long block of text into natural, readable sentences with proper line lengths for subtitles
SummaryCondenses the full transcript into a concise summary (great for long meetings or lectures)

Key Options ​

OptionDescription
ModeOne of the four modes above
Target LanguageFor Translation mode — the language to translate into
LLM Provider / ModelChoose which AI model to use (OpenAI GPT-4o, etc.)
Custom PromptAdd your own instructions to guide the LLM's behavior

Best Practices ​

TIP

Chain LLMProcess steps for complex tasks. For example: first Proofreading, then Translation. Each step refines the output of the previous one.

IMPORTANT

LLMProcess requires a valid API key configured in DuRT's settings for your chosen LLM provider.

  • Use Proofreading before Translation for the cleanest result — correcting errors first prevents them from propagating into the translation.
  • Sentence Splitting is especially useful after transcription of spoken content, which often lacks sentence boundaries.
  • For meeting recordings, combine Proofreading + Summary to get a clean transcript and a quick brief in one workflow.

Step 5 — TextToSpeech ​

Purpose: Convert the processed transcript into natural-sounding speech audio.

Voice Options ​

OptionDescription
Preset VoicesChoose from a library of built-in voices across multiple languages
Voice CloningUpload a short audio sample to clone a specific voice
Voice DesignAdjust parameters (pitch, style, accent) to create a custom voice

Key Options ​

OptionDescription
VoiceSelect preset, cloned, or designed voice
SpeedAdjust speaking rate (e.g., 0.75× for slow, 1.25× for fast)
Language / LocaleEnsure the TTS engine matches the text language
TTS ProviderChoose the provider (OpenAI TTS, system voices, etc.)

Best Practices ​

TIP

When using Voice Cloning, provide a clean recording of 10–30 seconds with minimal background noise for the best clone quality.

TIP

Pair TextToSpeech with a Translation step before it to produce dubbed audio in a different language.

  • Match the TTS voice language to the text language — a mismatch produces heavily accented or unintelligible output.
  • Use speed adjustment for accessibility: slower speech for language learners, faster for podcasts or narration.

Step 6 — ExportArtifact ​

Purpose: Save the workflow output to one or more files in your chosen format.

Subtitle Formats ​

FormatExtensionBest For
SRT.srtUniversal subtitles, video players, most platforms
VTT.vttWeb video (HTML5), YouTube, streaming
TXT.txtPlain text transcript, no timestamps

Audio Formats ​

FormatExtensionBest For
AAC.aacHigh quality, small file size
MP3.mp3Universal compatibility

Key Options ​

OptionDescription
FormatSelect one or more output formats
Output PathChoose the save location on disk
FilenameCustomize the output filename
Split by SpeakerExport separate files per speaker (requires diarization)

Best Practices ​

TIP

You can select multiple export formats in one step — for example, export both SRT and TXT simultaneously.

NOTE

Audio export (AAC/MP3) is only available when the workflow includes a TextToSpeech step upstream.

  • For video editing workflows, export SRT and drop it directly into your video editor.
  • For archiving meeting notes, TXT is the simplest and most portable format.
  • If your workflow includes Speaker Diarization, consider enabling Split by Speaker to produce per-participant transcripts.

Released under the MIT License.