AI Workflow Engine
The core innovation of DuRT is its modular workflow architecture. Instead of isolated tools, DuRT allows you to construct and execute automated pipelines that combine audio capture, speech recognition, LLM intelligence, voiceover, and artifact exporting.
The 6 Standard Workflow Steps
[ 1. Media Input ] ➔ [ 2. Speech ASR ] ➔ [ 3. LLM Polish ] ➔ [ 4. TTS Voice ] ➔ [ 5. Export Artifacts ]| Step Type | Icon | Responsibility |
|---|---|---|
RealtimeASR | 🎙️ | Capture live microphone or system audio stream in real-time |
MediaTranscription | 📁 | Ingest audio/video files or media URLs with diarization & word timestamps |
SubtitleInput | 📝 | Ingest existing subtitle files (.srt/.vtt) or history tasks for downstream AI |
LLMProcess | 🤖 | AI-driven proofreading, translation, summarization, or semantic segmentation |
TextToSpeech | 🗣️ | Synthesize multilingual audio voiceover from transcribed or polished text |
ExportArtifact | 💾 | Automatically save formatted subtitle files, audio tracks, or logs to disk |
Workflow Hub (Template Center)
DuRT comes with pre-built workflow templates for common scenarios:
- Live Meeting Scribe: Realtime ASR ➔ Live Subtitles ➔ LLM Meeting Summary ➔ Save Markdown.
- Video Localization Pipeline: Media Transcription ➔ LLM Translation ➔ TTS Dubbing ➔ Export Dual-language SRT.
- Podcast Cleaner: Media File ➔ Whisper Word Timestamps ➔ LLM Filler-word Removal ➔ Export Cleaned Text.
Custom Workflow Builder
You can customize existing templates or create your own pipelines:
- Pin your most frequent workflows to the Dashboard top row.
- Choose which engines power each step independently (e.g., Apple Speech for ASR + DeepSeek for LLM + OpenAI for TTS).
- Save and share workflow configurations across your team.
