Readable cues
Sentence boundaries provide a dependable base for review and playback.
SUBTITLES & CAPTIONS
Edit sentence cues in the transcript, then export SRT or VTT without rebuilding the subtitle track.
Clarity comes from keeping the cue close to the spoken sentence.
Every transcription returns a usable sentence-level timeline. When MAI returns word timestamps, CastTranscript records that capability. When it does not, the editor says so plainly and never invents alignment from another model.
What survives the edit
Changing the transcript should not detach the words from the media. Each CastTranscript segment keeps a start and end time, speaker label and editable text. Export uses that same project state, so a corrected name or sentence is reflected in the next SRT or VTT download.
Sentence boundaries provide a dependable base for review and playback.
Use SRT in most editors and platforms, or WebVTT for web players.
Keep Chinese and English in the same cue instead of forcing a second timeline.
Word-level controls only appear when the provider actually returned word times.
Before you publish
Start with the first cue, a fast exchange between speakers and the final minute. Check that captions appear after speech begins, remain long enough to read and do not carry a speaker name into the wrong turn. For video, inspect the exported file in the destination player because line wrapping varies by platform.
When a response contains only sentence timestamps, CastTranscript keeps the honest sentence-level track. It does not run another speech model and force the words from one result onto timing from another. That restraint matters when the subtitle file is the deliverable.
Read the complete audio-to-SRT workflowQuestions, answered
Both store subtitle text with cue times. SRT is widely accepted by editors and video platforms. WebVTT is designed for web video and supports browser-oriented metadata.
Yes. Sentence-level start and end times are enough to create SRT or VTT. Word-level highlighting is only enabled when the provider returns verified word timing.
They can. Rename speakers in the project before export, then include or remove those labels for the delivery format you need.
Yes. Mixed-language text remains editable in the same timed document. Review punctuation, names and line length before publishing.
Private by default
Upload the final recording, correct the timed transcript and download SRT or VTT.