Guides

Guides

Captions, voices and cutout

The AI features run on local models that download once from Settings and then work offline.

Nothing uploads. Captions, voices, cutout and enhance all run on your machine, on models that download from Settings the first time a feature needs one and never need the network again. Your footage never leaves your disk. Concat ships no model weights inside its installers: a cutout network, a voice bank and a caption model together outweigh the editor many times over, and most people never ask for all of them.

Auto-captions

Local Whisper transcribes the audio and puts styled captions on the timeline. Pick a size: bigger is more accurate and slower, and the larger ones want 16 GB of RAM. Each size comes in an English-only variant and a multilingual one that detects the language on its own.

ModelEnglishMultilingual
Tiny78 MB78 MB
Base148 MB148 MB
Small488 MB488 MB

Text-to-speech and voice cloning

Three families of voice, one Speech panel.

ModelWhat it doesSize
KokoroBuilt-in speakers in several languages. Quick.132 MB (int8) or 350 MB
Pocket TTSKyutai’s English voice. No speakers of its own: every voice is a few seconds of a recording. Two recordings come with it, and any clip in the bin can be the third.98 MB
Chatterbox TurboResemble AI’s voice, the closest to a studio read. Every voice is a recording, and it reads [laugh], [chuckle] and [sigh] as what they say. Slower, and the biggest download.About 1.1 GB, in nine files

Background removal and enhance

ModelWhat it doesSize
Person cutoutCuts a person out of the frame15 MB
Object cutoutCuts an object out of the frame179 MB
Cutout brushPaint the mask yourself40 MB
EnhanceRestores and upscales a clip’s picture, one clip at a time5 MB

Voice cleanup

One switch to denoise, enhance the voice and level the loudness. Plus chipmunk, robot, telephone and friends. These need no download.

Edit this page on GitHub