CanDoYa
EN

Auto Caption Generator

Runs entirely in your browser - no upload, no sign-up.

Add a video to generate private, editable captions on your device.

Share this tool

How can I generate captions for a video automatically?

Add a video and this auto caption generator runs Whisper speech recognition locally on your device without uploading the file. It creates timed, editable caption cues that you can preview and download as SRT or VTT. Your browser must download the AI model on first use.

How to use

  1. 1Choose a video. Drop an MP4, WebM, MOV, M4V, or MKV file onto the tool, or click the upload area to select one.
  2. 2Generate timed captions. Click Generate captions. On the first run, wait while the multilingual Whisper model downloads, then let your device process the audio track.
  3. 3Review every cue. Play captions from the editor, correct the wording, and adjust start or end times where the automatic timing feels early or late.
  4. 4Export the caption file. Download SRT for broad editor and platform support, or VTT for HTML video and web publishing.

Who it's for

This tool extracts the audio track inside your browser, converts it to the 16 kHz mono signal Whisper expects, and creates caption timings in a background worker. WebGPU accelerates supported computers; other browsers use a slower CPU fallback. The editor lets you correct both wording and cue boundaries before export.

FAQ

Is my video uploaded to a server?

No. The browser reads the selected file and passes decoded audio samples to a worker on the same device. The worker downloads Whisper model files from Hugging Face, but it does not upload your video, audio, or captions. CanDoYa does not receive or store the file.

Is the auto caption generator free?

Yes. There is no account, payment, caption minute quota, watermark, or export lock. Your device performs the transcription, so the practical costs are local processing time, memory use, and the first model download. Your internet provider may count that download on a metered connection.

What video size and length can I caption?

The uploader accepts files up to 500 MB. There is no fixed minute limit, but long videos require more memory and take longer because decoding and AI inference happen locally. For hour-long recordings, use a modern laptop, close memory-heavy tabs, or split the video into smaller sections first.

Which video formats work?

The tool accepts MP4, WebM, MOV, M4V, MKV, AVI, MPEG, and OGV containers. Actual audio decoding depends on the codecs built into your browser and operating system. MP4 with AAC audio and WebM with Opus audio usually have the widest browser support.

Can I edit the automatic captions?

Yes. Each cue has editable text, start time, and end time. Play a cue to jump the preview to its starting point, correct recognition errors, and adjust the boundary. Exports stay disabled if a cue is empty, ends before it starts, or extends beyond the video.

What is the difference between SRT and VTT?

SRT is a widely supported subtitle format used by video editors and publishing platforms. WebVTT uses a similar timed-text structure but is designed for the web and works directly with the HTML video track element. This tool exports both without adding visual styling or burning text into the video.

Does caption generation work without WebGPU?

Yes. The tool tries WebGPU first because a compatible graphics processor can run Whisper faster. If WebGPU is unavailable or fails, it retries with WebAssembly on the CPU. The fallback supports more browsers but can be much slower, especially for long videos on phones or older computers.

Are automatic captions accurate enough for accessibility?

They are a useful draft, not a finished accessibility deliverable. Noise, music, overlapping voices, accents, uncommon names, and technical terms can cause mistakes. Automatic speech recognition also omits many sound descriptions and speaker labels, so review the video and add meaningful non-speech cues before publishing formal closed captions.