Use case
Subtitles that match the audio, to the word.
Most short-form video is watched with the sound off, which makes the captions the soundtrack. Auto-generated ones are only useful if they are accurate to the word and can be made to look like you: MakeRoll transcribes with Whisper, times every word individually, and lets you restyle the result rather than choosing from a menu of presets.
How it works
Transcribe
Whisper runs on our GPUs over the audio of the file or link you brought, and returns a transcript with a timestamp on every word rather than on every line.
Style them like your channel
Font, size, colour, outline, position and the way each word lands are yours to set, and can be saved as a template so the next video inherits them.
Burn them in on export
Captions are rendered into the video itself, so they survive every platform's player. The export is pixel-identical to the preview you approved.
What you end up with
Word-level timing
Captions land with the speech instead of a line at a time, which is what makes the read feel like the audio rather than a subtitle track.
Editable, not just generated
Whisper is very good and not perfect. Names, jargon and crosstalk are fixed by typing over them in the transcript, and the caption follows.
Translated captions on Creator
English, French and Spanish. The same clip goes out to an audience that does not share your language, without a second edit.
What this does not do
Captions are burned into the exported video. MakeRoll does not currently hand back a separate .srt file to upload alongside it.
Credits are source minutes: a 60-minute video costs 60 credits, however many clips come out of it. Editing and exporting are free and unlimited. Pricing
Try it on your own footage.
An hour of source video free every month: full pipeline, full editor, no watermark, no card.