CAPTIONS GUIDE

How to Make Karaoke (Word-by-Word) Captions in DaVinci Resolve

By Antonio Pennino · Updated September 2026 · 6 min read

Short answer: the caption style where each word lights up as it is spoken, the TikTok and Reels look, cannot be done with the normal subtitle track. It shows one static line at a time. Word-by-word captions are built from Fusion Text+ nodes driven by word-level timing. You can place them by hand, which is slow and fiddly, or let a plugin transcribe and place them for you. Both routes work on the free version, because Fusion and Text+ are free.

What "karaoke captions" actually means

A normal subtitle shows a full line for a few seconds, then swaps to the next. A karaoke caption keeps the line on screen but highlights the one word being spoken right now, so the viewer's eye follows the voice. Some variants reveal words one at a time instead. Either way, the defining ingredient is word-level timing: the editor needs to know the start and end of every single word, not just the sentence.

That is the whole reason it takes more than the subtitle track. Regular subtitles carry sentence timing. Karaoke needs a timestamp per word, and something on screen that can restyle one word at a time.

Why the native subtitle track cannot do it

DaVinci Resolve's subtitle track is built for broadcast-style captions, one static line at a time. There is no setting to highlight a single word inside a line, and no per-word timing to drive it. Even the Studio-only Create Subtitles from Audio feature produces sentence-level subtitles, not karaoke. To animate individual words you have to drop down into Fusion and use Text+, which is where all the styling power in Resolve lives.

The manual route, by hand in Fusion

It is possible without any plugin, if you have patience:

  1. Transcribe your audio, by ear or with an external tool, and note the timing of each word.
  2. Add a Text+ node and design your look, font, size, colours, a highlight style for the active word.
  3. Use character-level styling or separate nodes so one word can be coloured differently from the rest.
  4. Keyframe the highlight so it lands on the right word on the right frame, for every word, across the whole video.

It works, and it teaches you a lot about Fusion. But for anything longer than a few seconds it becomes hours of keyframing, and one re-edit of the voiceover means redoing the timing.

The automated route, with SmartSubs Pro

SmartSubs Pro does the two hard parts for you: it gets word-level timing, and it places the nodes. It transcribes the timeline audio locally with Whisper, derives the timing of every word, then builds the captions on a Text+ template you designed, highlighting the active word on the exact frame it is spoken. Because it uses your template, the karaoke matches your brand instead of a generic preset.

  1. Open your timeline in Resolve and launch SmartSubs Pro from Workspace › Scripts › Utility.
  2. Pick the Viral Sentence Karaoke preset.
  3. Choose a Whisper model. medium is a good balance; the English-only .en models are more accurate on English speech.
  4. Click Generate Captions. The words land on your timeline, timed and styled.

Want a different look? Edit the Text+ template once and every caption follows it. That is covered in the styling guide.

By hand or automated?

What matters By hand in Fusion SmartSubs Pro
Word-level timing Manual Automatic
Time for a 60s clip Hours Minutes
Uses your own Text+ style
Works on free Resolve
Survives a voiceover re-edit Re-run it

Skip the keyframing

SmartSubs Pro turns your voice into animated word-by-word captions on your own Text+ template, in minutes. One-time payment, no subscription.

Get SmartSubs Pro for $39.99

30-day money-back guarantee · Studio or free Resolve, any version

Frequently asked questions