01
Add your sources
A script (.docx, .txt or .pdf), your voiceover (.mp3 or .wav), and a folder of pictures named by scene. Background music is optional.
VisualBeats
by Tubeneur
Professional Image & Video Editing, Simplified.
VisualBeats listens to your voiceover and changes the picture exactly when you start talking about it — with word-timed captions, zoom motion and export up to 4K. Everything runs on your own PC.
Windows desktop app · Currently v4.0.3, in testing · Pricing announced soon
How it works
01
A script (.docx, .txt or .pdf), your voiceover (.mp3 or .wav), and a folder of pictures named by scene. Background music is optional.
02
Pick a motion effect, 16:9 for YouTube or 9:16 for Reels and Shorts, and one of ten caption presets.
03
One button. Whisper reads the voiceover, finds the word each scene starts and ends on, and builds the timeline.
04
Play it back in sync inside the app, swap any picture that doesn't fit, then render an MP4 up to 4K.
0:00
The city was already awake.
0:07
Nobody noticed the first sign.
0:11
By morning, the story had changed.
0:19
And that is where it ends.
A diagram, not a screenshot. Each picture changes on the word the narration actually reaches — not on a fixed interval — and the highlighted word is how the burned-in captions track the voice.
Features
OpenAI Whisper listens to your narration and times every scene to the spoken word. Choose the Base, Small or Medium model; it runs on an NVIDIA GPU when there is one and falls back to the CPU on its own.
Karaoke-style captions burned into the export, highlighting each word as it is spoken. One line at a time or one word at a time, with sliders for position and size.
Static, Zoom In, Zoom Out, Zoom In-Out and Zoom Out-In, rendered smoothly without the judder that usually gives still images away.
Switch to 9:16 and overlay a mock of Instagram's or TikTok's own interface, so you can see what the platform will cover before you export. Preview only — it never reaches the file.
3840×2160 or 1080p, at HQ or Standard bitrate. H.264 MP4 at 30 fps with AAC audio, tagged and fast-started so YouTube, Instagram, TikTok and Facebook accept it cleanly.
NVIDIA NVENC when it is available, with automatic fallback to CPU encoding. For reference, an 80-second voiceover syncs in roughly 24 seconds on a GPU.
Point it at a folder of tracks, reorder them, save the order, and set the music volume when you export.
Scripts can carry a second language line per scene, each shown in its own native script — so you can assemble a video in a language you don't read. Voice sync works in any language Whisper supports.
Voice sync and rendering happen locally — your scripts, voice and images are never uploaded, and there is no per-video cost. Internet is only needed for the first model download, the licence check and updates.
Caption presets
Script Pro's bilingual Visual Storyboard export drops straight into VisualBeats as the script — the English line to edit from, the native line for the voice-over. Write and translate a channel in one tool, assemble the video in the other, in a language neither of you speaks.
System requirements
Pricing
Coming soon
The app is built and in testing at v4.0.3. How it will be sold is the last decision left — so there is no price here yet rather than a number we would have to change.
Current version 4.0.3 Windows desktop app Updates install themselves
If you want early access, or want to hear the moment pricing lands, get in touch and we'll let you know.