How to Convert an MP4 Video to Text (and Shrink the File First)

By the Speakmi team · Updated October 2, 2026 · 2 min read

Speech recognition only needs the sound track of a video. If your MP4 is large, extracting or compressing the audio first makes transcription faster and keeps you within upload limits.

Option 1: transcribe the video directly

If the file is under 25 MB, open the voice.speakmi transcriber, drop in the MP4, and press Transcribe. Your browser pulls out the audio locally, so the video is not uploaded anywhere.

Option 2: extract the audio first

Most phone and screen-recording videos are far larger than 25 MB because of the picture, not the speech. Removing the video track often shrinks the file by 90 percent or more.

Why the settings above work

Speech models work with 16 kHz mono audio, so stereo music-quality settings add size without adding accuracy. At 64 kbps, one minute of audio is roughly half a megabyte, which means a 20 minute recording is around 10 MB.

Long videos

For lectures or webinars longer than 20 minutes, cut the file into parts using any free editor, transcribe each part, and join the text afterwards. Timestamps in the SRT file are relative to each part, so note the start time of each segment if you need continuous subtitles.

After transcribing

Download the TXT file for notes, or the SRT and VTT files for subtitles. Read through the text once, because names, product terms, and numbers are the most likely to need correction. If the video has heavy background music, lower or remove it before transcribing for better results.

Related guides