---
title: "Audio to Video: Waveform Videos with Captions | Potto AI"
description: "Turn audio to video with an audiogram maker: upload a podcast or interview, and Potto picks a moment of up to 90 seconds and adds a waveform and captions."
canonical: "https://potto.ai/audio-to-video"
language: "en"
dateModified: "2026-10-10"
---

# Audio to Video: Audiogram Maker

Turn audio to video for feeds: upload a podcast, interview or talk of 10 minutes or less in MP3, WAV or M4A format, and Potto picks one continuous moment of 90 seconds at most, draws a waveform, adds captions, the show name, guest and cover art, and exports an MP4 with an SRT caption file to download.

## Audio to video samples from Potto's audiogram maker

Each audiogram started from one audio file and a sentence of direction: an interview in three frames, a solo talk, a recorded lecture and a literary reading. Select a poster to play it.

### Podcast interview audiogram

*1:1 · 60s*

A recorded interview as a square waveform video built around the guest's answer to what time actually is, with captions, a topic title, a Guest speaker label and two quote cards pulled from the answer.

### Same interview, vertical

*9:16 · 56s*

That same interview as a 56-second vertical audiogram for phones, with the Guest speaker label on screen from the first frame and a closing quote card.

### Same interview, wide

*16:9 · 90s*

A wide 16:9 version of the same answer, 90 seconds long, drawn with the wave style for a web page player.

### Solo talk, 90 seconds

*1:1 · 90s*

A 90-second square audiogram from one speaker's talk: the three legs of aeronautics research, computer analysis, wind tunnel testing and flight testing, under a wave-style waveform.

### Recorded lecture

*1:1 · 58s*

A public lecture as a square audiogram: one complete explanation of how big a super-eruption is, compared step by step with smaller eruptions, ending on a quote card.

### Literary reading

*9:16 · 23s*

The opening paragraph of Pride and Prejudice, read aloud, as a short vertical audiogram with a ring waveform and the title of the work as the show name.

*Sample audio: [NASA](https://www.nasa.gov/podcasts/houston-we-have-a-podcast/telling-time-on-other-worlds/) and [NASA Glenn Research Center](https://images.nasa.gov/details/GRC-2022-CM-0119), not subject to U.S. copyright; [USGS](https://www.usgs.gov/media/videos/pubtalk-112021-busting-myths-about-one-largest-volcanic-systems-world-top-10), public domain; [LibriVox](https://librivox.org/pride-and-prejudice-by-jane-austen-solo-project/), public domain. Potto is not affiliated with or endorsed by NASA, USGS or LibriVox. Each audiogram was generated from an uploaded audio file, not from a customer project.*

## From a full episode to one audiogram moment

On the left, the storyboard row behind a frame of the sample interview audiogram; on the right, that frame. When you turn audio to video here, Potto builds a new audiogram around one moment of your audio; it is not an MP3 to video conversion of a whole file.

*Storyboard row · interview audiogram*

Scene 1 of 6 · 9s. On-screen text

- Keeping Time on Other Planets
- What even is time?
- Guest speaker
- See, that is that is the magic question. And that's the question where some people are going to head nod and say, "I like that answer,"

*Video frame · 0:08 · 1:1*

### What the agent picked

One continuous stretch of the interview, never longer than 90 seconds, chosen from what the message asked for. The brief you confirm in chat says where it begins and ends. Nothing is removed from the middle or reordered, so the guest's answer plays as it was said. If the moment is wrong, ask for another in chat; there is no timeline handle to drag.

### What appears on screen

A waveform drawn from the audio, captions of what is being said, and the show name and speaker label you give it, with cover art when you add one. Captions appear sentence by sentence, so a viewer with the sound off can still follow the answer.

### What to check before you export

Read every caption for names, product terms and numbers, since speech can be transcribed wrongly. Check that the segment starts at the beginning of a sentence and ends on a complete thought, then download the SRT file from the export card if another player needs the captions.

## Pick the recording, then the moment worth hearing

To turn audio to video that people watch to the end, say who your audiogram is for and what it should make them do: play the full episode, follow the speaker or remember an idea. The kind of recording then decides which moment to ask for.

### Podcast audiograms

*1:1 · 60s*

Ask for the answer that makes the episode worth hearing, not the intro or a sponsor read. A podcast audiogram has room for one idea, so name it in your message, and check that the show details on screen match the episode you will link to.

### Interviews and guest spots

*1:1 · 90s*

Pick a moment where the guest says something specific: a number, a short story or a clear opinion. Put the guest's name on screen so it makes sense when shared on its own, and check that the segment opens with enough context to follow.

### Talks, classes and readings

*1:1 · 58s*

For a recorded talk, class or author reading, ask for one complete explanation or passage rather than a summary of everything. Spoken audio works best; music gives captions nothing to show. For a long recording, upload only the part that matters.

## How to turn audio to video

Use the audiogram maker in three steps with the creator above. You confirm the segment in chat before anything is exported, and credits are charged only when you export.

### 1. Upload audio and say what it is for

Upload audio in MP3, WAV or M4A format, up to 20 MB and 10 minutes long: a podcast episode, interview, talk or reading. Write a sentence on who will watch and which moment you want, such as the guest's answer about pricing.

### 2. Confirm the moment in the chat

Potto's agent picks a continuous segment of up to 90 seconds from your description, and the brief shows its start and end times. Confirm it, or describe a different moment in chat; there is no manual control for picking the segment.

### 3. Check the waveform video, then export

Check the waveform, captions and show details in a free preview. Export a 1080p MP4 at 30 fps, square by default or in another frame, and download the captions as SRT from the export card.

## What Potto's audio to video audiogram maker does

Many audio to video tools put a still image over a whole file, and an MP3 to video converter only changes the format of an audio file. Potto's audiogram maker generates a new, short waveform video from your audio instead: it picks a moment, draws a waveform around it and adds captions, so people scrolling with the sound off can still follow.

### A moment chosen from up to 10 minutes

Upload a recording and describe the moment you want; Potto's agent picks a single continuous stretch, 90 seconds at most, and the brief you confirm shows its timing. Ask for another moment if it is not the right one; the audio plays in order, with nothing removed from the middle.

- Audio up to 10 minutes long
- A continuous segment, 90 seconds max

*Opening of the interview sample*

Opening frame of a square audiogram made from a recorded interview

### A waveform video in three styles

The waveform moves with the voice in your audio and is drawn as bars, a wave or a ring. That motion tells a viewer at a glance that this is a recording worth turning up, which is the whole job of an audio waveform video.

- Bars, wave or ring styles
- Drawn from your own audio

*Wave-style waveform in the solo talk sample*

Audiogram frame with a wave-style waveform from a 90-second solo talk sample on aeronautics research

### Captions on screen, plus SRT

Captions are on by default: Potto turns the speech in your segment into captions shown sentence by sentence, and you can download them as an SRT file next to your MP4 for a player or site that takes caption files. Read them once for names and terms before you export.

- Captions shown sentence by sentence
- SRT caption file to download

*Captions in the vertical interview sample*

Vertical audiogram frame with a caption line under the waveform

### Show name, guest and cover art

Each audiogram carries your show details on screen, so a waveform video shared on its own still says which show it came from and who is speaking.

- Show name and guest on screen
- Cover art in the frame

*Show details in the reading sample*

Vertical audiogram frame from a reading, with a ring waveform, a Reader label and the title of the work as the show name

### Square by default, two more frames

The audiogram maker exports a 1080p MP4 at 30 fps. Square 1:1 (1080×1080) is the default for feeds; 16:9 (1920×1080) and 9:16 (1080×1920) are also available, so when you turn audio to video, one episode can have a wide version for its web page and a vertical one for phones.

- 1:1 by default, 16:9 and 9:16 too
- 1080p MP4, 30 fps

*Interview samples in 1:1, 9:16 and 16:9*

Three frames of the same interview audiogram in square, vertical and wide shapes

## Keep exploring Potto AI

An audio to video audiogram maker is one way to start. These pages start from topics, briefs and screenshots, or make images for your show.

- [Explainer Video](https://potto.ai/explainer-video.md): Start from a topic or notes, with AI voiceover.
- [AI Video Generator](https://potto.ai/ai-video-generator.md): Start from a brief and let the video type follow.
- [Product Launch Video](https://potto.ai/product-video.md): Show a real feature from your own screenshots.
- [AI Motion Graphics](https://potto.ai/ai-motion-graphics.md): Every Potto video style, from kinetic text to data stories.
- [AI Image Generator](https://potto.ai/ai-image-generator.md): Make square artwork for an episode from a prompt.
- [Uncrop](https://potto.ai/uncrop.md): Widen a square picture to fit a 16:9 frame.

## Audio to Video FAQs

What audio to video and audiograms are, which recordings work best, and the length limits, captions, export sizes and credits to know before you start.

### What is audio to video?

Audio to video is the process of turning a sound recording, such as a podcast, interview or talk, into a video people can watch and share where video is expected. The simplest audio to video method, the one most MP3 to video tools use, places one still image over the full recording. Potto makes an audiogram instead: it picks one continuous moment of 90 seconds at most from your audio and builds a new audiogram around it, with a moving waveform, captions plus your show details, exported as an MP4. It does not film footage or generate a digital presenter.

### What is an audiogram?

An audiogram, in podcasting and social media, is a short video that plays a piece of audio over a moving waveform, usually with captions and the show's artwork, so people can preview a recording in a feed. In medicine the same word names a hearing test chart; an audiogram generator makes the video kind. Potto's audio to video audiogram maker builds one from a segment of your own audio.

### How do I turn audio to video with Potto?

To turn audio to video, upload a WAV, M4A or MP3 file of 10 minutes or less in the creator at the top and write a sentence about who will watch and which moment you want. Potto's agent picks a continuous segment from your description, and the brief in chat shows where it begins and ends; confirm it or ask for a different moment. Then check the waveform, captions plus show details, export an MP4, and download the captions as SRT if you need them.

### Is Potto an MP3 to video converter?

Not in the file-format sense. MP3 files can be uploaded, but an MP3 to video converter wraps a whole audio file in a video container, often with one still image, so the result is as long as the recording. Potto's audiogram maker generates a new, shorter waveform video from your audio: a continuous segment no longer than 90 seconds, with captions and your show details. If you need a full episode as a single file with a static picture, an MP3 to video converter does that job; if you want a preview of the episode people can watch, turn audio to video with the creator on this page.

### What kind of audio works best for an audiogram?

Spoken audio works best: podcast episodes, interviews, panel talks, lectures and readings. Captions come from what is said, so clear speech with little background noise gives the most accurate caption text, and a recording with a strong answer or explanation gives the agent a clear moment to choose. Music is not what this audio to video audiogram maker is built for.

### How long can my audio and the audiogram be?

Your audio file can be up to 10 minutes long and 20 MB in size, and your audiogram is one continuous segment of up to 90 seconds. Because your own recording is the soundtrack, the length of the audiogram follows the segment that is chosen. For a longer episode, upload only the part of the audio you need.

### Can I choose which part of my audio is used?

Yes, through chat rather than a timeline. Describe the moment you want, such as the guest's answer about hiring or the closing minute of a talk, and Potto's agent chooses a segment and lists its start and end times in the brief you confirm. If it is not the right one, describe a different moment. There is no manual control for picking the segment yourself, and the segment always plays in its original order, with nothing removed from the middle.

### What is a waveform video?

A waveform video is a moving image in which a shape drawn from an audio signal moves with the sound, so viewers can see that someone is speaking even when their sound is off. Every Potto audiogram is an audio waveform video: the waveform follows the voice in your segment and is drawn in one of three styles, bars, wave or ring.

### Does my audiogram have captions, and can I get an SRT file?

Yes. Captions are on by default, and Potto adds captions of the speech in your segment, shown sentence by sentence. With captions on, you can also download them as an SRT file from the export card, for a video player or site that accepts separate caption files. Read the captions before you export, because names, product terms and numbers are the words most often heard wrongly.

### What sizes and formats can I export?

Potto's audiogram maker exports a 1080p MP4 at 30 fps. Square 1:1 (1080×1080) is the default for feeds, and 16:9 (1920×1080) for a website player or 9:16 (1080×1920) for phones are also available. A 4:5 frame and GIF export are not available.

### Can I upload a video podcast instead of audio?

Not at the moment. Potto currently cannot take an existing video as input, so for a video podcast, upload its audio track instead, within the same length limit. What you get is a new waveform video built around that audio, not a version of your original footage.

### How many credits does an audiogram cost?

Generating and previewing an audiogram is free. As with any Potto project, credits are charged only when you export, and the price is set by the second: 5 credits for each second of your exported audiogram, so a 45-second segment costs 225 credits, and a 60-second audiogram with the original audio exports for about 300 credits. When your original audio is the soundtrack there is no voiceover fee, and each AI illustration adds 23 credits. Each person can transcribe audio up to 20 times a day.

## Turn audio into an audiogram

Upload a podcast, interview or talk, confirm the moment in chat, and export an audio waveform video with captions as an MP4.
