Skip to content
INDXR.AI
long receding row of clear cassette tapes standing on sunlit sand, shadows trailing into the distance
Reference

Every format INDXR transcribes, and where each one comes from

IE
INDXR.AI Editorial
Published August 31, 2026 · Updated August 31, 2026

A file you want the words out of arrives in whatever format the thing that made it decided on. An iPhone gives you M4A. WhatsApp gives you OPUS. A download gives you MKV. A recorder set to high quality gives you WAV or FLAC. None of that is a choice you made, and none of it should be a step you have to undo before you can read what was said.

INDXR accepts fifteen formats, audio and video together, and you upload the file exactly as it reached you. There is no conversion step, no “export as MP3 first”, and no question about which container you have. It costs one credit per minute of audio, and a free account includes 50 credits, so a real recording can go through before you spend anything.

The full list

AudioMP3, MPGA, M4A, AAC, WAV, OGG, OPUS, FLAC
VideoMP4, MPEG, WEBM, MOV, FLV, AVI, MKV
Maximum file size500MB
Maximum length10 hours per file
Languages99, detected automatically

Video files are handled the same way audio files are: the audio track is taken out and the picture is discarded. Nothing about the image is read, so a screen recording gives you what was said about the screen, not an account of what was shown on it.

Audio formats

M4A — the iPhone voice memo

M4A is MPEG-4 audio: the same container family as MP4, holding audio alone. It is what the Voice Memos app on an iPhone produces, and what iTunes and Apple Music have used for years.

It is also the format people most often get stuck on. A file that plays perfectly on the phone that recorded it gets rejected by a converter that only knows MP3, and the obvious next move — find something that turns M4A into MP3 — adds a step, a quality loss and a second piece of software to a job that was one step to begin with.

Upload the M4A. A forty-minute memo costs forty credits and comes back punctuated, split by speaker and timestamped.

MP3 — the one everything can read

MPEG-1 Audio Layer III, and still the format most things fall back to. Podcast downloads, older recorders, anything exported “for compatibility”. Nothing to say about it except that it works, which is the point.

WAV — uncompressed, and usually large

WAV holds audio without compression, which is why a field recorder or an interview rig set to archival quality writes WAV and why the files are big. An hour of stereo WAV at CD quality runs to roughly 600MB, which is over the 500MB limit.

If a WAV file is too large, the fix is not to convert it to MP3 and lose the quality you recorded it at. Split it at a pause between sentences and upload the parts. You pay per minute of audio either way, so splitting costs nothing extra.

OPUS — the WhatsApp voice note

OPUS is a codec designed for speech over the internet, standardised by the IETF in 2012. WhatsApp uses it for voice messages, Discord uses it for calls, and a browser recording a microphone will often produce it.

Export the voice message from the chat and upload the file as it comes. It arrives with a .opus extension and goes through as it is.

Voice notes are short and the credit cost follows: a three-minute message costs three credits, which is about eight cents at Plus pricing.

OGG — the container OPUS usually travels in

OGG is an open container from the Xiph.Org Foundation, and the thing inside it is usually Vorbis or Opus. Where a .opus file is Opus audio labelled as such, a .ogg file is Opus or Vorbis audio in a container that could hold either.

Both upload. You do not need to know which one you have, or to check what is inside the container before sending it.

FLAC — lossless, for recordings that matter

FLAC compresses audio without discarding anything, which is why it is used for archival recordings, music masters and anything someone expects to still be working from in ten years. Files are smaller than WAV and larger than MP3.

There is no accuracy gain from uploading FLAC rather than a good MP3 of the same recording — speech recognition does not hear the difference. Upload the FLAC because it is the file you have, not because it will read better.

AAC — the successor MP3 never quite replaced

Advanced Audio Coding is what sits inside most M4A and MP4 files, and it also exists as a bare stream with an .aac extension: broadcast captures, some Android recorders, files pulled out of a video container. Both the wrapped and the bare version upload.

MPGA — MPEG audio under another name

MPGA is MPEG audio, in practice usually the same thing as MP3 with a different extension attached by whatever produced it. It goes through the same way.

Video formats

MP4 — the default for everything

MPEG-4 Part 14, the container most cameras, phones and editors write by default. The audio track inside is usually AAC. Upload the MP4 and the audio is taken out for you.

MOV — what an iPhone films

MOV is Apple’s QuickTime container and it is what an iPhone camera records. It behaves like MP4 in almost every respect, and the reason it exists as a separate line in the list is that plenty of tools accept one and not the other.

MKV — what downloads arrive as

Matroska is an open container that can hold nearly any combination of video, audio and subtitle tracks, which is why archives, rips and anything that needed more than one audio language ends up as MKV.

It is also the format that most often has nowhere to go, because handling it properly means dealing with a container that makes very few assumptions. Upload the MKV.

WEBM — the browser’s recording format

WEBM is the open container the WebM Project built around VP8 and VP9 video, and it is what a browser writes when a page records your screen or your camera. Loom-style tools, in-browser recorders and a good deal of what comes off the web is WEBM.

AVI — the old one that will not go away

Audio Video Interleave dates to 1992 and still turns up: old camcorder footage, screen captures from software that has not been updated in a decade, files that have been sitting on a drive since before phones had cameras. It uploads.

MPEG and FLV — the archive cases

MPEG-1 and MPEG-2 video, and Flash Video. Neither is something anyone produces on purpose now, and both are what you find when you go back through old material. They are in the list because a format nobody makes any more is exactly the kind of file that has no easy path anywhere else.

What you do not have to do first

Convert it. There is no format in the list that needs turning into another format in the list. Uploading an M4A is not worse than uploading an MP3 of the same recording, and converting between them costs you time and a little quality for no gain.

Extract the audio from the video. The audio track is taken out for you. You do not run the file through anything to separate them.

Check what is inside the container. MKV, OGG and MP4 can each hold several different codecs. Which one yours holds is not a question you need to answer.

Compress it to fit. The only reason to touch a file before uploading is if it is over 500MB or longer than ten hours, and the answer to both is to split it at a pause rather than to compress it.

The two limits

500MB per file. This binds in practice only on uncompressed audio and long high-resolution video. An hour of MP3 is around 60MB; an hour of stereo WAV is around 600MB.

Ten hours per file. Long enough for almost anything. A recording that runs longer gets split, and because you pay per minute rather than per file, splitting changes nothing about the cost.

What comes back

The same three things regardless of what you uploaded, because by the time transcription starts the format has stopped mattering:

The text, punctuated into real sentences, separated by speaker, and timestamped so you can jump back to the moment something was said. Exports as plain text, Markdown, CSV or JSON.

The subtitles, as SRT and VTT, rebuilt to broadcast conventions rather than left as raw fragments: lines capped at 42 characters across at most two lines, cues breaking on sentence boundaries, speaker names carried through.

The machine-readable version, a chunked JSON with metadata and timestamps shaped for a vector database. This is the only export that costs anything on top: 1 credit per ten minutes.

What it costs

One credit per minute of audio, rounded up, measured from the length detected after upload rather than the file size. A 200MB WAV of a thirty-minute meeting costs the same thirty credits as a 30MB MP3 of the same meeting.

LengthCreditsCost at Plus pricing
3 minutes3 credits€0.08
30 minutes30 credits€0.75
1 hour60 credits€1.50
2 hours120 credits€3.00

Plus is €25 for 1,000 credits. There is no subscription, and credits never expire.

Frequently asked questions

Do I need to convert my M4A to MP3 first?

No. M4A uploads as it is, and converting it to MP3 first would cost you a step and a little audio quality without making the transcript any better. Drop the file in as it came off the phone.

Can I transcribe a WhatsApp voice message?

Yes. WhatsApp exports voice messages as OPUS, which is one of the fifteen supported formats. Export the message from the chat and upload the file without renaming or converting it.

What if my file is bigger than 500MB?

Split it at a pause between sentences and upload the parts separately. Because you pay per minute of audio rather than per file, splitting costs you nothing extra. This comes up most with uncompressed WAV, where an hour of stereo runs to roughly 600MB.

Does the format affect how accurate the transcript is?

Not meaningfully. Recording quality matters — a clear microphone in a quiet room beats a phone across a table — but a lossless FLAC and a decent MP3 of the same recording transcribe to much the same text.

Which format should I upload if I have a choice?

Whichever one you already have. If you are choosing what to record in, anything at or above 128 kbps in a common format is well past the point where the model stops noticing.

What happens to my file after it is transcribed?

It is deleted as soon as transcription finishes, and only the text stays in your library. Everything is processed inside the EU, the transcription provider is opted out of training on your data, and its own retention is set to one day, the shortest it offers.

Sources

See also