Skip to content
INDXR.AI
row of worn leather books with foreign-script spines
Troubleshooting

Non-English YouTube transcripts: get the right language

IE
INDXR.AI Editorial
Published April 24, 2026 · Updated August 26, 2026

You open a Spanish video, click the transcript, and get Korean. Or Dutch. Or one auto-generated option in a language nobody in the video speaks.

This is not your settings, and it is not rare. It has been a known problem for years, and the text you end up reading is often a translation of a guess rather than what was said. Professional tools fail the same way: editors report English interviews transcribed as Mandarin, with auto-detect switched on, and the answer is always to set the language by hand and start over.

INDXR works from the audio. A free account comes with 50 credits.

Why YouTube gives you the wrong language

Two different things get confused under one word.

Auto-captions are a speech-recognition guess at what was said. They only exist in a limited set of languages, and they are wrong more often on accented or noisy audio.

Auto-translate takes those captions and runs them through machine translation. YouTube has combined the two since 2012 to produce subtitles in more than fifty languages, so what you see can be a translation of a guess, two steps away from the audio.

The transcript panel does not always tell you which of the two you are looking at, and your own account language influences what surfaces first. That is how a Spanish lecture ends up showing you English text that nobody said.

What INDXR does instead

We start from the audio, not from a label. A caption file carries a language tag, and that tag is often wrong or describes a translation rather than the video. So we do not trust it. We work from what is actually spoken.

We never hand you a translation dressed up as a transcript. If the only text available for a video is machine-translated, we treat that video as having no transcript at all and offer to transcribe the audio properly instead. A translation of a guess is worse than nothing, and it is the reason you are reading this.

The language is recognised, then checked. Recognition happens automatically across up to 99 languages, and the result is verified against the finished transcript before we label it. A talk in one language that is full of terms from another is exactly where automatic recognition slips, and that check is there to catch it.

You do not configure any of this. You paste a link or upload a file, and the transcript comes back in the language that was spoken.

The method chooser after pasting a link: free caption extraction alongside AI transcription, which reads the audio and detects the language automatically.
When a video's only captions are the wrong language, you transcribe the audio instead, and the language is recognised from what was spoken.

What you get back

An example from our own library: an 81-minute Arabic lecture, transcribed with AI, 917 segments, with a confidence score of 0.98. It opens like this:

السلام عليكم في هذه الحلقة سأعيد مرة أخرى مناقشة إنكار بعض فقهاء المعاصرين لمسألة المست الشيطاني والتي تقريباً أجمعت عليها الأمة ولم يخالف فيها إلا قلة قليلة فقط

The text reads right to left where the language does, speakers are separated where there is more than one voice, and timestamps work as they do in any other transcript. Every export format is available.

The detected language travels with the file. It sits in the VTT header, in the Markdown front matter and in the CSV metadata, so a player or a note app knows what it is reading. You also see it next to your transcript in the library.

Which languages

YouTube captionsAny language the video has its own caption track for
AI transcriptionUp to 99 languages, recognised automatically
Highest-quality model18 languages natively, including Arabic

Accuracy differs by language and we do not publish one number for all of them. The per-language accuracy bands are on the reference page, taken from the provider rather than from our own measurements, and honest about which languages are weaker.

What it does not do

It does not translate. INDXR transcribes what was spoken, in the language it was spoken in. If you need the text in another language, take the transcript to a translation tool of your choice. We would rather give you an accurate original than an automatic translation of a guess, which is the thing that sent you here.

You cannot pick the language. Recognition is automatic. Nothing to set, nothing to get wrong.

What it costs

Captions in the original languageFree
AI transcription1 credit per minute, in any language
Text and subtitle exportsIncluded

The Arabic lecture above cost 81 credits for its 81 minutes. There is no subscription, and credits never expire.

Try it

Frequently Asked Questions

Why does YouTube show captions in a language nobody spoke?
Because you are usually looking at an auto-translation of auto-captions rather than the original track, and your account language influences which one appears first.
Does INDXR use those translated captions?
No. If the original-language text is missing, the video counts as having no transcript and we transcribe the audio instead.
Can I choose the language before transcribing?
No. The language is recognised from the audio. There is nothing to set.
Can I correct the detected language afterwards?
Not at the moment. The language is verified against the finished transcript, which catches the common failures, but there is no manual override yet.
Does it translate the transcript?
No. You get the text in the language that was spoken. Take the export to a translation tool if you need another language.
What about a video with two languages in it?
One language is recognised for the whole transcript, so a bilingual video gets a single label. The words are still transcribed as spoken.
How accurate is it in my language?
It varies. The reference page lists accuracy bands per language and is honest about which ones are weaker.

Sources

See also