Tanglish Subtitles: Tamil Captions in English Letters
Captioning tools give you Tamil script or garbled English. Romanised Tamil runs 1.5 to 2 times longer per line, has no standard spelling, and needs the ta-Latn tag.
Tamil creators publishing to Reels and Shorts keep running into the same wall. The captioning tool hears Tamil correctly, and then writes it in Tamil script, which is not what most of the audience reads in a caption. The alternative most people try, forcing an English transcription, produces nonsense, because the model is trying to find English words in Tamil speech. What is wanted is the third thing: Tamil words spelled in Latin letters, the way people actually type to each other.
Why the mainstream tools do not give you this
This is not a gap in accuracy. Tools like VEED, Kapwing, Submagic and HappyScribe detect Tamil well and then output Tamil script, because that is the correct rendering of the language and the obvious thing for a subtitle file to contain. General transcription models pointed at English do the opposite and garble the audio into English words that were never said. Neither is wrong, exactly. Neither is Tanglish.
Romanising makes every line longer, and that breaks your line limits
This is the part that catches people out, and it is measurable. Tamil script is compact: one akshara carries a whole syllable. Latin letters spell that syllable out. The same words get substantially longer:
Tamil script characters romanised characters
வணக்கம் 5 vanakkam 8
நன்றி 3 nandri 6
சாப்பிட்டீங்களா 8 saapitteengala 14Roughly 1.5 to 2 times longer, consistently. Applied to a whole caption line, that is the difference between a line that fits and one that does not:
வணக்கம் நண்பர்களே இந்த வீடியோவில் நான் சொல்றேன்
30 characters, fits a 32-character vertical line
vanakkam nanbargale indha videovil naan solren
46 characters, does notSo a caption limit that was comfortable in Tamil script overflows the moment you switch to Tanglish, and the player wraps it somewhere you did not choose. Plan for shorter sentences rather than the same sentences in a different alphabet. Where the break should land once a line has to split is covered in caption line breaks for short-form video.
A second, invisible version of the same problem
That Tamil sentence is 30 characters to a reader and 47 UTF-16 code units to a computer. Most subtitle tools measure line length with a plain string length, so they see 47 and wrap a line that was never too long. If you work in both scripts, you will meet this in the Tamil-script file and not in the Tanglish one, which makes it look like the tool is behaving randomly. It is explained in subtitle character limits.
There is no standard spelling, so pick one and hold it
Tamil in Latin letters has no official orthography. Long vowels and retroflex consonants have no settled Latin spelling, so the same spoken word can reasonably be written several ways, and an automatic transcript will not be consistent about which one it picks. Across a three-minute video the same word can appear two or three ways, which reads as carelessness even though every spelling is defensible.
This is worth one pass before export rather than a judgement call per line. The subtitle spelling checker groups words that are near-identical across a file and lets you settle on one, so the decision is made once for the whole track instead of forty times.
Label the file for what it is
A Tanglish track is not Tamil and it is not English, and tagging it as either one misleads whatever consumes it next. The script subtag exists for exactly this: ta-Latn is Tamil written in Latin script, the same wayur-Latn is Roman Urdu. The reasoning, and where the tag actually goes in a file, is in what language a Roman-script subtitle file is.
A short checklist
- Generate the caption draft, then confirm it is romanised and not Tamil script.
- Expect lines to run 1.5 to 2 times longer than the same words in Tamil script.
- Re-wrap for the vertical frame, nearer 32 characters than 42.
- Settle repeated words on one spelling across the whole file.
- Tag the track
ta-Latnwhere the platform lets you.
AutoCaption produces an editable caption track before export, which is what makes the spelling and line-break passes above possible at all. Start in the caption generator, or use the Tamil subtitle generator when Tamil script is what the audience should read.
Quick answers
Why does my captioning tool write Tamil script instead of Tanglish?
Because Tamil script is the correct rendering of the language, and that is what a subtitle file is normally expected to contain. Tools that detect Tamil accurately still output the native script. Forcing an English transcription instead does not help, since the model then looks for English words in Tamil speech.
Do Tanglish captions need shorter sentences?
Yes. Latin letters spell out syllables that one Tamil character carries, so the same words run roughly 1.5 to 2 times longer. A sentence of 30 characters in Tamil script can reach 46 romanised, which overflows the 32-character line a vertical video wants.
What language code should a Tanglish subtitle file use?
ta-Latn, meaning Tamil written in Latin script. Tagging it as Tamil implies Tamil script, and tagging it as English is simply wrong. This is the documented use of the script subtag, not a workaround.