Word Count in Arabic, Hindi, Urdu, Korean and Other Languages
3 min read
Short answer
A word count depends on knowing where one word ends and the next begins. Most writing systems mark that with a space, and for those, counting works exactly as it does in English. A few do not, and they need a different approach.
The word counter handles all of the languages below. Paste your text and the count is right for the script it is written in.
Languages with spaces between words
Arabic, Urdu and Persian. Words are separated by spaces, so each space-separated word counts once. Diacritics (harakat) are part of the word and do not change the count. Arabic punctuation such as ، and ؟ is not counted as a word, and ؟ and the Urdu full stop ۔ end a sentence.
One thing to check in Urdu: text typed on some keyboards leaves out the space after letters that do not join, so two words can run together into one. If your count seems low, look for words that appear joined.
Hindi, Marathi, Nepali and other Devanagari languages. Spaces separate words, so the count is by space. The danda (।) that ends a sentence is punctuation, not a word, and the sentence counter treats it as a full stop.
Filipino (Tagalog), Indonesian, Malay, Swahili, Vietnamese and other Latin-alphabet languages. Counted exactly like English. In Vietnamese each syllable is written separately, so a two-syllable word such as “Việt Nam” counts as two. That matches Microsoft Word and most style guides.
Korean. Korean puts spaces between words (eojeol), so it is counted by spaces, not by syllable block. “나는 학교에 갑니다” is three words. Some older tools count every Hangul syllable as a word, which makes a Korean count several times too high.
Languages without spaces
Chinese. Chinese has no spaces, and word boundaries are open to interpretation, so length is measured in characters (字数). Each Chinese character counts as one word here, which is also how Microsoft Word counts. For a proper character count with punctuation separated out, use the Chinese character counter.
Japanese. Also written without spaces. Kanji, hiragana and katakana each count as one, which is how Microsoft Word reports “Asian characters”. Japanese length limits, like Chinese ones, are usually given in characters (文字数).
Thai, Lao, Khmer and Burmese. No spaces between words, but unlike Chinese, limits are usually given in words. The word counter splits these using the word dictionary built into your browser, so “ฉันพูดภาษาไทยได้” counts as five words. Dictionary splitting is very good but not perfect on names and new words, so allow a small margin.
Characters in other scripts
The character counter counts each letter as one character in any script, including accented Latin letters, Arabic and Devanagari. In Devanagari, a vowel sign (matra) is stored as a separate character from its consonant, so “हिंदी” is five characters to a computer even though it is written as two syllables. Word, Google Docs and most online forms count it the same way.
Frequently asked questions
Why does my Chinese word count match the character count? Because each Chinese character is counted as one word. That is the standard convention, and it is what Word does.
Is the count the same as Microsoft Word’s? For languages with spaces, and for Chinese and Japanese, it follows the same rules, so ordinary text gives the same total. Thai can differ by a word or two, because each program’s word dictionary splits a few words differently.
Do translators charge by word in these languages? For languages with spaces, usually per word. For Chinese, Japanese and Korean source text, usually per character.