Language Detector
Paste text and get an approximate guess at its language. Covers 14 Latin-script languages by common-word scoring, and recognizes Chinese, Japanese, Korean, Arabic, Cyrillic, and other scripts by their Unicode ranges. Works best on 20+ words of natural text.
Top candidates (approximate)
This is an approximation, not a certified identification. Latin-script detection scores your words against embedded lists of the most common words in English, Spanish, French, German, Portuguese, Italian, Dutch, Swedish, Polish, Turkish, Indonesian, Romanian, Tagalog, and Vietnamese, with a boost for language-specific accented characters. Short texts, names, code, and mixed-language text confuse it; 20+ words of ordinary prose gives the best results. Non-Latin scripts (Chinese, Japanese, Korean, Arabic, Hebrew, Cyrillic, Greek, Thai, Devanagari, Bengali, Tamil) are identified by Unicode character ranges, which detects the script reliably but cannot distinguish languages that share a script (for example Russian vs Ukrainian, or Arabic vs Persian).