What is the Turkish I problem?

Turkish has two letters for i: dotted i becomes İ in capitals, and dotless ı becomes I. Languages with locale-independent casing, such as JavaScript, apply a different default rule: i becomes I, and a capital İ can split into i plus a combining dot when lower-cased. The same code therefore gives different text depending on the environment's locale.

This matters for keyword comparisons, permission checks, file and URL names, e-mail matching and sorting. The “Turkey test” bug class comes from here: the code looks fine until it runs on a machine set to tr-TR.

What the tool shows

Paste one word per line above, and optionally a compare text in the second field. For every line the tool computes seven rows under both rules: upper case, lower case, length after lower-casing, NFC and NFD lengths, equality, regular expression and slug. The lines are also sorted with Array.prototype.sort() and Intl.Collator("tr"), and the lines that differ are flagged.

Differing results get a yellow background and a “Differs” label, so colour is never the only signal, and the differing code points are written out. U+0307 is invisible, so it appears as “U+0307 COMBINING DOT ABOVE”. To convert text, use the case converter; this tool only diagnoses.

JavaScript results for the example inputs
InputtoUpperCase()toLocaleUpperCase("tr-TR")toLowerCase()toLocaleLowerCase("tr-TR")
istanbulISTANBULİSTANBUListanbulistanbul
IğdırIĞDIRIĞDIRiğdırığdır
İzmirİZMIRİZMİRi + U+0307 + zmir (6 code units)izmir
fileFILEFİLEfilefile
TITLETITLETITLEtitletıtle

Four letters, two pairs: İ/i, I/ı and U+0130, U+0131, U+0307

Turkish has two pairs: i with İ, and ı with I; the English-like rule merges both into i and I. The last row below is the combining dot that lower-casing can leave behind.

Characters involved in the I problem
CharacterCode pointUnicode name
İU+0130LATIN CAPITAL LETTER I WITH DOT ABOVE
iU+0069LATIN SMALL LETTER I
IU+0049LATIN CAPITAL LETTER I
ıU+0131LATIN SMALL LETTER DOTLESS I
combining dotU+0307COMBINING DOT ABOVE

"İ".toLowerCase() returns the two code unit string "i\u0307" by default and the single letter i under tr-TR; on screen they look alike, but length and equality change.

The rule for identifiers and keys

Machine-owned identifiers (protocol names, HTTP headers, configuration keys, file extensions, user names) must not change with the user's language. Use a locale-independent or ordinal comparison for them.

  • Java: toUpperCase(Locale.ROOT); equalsIgnoreCase is locale-independent but not Turkish-aware.
  • C#/.NET: ToUpperInvariant() or StringComparison.OrdinalIgnoreCase.
  • JavaScript: toLowerCase() and toUpperCase() are already locale-independent; still define which characters you accept, for example constrained ASCII.

Restricting identifiers to ASCII is often the sturdiest option: you decide the accepted character set instead of assuming arbitrary Unicode identifiers are safely equivalent.

The rule for user-visible text

For titles, names and labels shown to Turkish readers, pass the locale explicitly: toLocaleUpperCase("tr-TR") in JavaScript, Locale.forLanguageTag("tr-TR") in Java, CultureInfo("tr-TR") in .NET. That turns istanbul into İSTANBUL on every machine, whatever the environment's default locale.

In Python, str.upper() is locale-independent, so Turkish upper-casing needs a letter map or an ICU-based library. The “Code patterns” panel lists the pitfall and the safe option per language.

Limits

The tool uses your browser's ICU data, so results vary by browser and version; without Turkish locale data it warns instead of showing a misleading result. The JavaScript results here were verified on 11 October 2026 with Node v20.15.1 and ICU 74.2, and Python's casing was run too.

The Java, C# and SQL rows were not executed; they rest on official documentation. Java, C# and SQL behaviour depends on the environment and collation settings, so verify it in your own environment. Input is capped at 10,000 lines and 256 KiB, the compare text at 200 characters; nothing you paste is executed.

Sources (checked on 11 October 2026): Java String, .NET string practices, Python str.casefold, PostgreSQL collation, MySQL Unicode sets.

Frequently asked questions

What is the difference between toUpperCase() and toLocaleUpperCase("tr-TR")?

toUpperCase() is locale-independent and turns istanbul into ISTANBUL; toLocaleUpperCase("tr-TR") gives İSTANBUL. Use the latter for Turkish text shown to users.

Why does lower-casing İ give two code units?

Unicode's default mapping lower-cases İ (U+0130) to i plus U+0307 COMBINING DOT ABOVE, so the length grows from 1 to 2. The tr-TR rule gives the single letter i.

What happens when the JVM locale is tr-TR?

Methods that take the locale implicitly change. Per the Java documentation, toUpperCase() without an argument uses the default locale, so on a tr-TR JVM "title".toUpperCase() gives TİTLE. In .NET, calls without a culture use the current culture. Verify it in your environment.

Does the regex i flag match Turkish letters?

No. /istanbul/i does not match İSTANBUL, with or without the u flag. Lower-case both sides with toLocaleLowerCase("tr-TR") or use localeCompare.

Which tool converts text to upper or lower case?

This tool only diagnoses. Use the case converter for Turkish-aware conversion and the slug generator for URLs.

Is there a longer article on this topic?

Yes: The Turkish İ problem: how İ/i and ı/I break software covers it at length with code examples.

Published: · Updated: