Four letters, two pairs

Turkish has two pairs of i letters: dotted i becomes İ in capitals, and dotless ı becomes I. Software written with English rules in mind knows only one pair: i and I. When the two assumptions meet, text breaks silently. istanbul becomes ISTANBUL in one place and İSTANBUL in another; a file name, a user name or a protocol key stops matching the way you expect. Developers often call this class of bug the “Turkey test”.

This article shows, with examples, the layers where the problem appears: case mapping, comparison and search, sorting, slugs and URLs, and identifiers and protocol keywords. I ran the JavaScript and Python examples myself and pasted the outputs as they came (Node v20.15.1 with ICU 74.2, Python 3.12.13; 11 October 2026). I did not run any Java, C#/.NET or SQL code; the statements about those languages rest on the official documentation and go no further than it does.

Characters involved in the I problem
CharacterCode pointUnicode nameTurkish counterpart
iU+0069LATIN SMALL LETTER Icapital is İ
İU+0130LATIN CAPITAL LETTER I WITH DOT ABOVEsmall is i
ıU+0131LATIN SMALL LETTER DOTLESS Icapital is I
IU+0049LATIN CAPITAL LETTER Ismall is ı

A fifth character is involved too: U+0307 COMBINING DOT ABOVE. When İ is lower-cased by the default rule, the result is i plus this dot mark. Because it is invisible, this article and its diagrams always write it as a code point.

Case mapping

Let us start with the simplest conversion. In JavaScript, toUpperCase() and toLowerCase() are locale-independent and apply Unicode's default mapping; toLocaleUpperCase("tr-TR") and toLocaleLowerCase("tr-TR") apply the Turkish rule. The program below converts five words under both rules. The combining dot that appears when İ is lower-cased is invisible, so the output prints it as <U+0307>.

case.mjs: five words, two rules.js
01const cps = (s) => [...s].map((c) => "U+" + c.codePointAt(0).toString(16).toUpperCase().padStart(4, "0")).join(" ");02const vis = (s) => s.replace(/\p{M}/gu, (m) => `<${cps(m)}>`);03 04const words = ["istanbul", "Iğdır", "İzmir", "file", "TITLE"];05console.log("word".padEnd(10), "upper()".padEnd(10), "upper(tr-TR)".padEnd(14), "lower()".padEnd(14), "lower(tr-TR)");06for (const w of words) {07  console.log(08    w.padEnd(10),09    w.toUpperCase().padEnd(10),10    w.toLocaleUpperCase("tr-TR").padEnd(14),11    vis(w.toLowerCase()).padEnd(14),12    w.toLocaleLowerCase("tr-TR"),13  );14}15 16console.log("İ lower()      :", cps("İ".toLowerCase()));17console.log("İ lower(tr-TR) :", cps("İ".toLocaleLowerCase("tr-TR")));18console.log("lengths        :", "İzmir".toLowerCase().length, "İzmir".toLocaleLowerCase("tr-TR").length);
Output of case.mjs (Node v20.15.1).text
01$ node case.mjs02word       upper()    upper(tr-TR)   lower()        lower(tr-TR)03istanbul   ISTANBUL   İSTANBUL       istanbul       istanbul04Iğdır      IĞDIR      IĞDIR          iğdır          ığdır05İzmir      İZMIR      İZMİR          i<U+0307>zmir  izmir06file       FILE       FİLE           file           file07TITLE      TITLE      TITLE          title          tıtle08İ lower()      : U+0069 U+030709İ lower(tr-TR) : U+006910lengths        : 6 5

Three rows of the table repay a closer look. istanbul becomes ISTANBUL under the default rule and İSTANBUL under the Turkish rule; code that compares the two results gets false. İzmir, lower-cased by the default rule, turns into i, the combining dot and zmir, and its length is 6; the Turkish rule gives the five-letter izmir. TITLE becomes tıtle under the Turkish rule, because I goes to dotless ı there. None of this is a bug; each rule is right in its own language. The bug is using the wrong rule in the wrong place.

A four-row comparison. Upper case: under the default rule i and ı give I; under tr-TR i gives İ and ı gives I. Lower case: under the default rule I gives i and İ gives i plus the U+0307 combining dot; under tr-TR I gives ı and İ gives i. Below, three word examples: istanbul becomes ISTANBUL or İSTANBUL in capitals; İzmir in lower case becomes i, U+0307, zmir (6 code units) or the five-letter izmir; TITLE in lower case becomes title or tıtle.
The same input gives different results under the two rules; the most surprising row is lower-casing İ: i plus U+0307 (two code units) by default, the single letter i under tr-TR.

Converting to capitals loses information. Under the default rule both i and ı go to the same I; from ISTANBUL you cannot work out which was meant. Lower-casing is not reversible either: the city Iğdır becomes iğdır under the default rule, although the right form is ığdır. Use the result of a conversion only as a comparison key and keep the original as it is.

The rule is written down in Unicode's SpecialCasing.txt. For İ (U+0130) the file gives the unconditional lower-case mapping 0069 0307, while in the tr and az sections the same letter goes to 0069. For I, the tr and az rule is 0131 (ı) when no combining dot follows. Microsoft's .NET documentation also notes that the same behaviour occurs in the Azerbaijani (az) culture, so the problem is not specific to Turkish.

The result is the same in Python: str.upper(), str.lower() and str.casefold() do not look at the locale; the documentation describes Unicode's default case conversion and folding. The program below asks Python the same questions.

case.py: Python's default mapping and Unicode equivalence.python
01import unicodedata02 03 04def cps(s):05    return " ".join(f"U+{ord(c):04X}" for c in s)06 07 08print("i".upper(), "ı".upper(), "I".lower(), "title".upper())09 10lowered = "İzmir".lower()11print(len(lowered), cps(lowered[:2]), len(unicodedata.normalize("NFC", lowered)))12print("İ".casefold() == "i\N{COMBINING DOT ABOVE}", "ı".casefold() == "ı", "I".casefold())13 14decomposed = unicodedata.normalize("NFD", "İ")15print(cps(decomposed), unicodedata.normalize("NFC", decomposed) == "İ")16print(unicodedata.name("İ"), "|", unicodedata.name("ı"))
Output of case.py (Python 3.12.13).text
01$ python3 case.py02I I i TITLE036 U+0069 U+0307 604True True i05U+0049 U+0307 True06LATIN CAPITAL LETTER I WITH DOT ABOVE | LATIN SMALL LETTER DOTLESS I

The first line is I I i TITLE: i and ı both become I, and I becomes i in lower case. The second line shows that lower-casing İzmir gives 6 code points and that normalisation does not reduce that. On the third line casefold() does the same: i plus the combining dot for İ. The last two lines recall Unicode equivalence: İ decomposes canonically into I plus U+0307 and NFC puts it back together, whereas i plus U+0307 does not combine into a single letter.

Comparison and search

The most common mistake is to test equality without thinking about case. The i flag of JavaScript regular expressions does not know Turkish: it returns false in both directions, and adding the u flag does not change that. The program below tests this and three Turkish-aware alternatives.

compare.mjs: regular expression, lower-casing key, localeCompare and folding.js
01console.log(/İSTANBUL/i.test("istanbul"), /istanbul/i.test("İSTANBUL"));02console.log(/İSTANBUL/iu.test("istanbul"), /istanbul/iu.test("İSTANBUL"));03 04const key = (s) => s.toLocaleLowerCase("tr-TR");05console.log(key("İSTANBUL") === key("istanbul"), key("ISTANBUL") === key("istanbul"));06 07const same = (a, b) => a.localeCompare(b, "tr", { sensitivity: "accent" }) === 0;08console.log(same("istanbul", "İSTANBUL"), same("ığdır", "IĞDIR"), same("ISTANBUL", "İSTANBUL"));09console.log("istanbul".localeCompare("İSTANBUL", "en", { sensitivity: "accent" }));10 11const MAP = { ç: "c", ğ: "g", ı: "i", ö: "o", ş: "s", ü: "u" };12const fold = (s) => s.toLocaleLowerCase("tr-TR").replace(/[çğıöşü]/g, (c) => MAP[c]);13console.log(["istanbul", "ISTANBUL", "İSTANBUL", "Istanbul"].map(fold).join(" "));
Output of compare.mjs (Node v20.15.1).text
01$ node compare.mjs02false false03false false04true false05true true false06-107istanbul istanbul istanbul istanbul

The first two lines are the results of the regular expressions: /İSTANBUL/i.test("istanbul") and /istanbul/i.test("İSTANBUL") both returned false, with the u flag too. On the third line both sides are lower-cased with toLocaleLowerCase("tr-TR"), so İSTANBUL and istanbul match (true); ISTANBUL does not (false), because under the Turkish rule that text becomes ıstanbul. The fourth line shows that the same function, which uses localeCompare with sensitivity: "accent", gives the same results; ığdır and IĞDIR match as well, while ISTANBUL and İSTANBUL do not. On the fifth line the same comparison with "en" returned -1: English rules treat İ as an accented I.

The delicate decision is this: when a user types ISTANBUL, did they mean İstanbul or ıstanbul? A user with an ASCII keyboard or caps lock on usually meant the first. In a search box, therefore, turn both sides into a key that folds Turkish letters to ASCII (fold on the sixth line): istanbul, ISTANBUL, İSTANBUL and Istanbul all land on the same key. The price is that distinct words merge: kır and kir get the same key. That loss is acceptable for search and unacceptable in authentication.

In Python, casefold() is not Turkish-aware. The small helpers below map the i/ı/İ/I pairs first and then call the standard methods.

helpers.py: Turkish upper-casing and lower-casing helpers.python
01TR_UPPER = str.maketrans({"i": "İ", "ı": "I"})02TR_LOWER = str.maketrans({"İ": "i", "I": "ı"})03 04 05def tr_upper(s):06    return s.translate(TR_UPPER).upper()07 08 09def tr_lower(s):10    return s.translate(TR_LOWER).lower()11 12 13print(tr_upper("istanbul"), tr_upper("ığdır"), tr_upper("file"))14print(tr_lower("TITLE"), tr_lower("İzmir"), len(tr_lower("İzmir")))15 16print("İSTANBUL".casefold() == "istanbul".casefold())17print(tr_lower("İSTANBUL") == tr_lower("istanbul"))18print(tr_lower("ISTANBUL") == tr_lower("istanbul"))
Output of helpers.py (Python 3.12.13).text
01$ python3 helpers.py02İSTANBUL IĞDIR FİLE03tıtle izmir 504False05True06False

tr_upper and tr_lower cover these precomposed-character examples; they do not handle contextual mappings such as NFD I + U+0307. Use ICU for full Unicode support. In the False, True, False output, casefold() does not equate the two spellings, tr_lower does, and ISTANBUL stays distinct.

Sorting: collation

Sorting is one level above comparison and breaks in the same place. Without a comparison function, Array.prototype.sort() orders strings by UTF-16 code unit; ç, ş, ü, İ and ı lie outside ASCII, so they fall to the end of the list. The Turkish alphabet, on the other hand, runs a b c ç d e f g ğ h ı i j k l m n o ö p r s ş t u ü v y z (29 letters).

sort.mjs: code unit order, English and Turkish collators.js
01const letters = ["ı", "i", "İ", "I", "ç", "c", "z", "ş", "s", "ü", "u"];02console.log("sort()           :", [...letters].sort().join(" "));03console.log('Collator("en")   :', [...letters].sort(new Intl.Collator("en").compare).join(" "));04console.log('Collator("tr")   :', [...letters].sort(new Intl.Collator("tr").compare).join(" "));05 06const words = ["şeker", "sefer", "Işık", "ılık", "iş", "İzmir", "çay", "cam"];07console.log("sort()           :", [...words].sort().join(" "));08console.log('Collator("tr")   :', [...words].sort(new Intl.Collator("tr").compare).join(" "));
Output of sort.mjs (Node v20.15.1, ICU 74.2).text
01$ node sort.mjs02sort()           : I c i s u z ç ü İ ı ş03Collator("en")   : c ç i I İ ı s ş u ü z04Collator("tr")   : c ç ı I i İ s ş u ü z05sort()           : Işık cam iş sefer çay İzmir ılık şeker06Collator("tr")   : cam çay ılık Işık iş İzmir sefer şeker
Input: ı i İ I ç c z ş s ü u. The default Array.prototype.sort() order is I c i s u z ç ü İ ı ş; the non-ASCII letters (ç ü İ ı ş) fall to the end, and each letter has its code point written below. The Intl.Collator("tr") order is c ç ı I i İ s ş u ü z; ı and I, and i and İ, stand together as the small and capital forms of one letter, and each letter has its position in the Turkish alphabet written below.
Code unit order throws the non-ASCII letters to the end of the list; Intl.Collator("tr") follows the Turkish alphabet. The results were taken with Node v20.15.1 and ICU 74.2 on 11 October 2026.

The first line is the default order: I c i s u z ç ü İ ı ş. The second line is the Intl.Collator("en") result; it is better than code unit order but not Turkish (c ç i I İ ı s ş u ü z). The third line is the Intl.Collator("tr") result: c ç ı I i İ s ş u ü z. In this order ı and I stand together as the small and capital forms of one letter; so do i and İ. With words the effect is more concrete: in Turkish order ılık and Işık come before iş and İzmir, while in the default order Işık goes first and ılık lands near the end.

In Python, sorted() gives the same code point order. The toy key below knows only the 29 letters: it applies the letter order and puts a lower-case letter before its capital. It does not sort digits, punctuation or accented letters such as â correctly; do not use it in a real application, I wrote it only to show the logic of the order.

sort.py: a toy sort key based on the Turkish alphabet.python
01ALPHABET = "abcçdefgğhıijklmnoöprsştuüvyz"02RANK = {letter: n for n, letter in enumerate(ALPHABET)}03 04TR_LOWER = str.maketrans({"İ": "i", "I": "ı"})05 06 07def tr_key(word):08    base = [RANK.get(c, len(ALPHABET) + ord(c)) for c in word.translate(TR_LOWER).lower()]09    return base, [c.isupper() for c in word]10 11 12letters = ["ı", "i", "İ", "I", "ç", "c", "z", "ş", "s", "ü", "u"]13print(len(ALPHABET))14print("sorted()  :", " ".join(sorted(letters)))15print("tr_key    :", " ".join(sorted(letters, key=tr_key)))16 17words = ["şeker", "sefer", "Işık", "ılık", "iş", "İzmir", "çay", "cam"]18print("sorted()  :", " ".join(sorted(words)))19print("tr_key    :", " ".join(sorted(words, key=tr_key)))
Output of sort.py (Python 3.12.13).text
01$ python3 sort.py022903sorted()  : I c i s u z ç ü İ ı ş04tr_key    : c ç ı I i İ s ş u ü z05sorted()  : Işık cam iş sefer çay İzmir ılık şeker06tr_key    : cam çay ılık Işık iş İzmir sefer şeker

You can also use the operating system's ordering through locale.strxfrm, but the result varies from library to library. The program below sorts the same letters under two locales.

locale_check.py: the system locale and locale.strxfrm.python
01import locale02import sys03 04try:05    locale.setlocale(locale.LC_ALL, "")06except locale.Error:07    sys.exit("locale not installed")08 09print(locale.setlocale(locale.LC_ALL))10print("title".upper(), "ı".upper(), "İ".lower() == "i\N{COMBINING DOT ABOVE}")11letters = ["ı", "i", "İ", "I", "ç", "c", "z", "ş", "s", "ü", "u"]12print("strxfrm :", " ".join(sorted(letters, key=locale.strxfrm)))
Output of locale_check.py under two locales (Python 3.12.13).text
01$ LC_ALL=en_US.UTF-8 python3 locale_check.py02en_US.UTF-803TITLE I True04strxfrm : c ç i I İ ı s ş u ü z05$ LC_ALL=tr_TR.UTF-8 python3 locale_check.py06tr_TR.UTF-807TITLE I True08strxfrm : c ç I ı İ i s ş u ü z

Two things stand out. The results of upper() and lower() did not change under tr_TR.UTF-8 either: title is still TITLE, ı is still I. The sort order, however, changed with the environment, and the Turkish order of the glibc on this machine (I ı İ i) differed from ICU's Intl.Collator("tr") order (ı I i İ). “Turkish sort order” is not one result but a decision of the library you use.

In databases this decision is made through the collation. According to the PostgreSQL documentation, a collation does not only determine sort order: case functions such as lower, upper and initcap and pattern-matching operators are affected by it too; the documentation describes the providers libc and icu. The MySQL documentation says language-specific collations are UCA-based and carry a language name or locale code in their names; for Turkish the specifier is tr or turkish, and utf8mb4_turkish_ci is such a name. SQL Server's table of default collations gives Turkish_CI_AS for Turkish (Türkiye); CI means case-insensitive and AS accent-sensitive. I did not run any of these three systems; test the result on your own database.

Slugs, URLs and search engines

The most common way to build a slug has three steps: lower-case the title, drop everything outside ASCII, join the rest with hyphens. For Turkish titles this method breaks silently. The program below compares it with one that first maps Turkish letters to their ASCII counterparts; it also contains two URL observations.

slug.mjs: naive and Turkish-aware slugs, percent-encoding, host names.js
01const naive = (s) => s.toLowerCase().replace(/[^a-z0-9]+/g, "-").replace(/^-|-$/g, "");02 03const MAP = { ç: "c", ğ: "g", ı: "i", ö: "o", ş: "s", ü: "u" };04const fold = (s) => s.toLocaleLowerCase("tr-TR").replace(/[çğıöşü]/g, (c) => MAP[c]);05const slug = (s) => fold(s).replace(/[^a-z0-9]+/g, "-").replace(/^-|-$/g, "");06 07for (const title of ["İstanbul Işık", "Çalışma Ağacı", "Iğdır"]) {08  console.log(title.padEnd(16), "naive:", naive(title).padEnd(16), "tr:", slug(title));09}10 11console.log(encodeURIComponent("İstanbul"), encodeURIComponent("ığdır"));12for (const url of ["https://İSTANBUL.example/", "https://ISTANBUL.example/"]) {13  console.log(url, "->", new URL(url).hostname);14}
Output of slug.mjs (Node v20.15.1).text
01$ node slug.mjs02İstanbul Işık    naive: i-stanbul-i-k    tr: istanbul-isik03Çalışma Ağacı    naive: al-ma-a-ac       tr: calisma-agaci04Iğdır            naive: i-d-r            tr: igdir05%C4%B0stanbul %C4%B1%C4%9Fd%C4%B1r06https://İSTANBUL.example/ -> xn--istanbul-o0e.example07https://ISTANBUL.example/ -> istanbul.example

The naive method turns İstanbul Işık into i-stanbul-i-k: lower-casing İ leaves i plus a combining dot, the dot is not ASCII so it becomes a hyphen and splits the word in the middle; ş and ı also turn into separators and cut Işık apart. Çalışma Ağacı becomes al-ma-a-ac and Iğdır becomes i-d-r. The Turkish-aware method turns all three into readable addresses: istanbul-isik, calisma-agaci, igdir. When I ran the slugify function used by the site's slug generator locally on the same İstanbul Işık input, I also got istanbul-isik.

There are two separate topics in URLs. The first is percent-encoding: encodeURIComponent produces %C4%B0stanbul for İstanbul and %C4%B1%C4%9Fd%C4%B1r for ığdır. Such addresses are valid but unreadable for people. The second is host names: Node's URL class returned xn--istanbul-o0e.example as the host name of https://İSTANBUL.example/, while for https://ISTANBUL.example/ it returned istanbul.example. So a host name containing a capital İ did not end up as the same host name as its lower-case ASCII version. This is not a browser guarantee, only an observation from this run; but it is a good reason to pass host names that come from users through a standard parser such as URL before comparing them.

I have no measured knowledge of how search engines treat İ/ı, so I make no claim about it. The safe options are on your own side: build the permanent address as ASCII and lower case, set up a redirect from the old address to the new one if you change it later, and build the on-site search key (fold in section 3) with the same logic as the slug. If you want a ready-made tool, the slug generator converts Turkish letters to ASCII; for converting text with the Turkish rule there is the case converter.

Identifiers and protocol keywords: locale-independent operations

Machine-owned strings (HTTP header names, URL schemes, configuration keys, file extensions, HTML tags, identifiers) must not change with the user's language. The Java documentation says so explicitly: toLowerCase() and toUpperCase() without an argument are locale-sensitive and may give unexpected results for strings such as identifiers, protocol keys and HTML tags. The documentation states that "title".toUpperCase() returns "TİTLE" in a Turkish locale and says to use Locale.ROOT.

Microsoft's .NET documentation shows the security side: a call such as IsFileURI("file:") can return true under the U.S. English culture and false under the Turkish culture, so a check that blocks addresses beginning with FILE: could be bypassed on Turkish systems. For comparisons that need no linguistic knowledge, the documentation recommends StringComparison.Ordinal or OrdinalIgnoreCase.

In JavaScript, toLowerCase() and toUpperCase() are already locale-independent; the danger is code written for Turkish leaking into identifier comparison. The program below shows exactly this mistake: the same FILE: string becomes file: with toLowerCase() and fıle: under the tr-TR rule.

identifiers.mjs: a protocol scheme and ASCII-limited lower-casing.js
01const scheme = "FILE:";02console.log(scheme.toLowerCase() === "file:");03console.log(scheme.toLocaleLowerCase("tr-TR") === "file:", scheme.toLocaleLowerCase("tr-TR"));04 05const asciiLower = (s) => s.replace(/[A-Z]+/g, (m) => m.toLowerCase());06console.log(asciiLower("ID"), asciiLower("İD"), asciiLower("ID") === "id", asciiLower("İD") === "id");
Output of identifiers.mjs (Node v20.15.1).text
01$ node identifiers.mjs02true03false fıle:04id İd true false

The second line gives false fıle:. The last line is an alternative policy for identifiers: asciiLower lowers only the letters A-Z and consults no locale; İD therefore stays İd and does not match id. That is deliberate: for identifiers you decide which characters are accepted. The sturdiest route is often to restrict the identifier at registration to a constrained ASCII set (for example [a-z0-9_-]).

  • Machine-owned identifiers and keys: a locale-independent or ordinal operation; where possible, limit the accepted character set to ASCII.
  • Turkish text shown to users: write the locale explicitly in the code ("tr-TR"); do not inherit it from the environment.
  • Comparison key: keep the original value and fold only the key.
  • At system boundaries (HTTP, file system, database), write down which side applies which rule.

A checklist, language by language

The table below summarises the pitfall and two safe routes for each language. Mind the last column: I ran the JavaScript and Python rows with the programs in this article; the Java, C# and SQL rows rest only on the official documentation and were not run.

Dotted and dotless I pitfalls per language, with safe options
LanguagePitfallFor identifiers and keysFor Turkish textVerification
JavaScripttoUpperCase()/toLowerCase() do not know Turkish; the /i regular expression does not match İ and ı; localeCompare without an argument followed the environmenttoLowerCase(), identifiers limited to ASCIItoLocaleUpperCase("tr-TR"), Intl.Collator("tr"), localeCompare(…, "tr", {sensitivity})Run: Node v20.15.1, ICU 74.2
JavatoUpperCase()/toLowerCase() without an argument use the default locale; in a Turkish locale "title".toUpperCase() gives "TİTLE". equalsIgnoreCase ignores the locale but does not know TurkishtoUpperCase(Locale.ROOT), toLowerCase(Locale.ROOT)Locale.forLanguageTag("tr-TR"); Collator for locale-sensitive comparisonNot run: Java SE 21 documentation
C#/.NETToUpper()/ToLower() use the current culture; Compare defaults to culture-sensitive comparison, Equals to ordinal comparison. A culture-sensitive "file:" check can fail under TurkishToUpperInvariant(), StringComparison.Ordinal/OrdinalIgnoreCaseToUpper(culture) with new CultureInfo("tr-TR"); comparisons that name the culture explicitlyNot run: Microsoft Learn
Pythonupper(), lower() and casefold() apply the Unicode default: "i".upper() gives I, "İ".lower() gives two characters; sorted() is code point orderlower()/casefold() and identifiers limited to ASCIIA prior i/ı/İ/I mapping (tr_upper, tr_lower); an ICU-based library or locale.strxfrm for sortingRun: Python 3.12.13
SQLThe result depends on the collation of the column or database; in PostgreSQL the provider (libc or icu) matters as wellFor identifier columns, a language-independent collation and a single-form value (e.g. lower-case ASCII)Turkish collations such as MySQL utf8mb4_turkish_ci or SQL Server Turkish_CI_AS; a Turkish ICU or libc collation in PostgreSQLNot run: PostgreSQL, MySQL and SQL Server documentation

The table is a decision summary, not a measurement result. Java, C# and SQL behaviour depends on the environment and the collation settings; verify it in your own environment.

Java: a documented pattern, not run.java
01String key = header.toUpperCase(Locale.ROOT);02String label = title.toUpperCase(Locale.forLanguageTag("tr-TR"));
C#: a documented pattern, not run.csharp
01string key = header.ToUpperInvariant();02bool same = string.Equals(a, b, StringComparison.OrdinalIgnoreCase);03string label = title.ToUpper(new CultureInfo("tr-TR"));

Test strategy

Most Turkey-test bugs blow up not on the developer's machine but in a customer's or a CI server's Turkish-locale environment. Turn that around: supply Turkish yourself, both as test input and as test environment.

  • Keep a fixed set of Turkish test strings: istanbul, ISTANBUL, İSTANBUL, Iğdır, ığdır, IĞDIR, İzmir, TITLE, file, FILE:.
  • Try the same set in NFC and NFD form: İ can arrive as a single code point or as I plus U+0307.
  • Write the expectation separately for each operation: upper case, lower case, length, equality, sorting, slug and regular expression.
  • Label separately the places that name the locale explicitly and those that read it from the environment.
  • Run the tests in at least two environments: an English and a Turkish (tr-TR) locale.

The small check program below shows the idea. The first four checks test calls that name the locale explicitly or use none at all. The last two stand for code that trusts the environment's locale and deliberately assume en-US. I ran the program in two environments.

ci-check.mjs: checks with an explicit locale and with the ambient one.js
01console.log("default locale:", Intl.DateTimeFormat().resolvedOptions().locale);02console.log("title".toLocaleUpperCase(), "(toLocaleUpperCase, no argument)");03 04const letters = ["ı", "i", "İ", "I"];05const checks = [06  ['"TITLE".toLowerCase() is "title"', "TITLE".toLowerCase() === "title"],07  ['"file".toUpperCase() is "FILE"', "file".toUpperCase() === "FILE"],08  ['"istanbul" to tr-TR upper is "İSTANBUL"', "istanbul".toLocaleUpperCase("tr-TR") === "İSTANBUL"],09  ["Collator('tr') order is ı I i İ", [...letters].sort(new Intl.Collator("tr").compare).join(" ") === "ı I i İ"],10  ["ambient sort order is i I İ ı", [...letters].sort(new Intl.Collator().compare).join(" ") === "i I İ ı"],11  ["ambient accent-equal istanbul/İSTANBUL is false", "istanbul".localeCompare("İSTANBUL", undefined, { sensitivity: "accent" }) !== 0],12];13for (const [name, ok] of checks) console.log(ok ? "ok  " : "FAIL", name);14process.exitCode = checks.every(([, ok]) => ok) ? 0 : 1;
Output of ci-check.mjs in two environments (Node v20.15.1).text
01$ LC_ALL=en_US.UTF-8 node ci-check.mjs02default locale: en-US03TITLE (toLocaleUpperCase, no argument)04ok   "TITLE".toLowerCase() is "title"05ok   "file".toUpperCase() is "FILE"06ok   "istanbul" to tr-TR upper is "İSTANBUL"07ok   Collator('tr') order is ı I i İ08ok   ambient sort order is i I İ ı09ok   ambient accent-equal istanbul/İSTANBUL is false10$ LC_ALL=tr_TR.UTF-8 node ci-check.mjs11default locale: tr-TR12TITLE (toLocaleUpperCase, no argument)13ok   "TITLE".toLowerCase() is "title"14ok   "file".toUpperCase() is "FILE"15ok   "istanbul" to tr-TR upper is "İSTANBUL"16ok   Collator('tr') order is ı I i İ17FAIL ambient sort order is i I İ ı18FAIL ambient accent-equal istanbul/İSTANBUL is false19# exit 1

In the English environment all six checks gave ok. With LC_ALL=tr_TR.UTF-8 the first four checks still passed; the last two, which rely on the environment's locale, gave FAIL and the exit code was 1: Intl.Collator() and localeCompare without an argument saw that the environment was tr-TR. One more detail: "title".toLocaleUpperCase() without an argument returned TITLE in both environments, so the behaviour of “following the environment” can vary with the API and the engine. That is a reason to assume neither that the environment is Turkish nor that it is not: write the language tag wherever you want Turkish. The FAIL lines are not bugs but proof that the code depends on the environment.

In other runtimes the same idea is set up differently; these were not run. The example in the .NET documentation changes the culture with Thread.CurrentThread.CurrentCulture = new CultureInfo("tr-TR"). In Java you can switch the default locale to Turkish at the start of a test with Locale.setDefault and restore the old value afterwards. In Python, locale.setlocale selects a Turkish locale installed on the system (like locale_check.py above). In database tests, name the column's collation explicitly and query with rows that contain both a capital İ and a dotless ı.

Try it with your own text

To try the examples in this article with your own text, you can use the site's Turkish I test tool (the link at the end of the article goes there too). The tool answers the same question under two rules side by side: the root (default) locale and tr-TR. The calculation runs in your browser and none of the code you paste is executed.

A short guide: paste text into the upper field, one word per line, or use the Load example button; if you like, type a compare text into the second field. For every line the tool shows upper case, lower case, length after lower-casing, NFC/NFD length, equality, regular expression and slug results in two columns. Differing results are flagged with the Differs label, and U+0307 is written with its name COMBINING DOT ABOVE. At the bottom, the two sort orders (Array.prototype.sort() and Intl.Collator("tr")) are compared.

The tool is for diagnosis only. To convert text, use the case converter; to build addresses, use the slug generator. The tool's results depend on your browser's ICU data and can vary with the browser or its version; verify the Java, C# and SQL side separately in your own environment.

In short: the Turkish İ problem is not a bug but two correct rules meeting in the wrong place. Write the rule explicitly in the code, keep identifiers apart from the locale, and put Turkish into your test environment.

Official sources and further reading

The language behaviour and documentation statements in this article were checked against these sources on 11 October 2026. The JavaScript and Python examples were also run; the Java, C#/.NET and SQL statements rest only on these documents:

  1. Oracle Java SE 21: java.lang.String (toUpperCase, toLowerCase, equalsIgnoreCase)
  2. Microsoft Learn: Best practices for comparing strings in .NET
  3. Microsoft Learn: String.ToUpper method
  4. Python documentation: str.lower, str.upper and str.casefold
  5. PostgreSQL documentation: Collation Support
  6. MySQL 8.4 Reference Manual: Unicode Character Sets
  7. Microsoft Learn: Collation and Unicode support (SQL Server)
  8. MDN: String.prototype.toLocaleUpperCase()
  9. Unicode Character Database: SpecialCasing.txt

The source pages are living documents and can change. The JavaScript results were taken with Node v20.15.1 (ICU 74.2, Unicode 15.1) and the Python results with Python 3.12.13 (unicodedata 15.0.0) on a Linux machine; other browser, engine or library versions may give different results. The Java, C#/.NET and SQL statements were not run. “Turkey test” is a commonly used name for this class of bug; it is not attributed to any source. The links to the site's tools are in sections 5 and 9; the Turkish I test tool can also be reached from the button at the end of the article.

✳

Four letters, two pairs: the problem is code that does not say which rule it uses.

Open the Turkish I test ↗