Four letters, two pairs
Turkish has two pairs of i letters: dotted i becomes İ in capitals, and dotless ı becomes I. Software written with English rules in mind knows only one pair: i and I. When the two assumptions meet, text breaks silently. istanbul becomes ISTANBUL in one place and İSTANBUL in another; a file name, a user name or a protocol key stops matching the way you expect. Developers often call this class of bug the “Turkey test”.
This article shows, with examples, the layers where the problem appears: case mapping, comparison and search, sorting, slugs and URLs, and identifiers and protocol keywords. I ran the JavaScript and Python examples myself and pasted the outputs as they came (Node v20.15.1 with ICU 74.2, Python 3.12.13; 11 October 2026). I did not run any Java, C#/.NET or SQL code; the statements about those languages rest on the official documentation and go no further than it does.
| Character | Code point | Unicode name | Turkish counterpart |
|---|---|---|---|
i | U+0069 | LATIN SMALL LETTER I | capital is İ |
İ | U+0130 | LATIN CAPITAL LETTER I WITH DOT ABOVE | small is i |
ı | U+0131 | LATIN SMALL LETTER DOTLESS I | capital is I |
I | U+0049 | LATIN CAPITAL LETTER I | small is ı |
A fifth character is involved too: U+0307 COMBINING DOT ABOVE. When İ is lower-cased by the default rule, the result is i plus this dot mark. Because it is invisible, this article and its diagrams always write it as a code point.
Case mapping
Let us start with the simplest conversion. In JavaScript, toUpperCase() and toLowerCase() are locale-independent and apply Unicode's default mapping; toLocaleUpperCase("tr-TR") and toLocaleLowerCase("tr-TR") apply the Turkish rule. The program below converts five words under both rules. The combining dot that appears when İ is lower-cased is invisible, so the output prints it as <U+0307>.
01const cps = (s) => [...s].map((c) => "U+" + c.codePointAt(0).toString(16).toUpperCase().padStart(4, "0")).join(" ");02const vis = (s) => s.replace(/\p{M}/gu, (m) => `<${cps(m)}>`);03 04const words = ["istanbul", "Iğdır", "İzmir", "file", "TITLE"];05console.log("word".padEnd(10), "upper()".padEnd(10), "upper(tr-TR)".padEnd(14), "lower()".padEnd(14), "lower(tr-TR)");06for (const w of words) {07 console.log(08 w.padEnd(10),09 w.toUpperCase().padEnd(10),10 w.toLocaleUpperCase("tr-TR").padEnd(14),11 vis(w.toLowerCase()).padEnd(14),12 w.toLocaleLowerCase("tr-TR"),13 );14}15 16console.log("İ lower() :", cps("İ".toLowerCase()));17console.log("İ lower(tr-TR) :", cps("İ".toLocaleLowerCase("tr-TR")));18console.log("lengths :", "İzmir".toLowerCase().length, "İzmir".toLocaleLowerCase("tr-TR").length);01$ node case.mjs02word upper() upper(tr-TR) lower() lower(tr-TR)03istanbul ISTANBUL İSTANBUL istanbul istanbul04Iğdır IĞDIR IĞDIR iğdır ığdır05İzmir İZMIR İZMİR i<U+0307>zmir izmir06file FILE FİLE file file07TITLE TITLE TITLE title tıtle08İ lower() : U+0069 U+030709İ lower(tr-TR) : U+006910lengths : 6 5Three rows of the table repay a closer look. istanbul becomes ISTANBUL under the default rule and İSTANBUL under the Turkish rule; code that compares the two results gets false. İzmir, lower-cased by the default rule, turns into i, the combining dot and zmir, and its length is 6; the Turkish rule gives the five-letter izmir. TITLE becomes tıtle under the Turkish rule, because I goes to dotless ı there. None of this is a bug; each rule is right in its own language. The bug is using the wrong rule in the wrong place.
Converting to capitals loses information. Under the default rule both i and ı go to the same I; from ISTANBUL you cannot work out which was meant. Lower-casing is not reversible either: the city Iğdır becomes iğdır under the default rule, although the right form is ığdır. Use the result of a conversion only as a comparison key and keep the original as it is.
The rule is written down in Unicode's SpecialCasing.txt. For İ (U+0130) the file gives the unconditional lower-case mapping 0069 0307, while in the tr and az sections the same letter goes to 0069. For I, the tr and az rule is 0131 (ı) when no combining dot follows. Microsoft's .NET documentation also notes that the same behaviour occurs in the Azerbaijani (az) culture, so the problem is not specific to Turkish.
The result is the same in Python: str.upper(), str.lower() and str.casefold() do not look at the locale; the documentation describes Unicode's default case conversion and folding. The program below asks Python the same questions.
01import unicodedata02 03 04def cps(s):05 return " ".join(f"U+{ord(c):04X}" for c in s)06 07 08print("i".upper(), "ı".upper(), "I".lower(), "title".upper())09 10lowered = "İzmir".lower()11print(len(lowered), cps(lowered[:2]), len(unicodedata.normalize("NFC", lowered)))12print("İ".casefold() == "i\N{COMBINING DOT ABOVE}", "ı".casefold() == "ı", "I".casefold())13 14decomposed = unicodedata.normalize("NFD", "İ")15print(cps(decomposed), unicodedata.normalize("NFC", decomposed) == "İ")16print(unicodedata.name("İ"), "|", unicodedata.name("ı"))01$ python3 case.py02I I i TITLE036 U+0069 U+0307 604True True i05U+0049 U+0307 True06LATIN CAPITAL LETTER I WITH DOT ABOVE | LATIN SMALL LETTER DOTLESS IThe first line is I I i TITLE: i and ı both become I, and I becomes i in lower case. The second line shows that lower-casing İzmir gives 6 code points and that normalisation does not reduce that. On the third line casefold() does the same: i plus the combining dot for İ. The last two lines recall Unicode equivalence: İ decomposes canonically into I plus U+0307 and NFC puts it back together, whereas i plus U+0307 does not combine into a single letter.
Comparison and search
The most common mistake is to test equality without thinking about case. The i flag of JavaScript regular expressions does not know Turkish: it returns false in both directions, and adding the u flag does not change that. The program below tests this and three Turkish-aware alternatives.
01console.log(/İSTANBUL/i.test("istanbul"), /istanbul/i.test("İSTANBUL"));02console.log(/İSTANBUL/iu.test("istanbul"), /istanbul/iu.test("İSTANBUL"));03 04const key = (s) => s.toLocaleLowerCase("tr-TR");05console.log(key("İSTANBUL") === key("istanbul"), key("ISTANBUL") === key("istanbul"));06 07const same = (a, b) => a.localeCompare(b, "tr", { sensitivity: "accent" }) === 0;08console.log(same("istanbul", "İSTANBUL"), same("ığdır", "IĞDIR"), same("ISTANBUL", "İSTANBUL"));09console.log("istanbul".localeCompare("İSTANBUL", "en", { sensitivity: "accent" }));10 11const MAP = { ç: "c", ğ: "g", ı: "i", ö: "o", ş: "s", ü: "u" };12const fold = (s) => s.toLocaleLowerCase("tr-TR").replace(/[çğıöşü]/g, (c) => MAP[c]);13console.log(["istanbul", "ISTANBUL", "İSTANBUL", "Istanbul"].map(fold).join(" "));01$ node compare.mjs02false false03false false04true false05true true false06-107istanbul istanbul istanbul istanbulThe first two lines are the results of the regular expressions: /İSTANBUL/i.test("istanbul") and /istanbul/i.test("İSTANBUL") both returned false, with the u flag too. On the third line both sides are lower-cased with toLocaleLowerCase("tr-TR"), so İSTANBUL and istanbul match (true); ISTANBUL does not (false), because under the Turkish rule that text becomes ıstanbul. The fourth line shows that the same function, which uses localeCompare with sensitivity: "accent", gives the same results; ığdır and IĞDIR match as well, while ISTANBUL and İSTANBUL do not. On the fifth line the same comparison with "en" returned -1: English rules treat İ as an accented I.
The delicate decision is this: when a user types ISTANBUL, did they mean İstanbul or ıstanbul? A user with an ASCII keyboard or caps lock on usually meant the first. In a search box, therefore, turn both sides into a key that folds Turkish letters to ASCII (fold on the sixth line): istanbul, ISTANBUL, İSTANBUL and Istanbul all land on the same key. The price is that distinct words merge: kır and kir get the same key. That loss is acceptable for search and unacceptable in authentication.
In Python, casefold() is not Turkish-aware. The small helpers below map the i/ı/İ/I pairs first and then call the standard methods.
01TR_UPPER = str.maketrans({"i": "İ", "ı": "I"})02TR_LOWER = str.maketrans({"İ": "i", "I": "ı"})03 04 05def tr_upper(s):06 return s.translate(TR_UPPER).upper()07 08 09def tr_lower(s):10 return s.translate(TR_LOWER).lower()11 12 13print(tr_upper("istanbul"), tr_upper("ığdır"), tr_upper("file"))14print(tr_lower("TITLE"), tr_lower("İzmir"), len(tr_lower("İzmir")))15 16print("İSTANBUL".casefold() == "istanbul".casefold())17print(tr_lower("İSTANBUL") == tr_lower("istanbul"))18print(tr_lower("ISTANBUL") == tr_lower("istanbul"))01$ python3 helpers.py02İSTANBUL IĞDIR FİLE03tıtle izmir 504False05True06Falsetr_upper and tr_lower cover these precomposed-character examples; they do not handle contextual mappings such as NFD I + U+0307. Use ICU for full Unicode support. In the False, True, False output, casefold() does not equate the two spellings, tr_lower does, and ISTANBUL stays distinct.
Sorting: collation
Sorting is one level above comparison and breaks in the same place. Without a comparison function, Array.prototype.sort() orders strings by UTF-16 code unit; ç, ş, ü, İ and ı lie outside ASCII, so they fall to the end of the list. The Turkish alphabet, on the other hand, runs a b c ç d e f g ğ h ı i j k l m n o ö p r s ş t u ü v y z (29 letters).
01const letters = ["ı", "i", "İ", "I", "ç", "c", "z", "ş", "s", "ü", "u"];02console.log("sort() :", [...letters].sort().join(" "));03console.log('Collator("en") :', [...letters].sort(new Intl.Collator("en").compare).join(" "));04console.log('Collator("tr") :', [...letters].sort(new Intl.Collator("tr").compare).join(" "));05 06const words = ["şeker", "sefer", "Işık", "ılık", "iş", "İzmir", "çay", "cam"];07console.log("sort() :", [...words].sort().join(" "));08console.log('Collator("tr") :', [...words].sort(new Intl.Collator("tr").compare).join(" "));01$ node sort.mjs02sort() : I c i s u z ç ü İ ı ş03Collator("en") : c ç i I İ ı s ş u ü z04Collator("tr") : c ç ı I i İ s ş u ü z05sort() : Işık cam iş sefer çay İzmir ılık şeker06Collator("tr") : cam çay ılık Işık iş İzmir sefer şekerThe first line is the default order: I c i s u z ç ü İ ı ş. The second line is the Intl.Collator("en") result; it is better than code unit order but not Turkish (c ç i I İ ı s ş u ü z). The third line is the Intl.Collator("tr") result: c ç ı I i İ s ş u ü z. In this order ı and I stand together as the small and capital forms of one letter; so do i and İ. With words the effect is more concrete: in Turkish order ılık and Işık come before iş and İzmir, while in the default order Işık goes first and ılık lands near the end.
In Python, sorted() gives the same code point order. The toy key below knows only the 29 letters: it applies the letter order and puts a lower-case letter before its capital. It does not sort digits, punctuation or accented letters such as â correctly; do not use it in a real application, I wrote it only to show the logic of the order.
01ALPHABET = "abcçdefgğhıijklmnoöprsştuüvyz"02RANK = {letter: n for n, letter in enumerate(ALPHABET)}03 04TR_LOWER = str.maketrans({"İ": "i", "I": "ı"})05 06 07def tr_key(word):08 base = [RANK.get(c, len(ALPHABET) + ord(c)) for c in word.translate(TR_LOWER).lower()]09 return base, [c.isupper() for c in word]10 11 12letters = ["ı", "i", "İ", "I", "ç", "c", "z", "ş", "s", "ü", "u"]13print(len(ALPHABET))14print("sorted() :", " ".join(sorted(letters)))15print("tr_key :", " ".join(sorted(letters, key=tr_key)))16 17words = ["şeker", "sefer", "Işık", "ılık", "iş", "İzmir", "çay", "cam"]18print("sorted() :", " ".join(sorted(words)))19print("tr_key :", " ".join(sorted(words, key=tr_key)))01$ python3 sort.py022903sorted() : I c i s u z ç ü İ ı ş04tr_key : c ç ı I i İ s ş u ü z05sorted() : Işık cam iş sefer çay İzmir ılık şeker06tr_key : cam çay ılık Işık iş İzmir sefer şekerYou can also use the operating system's ordering through locale.strxfrm, but the result varies from library to library. The program below sorts the same letters under two locales.
01import locale02import sys03 04try:05 locale.setlocale(locale.LC_ALL, "")06except locale.Error:07 sys.exit("locale not installed")08 09print(locale.setlocale(locale.LC_ALL))10print("title".upper(), "ı".upper(), "İ".lower() == "i\N{COMBINING DOT ABOVE}")11letters = ["ı", "i", "İ", "I", "ç", "c", "z", "ş", "s", "ü", "u"]12print("strxfrm :", " ".join(sorted(letters, key=locale.strxfrm)))01$ LC_ALL=en_US.UTF-8 python3 locale_check.py02en_US.UTF-803TITLE I True04strxfrm : c ç i I İ ı s ş u ü z05$ LC_ALL=tr_TR.UTF-8 python3 locale_check.py06tr_TR.UTF-807TITLE I True08strxfrm : c ç I ı İ i s ş u ü zTwo things stand out. The results of upper() and lower() did not change under tr_TR.UTF-8 either: title is still TITLE, ı is still I. The sort order, however, changed with the environment, and the Turkish order of the glibc on this machine (I ı İ i) differed from ICU's Intl.Collator("tr") order (ı I i İ). “Turkish sort order” is not one result but a decision of the library you use.
In databases this decision is made through the collation. According to the PostgreSQL documentation, a collation does not only determine sort order: case functions such as lower, upper and initcap and pattern-matching operators are affected by it too; the documentation describes the providers libc and icu. The MySQL documentation says language-specific collations are UCA-based and carry a language name or locale code in their names; for Turkish the specifier is tr or turkish, and utf8mb4_turkish_ci is such a name. SQL Server's table of default collations gives Turkish_CI_AS for Turkish (Türkiye); CI means case-insensitive and AS accent-sensitive. I did not run any of these three systems; test the result on your own database.
Slugs, URLs and search engines
The most common way to build a slug has three steps: lower-case the title, drop everything outside ASCII, join the rest with hyphens. For Turkish titles this method breaks silently. The program below compares it with one that first maps Turkish letters to their ASCII counterparts; it also contains two URL observations.
01const naive = (s) => s.toLowerCase().replace(/[^a-z0-9]+/g, "-").replace(/^-|-$/g, "");02 03const MAP = { ç: "c", ğ: "g", ı: "i", ö: "o", ş: "s", ü: "u" };04const fold = (s) => s.toLocaleLowerCase("tr-TR").replace(/[çğıöşü]/g, (c) => MAP[c]);05const slug = (s) => fold(s).replace(/[^a-z0-9]+/g, "-").replace(/^-|-$/g, "");06 07for (const title of ["İstanbul Işık", "Çalışma Ağacı", "Iğdır"]) {08 console.log(title.padEnd(16), "naive:", naive(title).padEnd(16), "tr:", slug(title));09}10 11console.log(encodeURIComponent("İstanbul"), encodeURIComponent("ığdır"));12for (const url of ["https://İSTANBUL.example/", "https://ISTANBUL.example/"]) {13 console.log(url, "->", new URL(url).hostname);14}01$ node slug.mjs02İstanbul Işık naive: i-stanbul-i-k tr: istanbul-isik03Çalışma Ağacı naive: al-ma-a-ac tr: calisma-agaci04Iğdır naive: i-d-r tr: igdir05%C4%B0stanbul %C4%B1%C4%9Fd%C4%B1r06https://İSTANBUL.example/ -> xn--istanbul-o0e.example07https://ISTANBUL.example/ -> istanbul.exampleThe naive method turns İstanbul Işık into i-stanbul-i-k: lower-casing İ leaves i plus a combining dot, the dot is not ASCII so it becomes a hyphen and splits the word in the middle; ş and ı also turn into separators and cut Işık apart. Çalışma Ağacı becomes al-ma-a-ac and Iğdır becomes i-d-r. The Turkish-aware method turns all three into readable addresses: istanbul-isik, calisma-agaci, igdir. When I ran the slugify function used by the site's slug generator locally on the same İstanbul Işık input, I also got istanbul-isik.
There are two separate topics in URLs. The first is percent-encoding: encodeURIComponent produces %C4%B0stanbul for İstanbul and %C4%B1%C4%9Fd%C4%B1r for ığdır. Such addresses are valid but unreadable for people. The second is host names: Node's URL class returned xn--istanbul-o0e.example as the host name of https://İSTANBUL.example/, while for https://ISTANBUL.example/ it returned istanbul.example. So a host name containing a capital İ did not end up as the same host name as its lower-case ASCII version. This is not a browser guarantee, only an observation from this run; but it is a good reason to pass host names that come from users through a standard parser such as URL before comparing them.
I have no measured knowledge of how search engines treat İ/ı, so I make no claim about it. The safe options are on your own side: build the permanent address as ASCII and lower case, set up a redirect from the old address to the new one if you change it later, and build the on-site search key (fold in section 3) with the same logic as the slug. If you want a ready-made tool, the slug generator converts Turkish letters to ASCII; for converting text with the Turkish rule there is the case converter.
Identifiers and protocol keywords: locale-independent operations
Machine-owned strings (HTTP header names, URL schemes, configuration keys, file extensions, HTML tags, identifiers) must not change with the user's language. The Java documentation says so explicitly: toLowerCase() and toUpperCase() without an argument are locale-sensitive and may give unexpected results for strings such as identifiers, protocol keys and HTML tags. The documentation states that "title".toUpperCase() returns "TİTLE" in a Turkish locale and says to use Locale.ROOT.
Microsoft's .NET documentation shows the security side: a call such as IsFileURI("file:") can return true under the U.S. English culture and false under the Turkish culture, so a check that blocks addresses beginning with FILE: could be bypassed on Turkish systems. For comparisons that need no linguistic knowledge, the documentation recommends StringComparison.Ordinal or OrdinalIgnoreCase.
In JavaScript, toLowerCase() and toUpperCase() are already locale-independent; the danger is code written for Turkish leaking into identifier comparison. The program below shows exactly this mistake: the same FILE: string becomes file: with toLowerCase() and fıle: under the tr-TR rule.
01const scheme = "FILE:";02console.log(scheme.toLowerCase() === "file:");03console.log(scheme.toLocaleLowerCase("tr-TR") === "file:", scheme.toLocaleLowerCase("tr-TR"));04 05const asciiLower = (s) => s.replace(/[A-Z]+/g, (m) => m.toLowerCase());06console.log(asciiLower("ID"), asciiLower("İD"), asciiLower("ID") === "id", asciiLower("İD") === "id");01$ node identifiers.mjs02true03false fıle:04id İd true falseThe second line gives false fıle:. The last line is an alternative policy for identifiers: asciiLower lowers only the letters A-Z and consults no locale; İD therefore stays İd and does not match id. That is deliberate: for identifiers you decide which characters are accepted. The sturdiest route is often to restrict the identifier at registration to a constrained ASCII set (for example [a-z0-9_-]).
- Machine-owned identifiers and keys: a locale-independent or ordinal operation; where possible, limit the accepted character set to ASCII.
- Turkish text shown to users: write the locale explicitly in the code (
"tr-TR"); do not inherit it from the environment. - Comparison key: keep the original value and fold only the key.
- At system boundaries (HTTP, file system, database), write down which side applies which rule.
A checklist, language by language
The table below summarises the pitfall and two safe routes for each language. Mind the last column: I ran the JavaScript and Python rows with the programs in this article; the Java, C# and SQL rows rest only on the official documentation and were not run.
| Language | Pitfall | For identifiers and keys | For Turkish text | Verification |
|---|---|---|---|---|
| JavaScript | toUpperCase()/toLowerCase() do not know Turkish; the /i regular expression does not match İ and ı; localeCompare without an argument followed the environment | toLowerCase(), identifiers limited to ASCII | toLocaleUpperCase("tr-TR"), Intl.Collator("tr"), localeCompare(…, "tr", {sensitivity}) | Run: Node v20.15.1, ICU 74.2 |
| Java | toUpperCase()/toLowerCase() without an argument use the default locale; in a Turkish locale "title".toUpperCase() gives "TİTLE". equalsIgnoreCase ignores the locale but does not know Turkish | toUpperCase(Locale.ROOT), toLowerCase(Locale.ROOT) | Locale.forLanguageTag("tr-TR"); Collator for locale-sensitive comparison | Not run: Java SE 21 documentation |
| C#/.NET | ToUpper()/ToLower() use the current culture; Compare defaults to culture-sensitive comparison, Equals to ordinal comparison. A culture-sensitive "file:" check can fail under Turkish | ToUpperInvariant(), StringComparison.Ordinal/OrdinalIgnoreCase | ToUpper(culture) with new CultureInfo("tr-TR"); comparisons that name the culture explicitly | Not run: Microsoft Learn |
| Python | upper(), lower() and casefold() apply the Unicode default: "i".upper() gives I, "İ".lower() gives two characters; sorted() is code point order | lower()/casefold() and identifiers limited to ASCII | A prior i/ı/İ/I mapping (tr_upper, tr_lower); an ICU-based library or locale.strxfrm for sorting | Run: Python 3.12.13 |
| SQL | The result depends on the collation of the column or database; in PostgreSQL the provider (libc or icu) matters as well | For identifier columns, a language-independent collation and a single-form value (e.g. lower-case ASCII) | Turkish collations such as MySQL utf8mb4_turkish_ci or SQL Server Turkish_CI_AS; a Turkish ICU or libc collation in PostgreSQL | Not run: PostgreSQL, MySQL and SQL Server documentation |
The table is a decision summary, not a measurement result. Java, C# and SQL behaviour depends on the environment and the collation settings; verify it in your own environment.
01String key = header.toUpperCase(Locale.ROOT);02String label = title.toUpperCase(Locale.forLanguageTag("tr-TR"));01string key = header.ToUpperInvariant();02bool same = string.Equals(a, b, StringComparison.OrdinalIgnoreCase);03string label = title.ToUpper(new CultureInfo("tr-TR"));Test strategy
Most Turkey-test bugs blow up not on the developer's machine but in a customer's or a CI server's Turkish-locale environment. Turn that around: supply Turkish yourself, both as test input and as test environment.
- Keep a fixed set of Turkish test strings:
istanbul,ISTANBUL,İSTANBUL,Iğdır,ığdır,IĞDIR,İzmir,TITLE,file,FILE:. - Try the same set in NFC and NFD form:
İcan arrive as a single code point or asIplus U+0307. - Write the expectation separately for each operation: upper case, lower case, length, equality, sorting, slug and regular expression.
- Label separately the places that name the locale explicitly and those that read it from the environment.
- Run the tests in at least two environments: an English and a Turkish (
tr-TR) locale.
The small check program below shows the idea. The first four checks test calls that name the locale explicitly or use none at all. The last two stand for code that trusts the environment's locale and deliberately assume en-US. I ran the program in two environments.
01console.log("default locale:", Intl.DateTimeFormat().resolvedOptions().locale);02console.log("title".toLocaleUpperCase(), "(toLocaleUpperCase, no argument)");03 04const letters = ["ı", "i", "İ", "I"];05const checks = [06 ['"TITLE".toLowerCase() is "title"', "TITLE".toLowerCase() === "title"],07 ['"file".toUpperCase() is "FILE"', "file".toUpperCase() === "FILE"],08 ['"istanbul" to tr-TR upper is "İSTANBUL"', "istanbul".toLocaleUpperCase("tr-TR") === "İSTANBUL"],09 ["Collator('tr') order is ı I i İ", [...letters].sort(new Intl.Collator("tr").compare).join(" ") === "ı I i İ"],10 ["ambient sort order is i I İ ı", [...letters].sort(new Intl.Collator().compare).join(" ") === "i I İ ı"],11 ["ambient accent-equal istanbul/İSTANBUL is false", "istanbul".localeCompare("İSTANBUL", undefined, { sensitivity: "accent" }) !== 0],12];13for (const [name, ok] of checks) console.log(ok ? "ok " : "FAIL", name);14process.exitCode = checks.every(([, ok]) => ok) ? 0 : 1;01$ LC_ALL=en_US.UTF-8 node ci-check.mjs02default locale: en-US03TITLE (toLocaleUpperCase, no argument)04ok "TITLE".toLowerCase() is "title"05ok "file".toUpperCase() is "FILE"06ok "istanbul" to tr-TR upper is "İSTANBUL"07ok Collator('tr') order is ı I i İ08ok ambient sort order is i I İ ı09ok ambient accent-equal istanbul/İSTANBUL is false10$ LC_ALL=tr_TR.UTF-8 node ci-check.mjs11default locale: tr-TR12TITLE (toLocaleUpperCase, no argument)13ok "TITLE".toLowerCase() is "title"14ok "file".toUpperCase() is "FILE"15ok "istanbul" to tr-TR upper is "İSTANBUL"16ok Collator('tr') order is ı I i İ17FAIL ambient sort order is i I İ ı18FAIL ambient accent-equal istanbul/İSTANBUL is false19# exit 1In the English environment all six checks gave ok. With LC_ALL=tr_TR.UTF-8 the first four checks still passed; the last two, which rely on the environment's locale, gave FAIL and the exit code was 1: Intl.Collator() and localeCompare without an argument saw that the environment was tr-TR. One more detail: "title".toLocaleUpperCase() without an argument returned TITLE in both environments, so the behaviour of “following the environment” can vary with the API and the engine. That is a reason to assume neither that the environment is Turkish nor that it is not: write the language tag wherever you want Turkish. The FAIL lines are not bugs but proof that the code depends on the environment.
In other runtimes the same idea is set up differently; these were not run. The example in the .NET documentation changes the culture with Thread.CurrentThread.CurrentCulture = new CultureInfo("tr-TR"). In Java you can switch the default locale to Turkish at the start of a test with Locale.setDefault and restore the old value afterwards. In Python, locale.setlocale selects a Turkish locale installed on the system (like locale_check.py above). In database tests, name the column's collation explicitly and query with rows that contain both a capital İ and a dotless ı.
Try it with your own text
To try the examples in this article with your own text, you can use the site's Turkish I test tool (the link at the end of the article goes there too). The tool answers the same question under two rules side by side: the root (default) locale and tr-TR. The calculation runs in your browser and none of the code you paste is executed.
A short guide: paste text into the upper field, one word per line, or use the Load example button; if you like, type a compare text into the second field. For every line the tool shows upper case, lower case, length after lower-casing, NFC/NFD length, equality, regular expression and slug results in two columns. Differing results are flagged with the Differs label, and U+0307 is written with its name COMBINING DOT ABOVE. At the bottom, the two sort orders (Array.prototype.sort() and Intl.Collator("tr")) are compared.
The tool is for diagnosis only. To convert text, use the case converter; to build addresses, use the slug generator. The tool's results depend on your browser's ICU data and can vary with the browser or its version; verify the Java, C# and SQL side separately in your own environment.
In short: the Turkish İ problem is not a bug but two correct rules meeting in the wrong place. Write the rule explicitly in the code, keep identifiers apart from the locale, and put Turkish into your test environment.
Official sources and further reading
The language behaviour and documentation statements in this article were checked against these sources on 11 October 2026. The JavaScript and Python examples were also run; the Java, C#/.NET and SQL statements rest only on these documents:
- Oracle Java SE 21: java.lang.String (toUpperCase, toLowerCase, equalsIgnoreCase)
- Microsoft Learn: Best practices for comparing strings in .NET
- Microsoft Learn: String.ToUpper method
- Python documentation: str.lower, str.upper and str.casefold
- PostgreSQL documentation: Collation Support
- MySQL 8.4 Reference Manual: Unicode Character Sets
- Microsoft Learn: Collation and Unicode support (SQL Server)
- MDN: String.prototype.toLocaleUpperCase()
- Unicode Character Database: SpecialCasing.txt
The source pages are living documents and can change. The JavaScript results were taken with Node v20.15.1 (ICU 74.2, Unicode 15.1) and the Python results with Python 3.12.13 (unicodedata 15.0.0) on a Linux machine; other browser, engine or library versions may give different results. The Java, C#/.NET and SQL statements were not run. “Turkey test” is a commonly used name for this class of bug; it is not attributed to any source. The links to the site's tools are in sections 5 and 9; the Turkish I test tool can also be reached from the button at the end of the article.
Four letters, two pairs: the problem is code that does not say which rule it uses.
Open the Turkish I test ↗