What does it do, and who is it for?
The keyword density checker lists which words and phrases a text repeats and how often. Each row shows the phrase, its count, its percentage and the word position where it first appears. You can download the full result as CSV.
Writers use it to spot a phrase they overuse in a draft, editors to see whether a page actually stays on topic, and SEO specialists to review phrase distribution on older or competing pages. It does not produce a score; it is a measurement that leaves the judgement to the person reading the text.
How to use it
- Paste your text. The table refreshes a moment after you stop typing.
- Choose the analysis language. English is the default on this page, but you can switch to Turkish to analyse Turkish text with Turkish casing rules.
- Choose the phrase length: single words, two-word or three-word phrases.
- Optionally hide stop words. The number of hidden rows appears in the summary.
- To count inflected forms together, turn on variant grouping and fill in the alias table.
- Copy the table or download it as CSV.
For total length and reading time of the same text, pair it with the word counter.
How is the denominator calculated?
For single words, the percentage is count / total words × 100. For two- and three-word phrases, the denominator is the number of valid phrase windows in the text: each sentence contributes words − n + 1 windows, and these are added up. That way every row in the table is measured against the same base.
Words are split with the browser's Intl.Segmenter. Punctuation and emoji are not words, and diacritics are kept, so café and cafe are separate rows. Percentages are displayed with two decimals but sorted with full precision: highest count first, ties in alphabetical order for the chosen language.
How n-grams are counted
An n-gram is a run of n consecutive words. Windows overlap: in seo seo seo there are two two-word windows, and seo seo appears twice, which is 100%. Phrases never cross a sentence boundary, so the last word of one sentence and the first word of the next do not form a phrase. Line breaks count as boundaries too, so a heading does not merge with the first paragraph.
Stop-word filter
Stop words are very frequent function words such as the, and, of or to. The filter does not remove them from the text; it only hides rows in the result. For single words, stop-word rows are hidden. For longer phrases, a row is hidden when its first or last word is a stop word, while phrases like search and content stay visible. The denominator and word adjacency never change, so search and content does not produce an artificial search content phrase.
Turkish variant grouping
Turkish is agglutinative: kitap (book), kitaplar (books) and kitapları (the books) refer to the same topic but count as different words. The same applies to English plurals such as tool and tools. The checker does not claim to be a general stemmer and never strips suffixes blindly. Instead it uses an alias table that is off by default, visible and editable. Write one variant => canonical rule per line; separate several variants with commas.
kitaplar, kitapları => kitap
tools => toolWith grouping on, each row shows the canonical phrase together with the variants that produced it. Mapping one variant to two different canonical words is reported as an error.
Worked example
Input: SEO content. SEO content strategy and content plan. The text has two sentences and 8 words in total.
| Setting | Result |
|---|---|
| Single words | content 3 (37.50%), seo 2 (25.00%), others 1 (12.50%) |
| Two words | 6 windows; seo content 2 (33.33%) |
| Two words + stop-word filter | 2 rows hidden, denominator still 6 |
In the two-word view, content seo never appears because the first sentence ends with a full stop. With the filter on, strategy and and and content are hidden, yet seo content stays at 33.33%.
Keyword stuffing and limits
Repeating a phrase without adding value for the reader is known as keyword stuffing, and it is listed in Google's spam policies. This tool gives no “ideal density” or SEO score, because no such fixed ratio exists. A high percentage is simply a prompt to reread the text.
- Input is limited to 100,000 Unicode code points; longer text is not cut silently, you get an error instead.
- The table shows the first 200 rows; the CSV file contains every row.
- Hyphenated compounds may be split into two words by the segmenter.
- The alias table only groups the words you list; other inflected forms stay separate.
Frequently asked questions
What is the ideal keyword density?
There is no universal ideal percentage. Reading naturally and covering the topic well matters more than hitting a particular number.
Why are “SEO” and “seo” in the same row?
The analysis is case-insensitive. With Turkish selected, İ lower-cases to i and I to ı, so IŞIK and ışık are grouped while ilk and ılık stay apart.
Does it merge plurals or suffixes automatically?
No. To avoid wrong merges, suffixes are never stripped automatically. Turn on variant grouping and write your own rules in the alias table.
Does the stop-word filter change the percentages?
No. It only hides rows; total words, window count and percentages stay the same.
Is my text sent to a server?
No. The calculation runs in your browser. The text only ends up in a URL if you choose to include inputs in a share link.