Text Frequency Analyzer
Find repeated words and adjacent phrases, with counts and percentages that keep their denominators clear.
Analysis runs in a locally bundled browser worker. No account, text upload or saved draft is needed.Source text
0 / 2,097,152 UTF-8 bytes2 MiB input · 100,000 tokens · local worker processing. Refreshing clears drafts. Imported CRLF, LF and CR all separate phrases.
1000 entries maximum, 100 code points each; 512 KiB list limit. Phrases containing a stop word are excluded, never joined across it.
Frequency results
Paste text and choose Analyze. Editing source or analysis options discards old results.
How to analyze word and phrase frequency
- Paste plain text, load the example, or import one UTF-8 .txt file. Import removes one leading BOM and rejects invalid UTF-8 without replacing your current source.
- Choose whether to ignore case and which stop words to exclude. Choose Analyze.
- Switch between single words, two-word phrases and three-word phrases. Adjust minimum count, maximum visible rows, sort order or the term filter, then choose Update table. These view operations do not recount the source.
- Select rows to copy, copy every matching row, or prepare a UTF-8 TXT download. The complete JSON report includes all three tabs, without display filters.
A worked example
Red blue red Ignoring case: red → 2 occurrences → 66.67% blue → 1 occurrence → 33.33% Two-word windows: red blue; blue red Each occurs once → 50% of 2 eligible windows
With case matching enabled, Red, blue and red are three distinct words. For a a a, overlapping windows give two a a occurrences and one a a a occurrence.
Tokenization and case rules
A token starts with a Unicode letter or number and includes attached combining marks. Internal straight/curly apostrophes and hyphen-minus join runs: don't, l'été and re-entry remain single tokens. Emoji and isolated punctuation do not count. Other dashes, including Unicode hyphen U+2010 and nonbreaking hyphen U+2011, separate tokens.
Ignore case uses locale-independent Unicode lowercasing, not accent folding, stemming or Unicode normalization. Composed é and decomposed e plus a combining accent remain distinct. Lowercasing can change character length, so the tool retains first-seen spelling but does not highlight source positions. Continuous text in scripts without word separators may form one long token; this is not language-aware word segmentation.
The existing word counter also supports Unicode hyphens and optional counting rules. This analyzer deliberately uses the narrower baseline above without changing the counter's behavior.
Phrase boundaries and stop words
Phrases are overlapping adjacent token windows within a physical line. They may cross punctuation on that line: red, blue produces red blue. They never cross CRLF, LF or standalone CR. This simple policy is not sentence-aware; punctuation between sentences on one line does not stop a window.
A stop word excludes a single-word occurrence and every original phrase window containing it. Stop words are not deleted before building windows: red and blue with and excluded leaves two eligible singles and no eligible two- or three-word windows. It never invents red blue.
The optional, locally bundled English preset contains 23 words: a, an, and, are, as, at, be, by, for, from, in, is, it, of, on, or, that, the, this, to, was, were, with. It does not detect your language. Custom entries are always applied and must be one complete token each, separated by commas or line breaks. Rejected entries are shown instead of silently splitting phrases. Stop words use the same case option; with case matching enabled the lowercase preset does not remove The.
Counts, percentages and display filters
Total tokens includes stop words. Eligible words counts only non-stop occurrences. Each tab shows original window count, eligible denominator and distinct terms before display filtering. Percentages divide a term's exact integer count by that tab's eligible denominator. All shares sum to 100% before display filtering, apart from display rounding to at most two decimals. A zero denominator produces an empty result, never an invalid percentage.
Minimum count, term search and the top-row limit do not change denominators. Term search is a literal substring match using the chosen case option. Frequency sorting breaks ties by Unicode code-point order; alphabetical sorting uses that same deterministic order, not locale dictionary order.
Exports, limits and privacy
Input is limited to 2 MiB UTF-8 and 100,000 tokens. Custom stop words are limited to 1000 entries, 100 code points each, and 512 KiB total. Only 1–1000 rows are displayed; long table terms are visibly shortened. Copy and TXT download include every matching row, ignoring the display limit. Selected-row copy uses table order. The selectable preview is limited to 65,536 UTF-16 code units and says when shortened.
JSON includes tokenizer version, analysis options, stop words, all tab denominators, token arrays, exact counts, unrounded percentage shares and first-seen spelling/spans. It omits the original source document. Export size is bounded at 64 MiB; an oversized report fails explicitly rather than offering a partial file.
Analysis, table sorting and exports run in a worker. Cancel or edits to source/counting options terminate obsolete work and discard results. View changes hide old tables until updated. Clear, example loading and valid imports offer in-memory Undo for the source and options; analyze again after restoring. Refreshing clears everything. Clipboard failure leaves selectable plain text and manual-copy or full-download guidance.
Literal HTML remains text. No input is placed in URLs, analytics, application logs, network requests or browser storage. Frequency alone does not determine writing quality or search ranking. This tool does not detect semantic topics, plagiarism, keyword search demand or recommended SEO density.
What should we improve next?
Help shape the next update to Text frequency analyzer. Tell us what would make it better for you.
Prefer email? feedback@tooltulip.com