Text statistics
InspectExperimentalLimited support. Verify anything critical.
Text statistics is a browser-local measurement tool for counting UTF-16 characters, Unicode code points, grapheme clusters, whitespace-delimited words, and line breaks. It helps developers compare different notions of text length, diagnose differences between JavaScript string metrics and user-perceived characters, and inspect line structure while keeping the complete input in local browser memory. Results are calculated on your device, with no pasted content sent to a processing server and no input stored for later retrieval.
Skip to the toolThis tool processes text on your device. The text is not uploaded.
How to use Text statistics
What is Text statistics?
Text statistics is a comprehensive measurement tool designed to analyze string length across multiple layers of abstraction. It goes beyond simple character counting by providing real-time metrics for standard UTF-16 code units (essential for database column constraints), grapheme clusters (the actual visual characters perceived by human readers), whitespace-delimited words, and line breaks. Whether you are typing or pasting large blocks of text, the counts update instantly. Most importantly, all text processing happens entirely locally within your browser’s memory; your data is never sent over a network, guaranteeing complete privacy and ensuring that sensitive strings never leave your device.
The Multi-Layered Nature of Text Measurement
In modern internationalized software development, determining the “length” of a text string requires an understanding of three distinct layers of Unicode abstraction:
- UTF-16 Code Units (
characters): JavaScript’s native string.lengthproperty counts 16-bit code units. Characters in the Basic Multilingual Plane (BMP) occupy a single code unit. However, supplementary characters, such as modern emojis or rare historical symbols, require “surrogate pairs” and thus count as two code units. - Unicode Code Points (
codePoints): Each individual Unicode character represents a unique scalar value. For instance, the letterAis represented byU+0041(one code point), whereas the rocket emoji🚀(U+1F680) also represents one code point despite consuming two UTF-16 code units. - Extended Grapheme Clusters (
graphemes): What a human reader naturally perceives as a single visual character can actually consist of multiple combined code points. Consider the family emoji👨👩👧👦. It is constructed from four individual person emojis joined by Zero-Width Joiners (ZWJ). While it renders as one grapheme cluster (one visual character), it is composed of seven Unicode code points and eleven UTF-16 code units.
Practical developer use cases
- Database Column Sizing: Validate string lengths against relational database
VARCHAR(N)orCHAR(N)column constraints before running bulk data imports, preventing unexpected truncation errors. - SEO and Content Metadata: Ensure web page
<title>tags remain under search engines’ desktop truncation limits (typically around 60 characters) and meta descriptions adhere to the optimal 155–160 character boundary. - Social Media and Ad Copy: Audit text length for platform-specific constraints, such as X/Twitter (280 characters), LinkedIn posts, or strict SMS messaging limits.
- LLM Token Estimation: Quickly estimate token consumption for generative AI prompts. A common rule of thumb for English text is that one token roughly equals four characters or 0.75 words.
Best practices
- Understand your boundaries: When working with standard APIs or database engines, know whether they evaluate length by byte size, UTF-16 code units, or code points. Use the UTF-16 code unit count for JavaScript environments.
- Use graphemes for UI constraints: If you are limiting how much space a user’s input takes up visually (such as a username display), rely on the grapheme cluster count to prevent cutting off complex characters or emojis.
- Handle word counting nuances: Be aware that word counting in this tool relies on whitespace splitting. Hyphenated compounds (like “state-of-the-art”) are counted as a single word. If your application requires rigorous linguistic word tokenization, consider specialized NLP libraries.
- Normalize line endings: The tool gracefully handles Unix (
\n), Windows (\r\n), and classic Mac (\r) line terminators. When measuring lines for legacy systems, make sure you account for these differing formats.
Security considerations
- Zero-network telemetry: Because this tool operates 100% locally in your web browser, it is perfectly safe to use for measuring sensitive, proprietary, or classified text. No server-side storage or external API calls are made.
- Client-side performance: Be mindful of extremely large payloads (e.g., gigabyte-sized log files). Memory usage scales linearly with the input size, which may cause browser tab slow-downs or crashes on low-resource devices.
- Cross-Site Scripting (XSS): While this tool simply counts and does not execute the text, always ensure that text input copied from external sources and pasted into your own systems is properly sanitized to prevent injection attacks.
Code examples
If you need to implement similar text statistics in your own projects, you can use these reference implementations.
JavaScript / TypeScript (Modern Unicode Standard)
Modern JavaScript provides Intl.Segmenter to accurately count grapheme clusters, properly handling complex emojis and language-specific ligatures.
function analyzeText(text: string) {
// Use Intl.Segmenter to count visual grapheme clusters
const segmenter = new Intl.Segmenter("en", { granularity: "grapheme" });
const graphemes = Array.from(segmenter.segment(text)).length;
// Array.from splits by code points, not code units
const codePoints = Array.from(text).length;
// Standard length returns UTF-16 code units
const characters = text.length;
// Basic whitespace-delimited word count
const words = text.trim() ? text.trim().split(/\s+/).length : 0;
// Universal line ending split
const lines = text ? text.split(/\r\n|\r|\n/).length : 0;
return { characters, codePoints, graphemes, words, lines };
}Python 3
In Python, the built-in len() function accurately counts Unicode code points. To count extended grapheme clusters, you can use the third-party regex library (which supports the \X pattern).
import regex
def analyze_text(text: str) -> dict:
# \X matches a Unicode extended grapheme cluster
graphemes = len(regex.findall(r"\X", text))
# Split by any whitespace character
words = len(text.strip().split()) if text.strip() else 0
# splitlines() handles universal newlines automatically
lines = len(text.splitlines()) if text else 0
# In Python 3, len() on a string returns the number of code points
return {
"characters": len(text),
"graphemes": graphemes,
"words": words,
"lines": lines,
}How it works
- Enter what you haveType or pick your text. Nothing is submitted anywhere.
- It runs in this tabThe calculation happens on your device, using your browser's own data.
- Take the resultRead the text, then copy, download, or share a link.
Reimplemented locally. Not derived from IT-Tools source.
- Basis
- independent
- Licence
- MIT
- Last reviewed
Frequently asked questions
What is the difference between UTF-16 code units, code points, and grapheme clusters?
UTF-16 code units reflect standard JavaScript string length (where surrogate pairs count as 2). Code points represent individual Unicode scalar values. Grapheme clusters represent human-perceived visual characters (e.g. complex emojis with skin-tone modifiers or zero-width joiners).
How are words counted?
Words are calculated by trimming whitespace and splitting on consecutive whitespace sequences (\s+). Hyphenated compounds (e.g. 'state-of-the-art') count as a single word.
How does the tool count multi-line text and different newline formats?
The line counter handles Unix (\n), Windows (\r\n), and classic Mac (\r) line terminators consistently, reporting the total number of non-empty and empty lines.
Is my pasted text sent to an external server for analysis?
No. All string segmentation, counting, and metric calculations run entirely in your local browser memory with zero network telemetry.