Run it
The same visible text three ways, with all three counts printed for each. The two strings in the middle look identical in your browser and are not equal.
The same text, two lengths
"é" has two valid Unicode representations:
| Form | Code points | len |
|---|
| Precomposed (NFC) | U+00E9 | 1 |
| Decomposed (NFD) | U+0065 U+0301 — e + combining accent | 2 |
Both display identically. Both are "correct". They are not equal:
This is a real bug source, not a curiosity. macOS filesystems historically store filenames in NFD while Linux stores what it is given, so the same filename read from two systems compares unequal.
The fix is normalisation before comparing:
| Form | Does what | Use for |
|---|
| NFC | Compose — shortest form | Storage and comparison |
| NFD | Decompose — base plus marks | Stripping accents |
| NFKC | Compose plus compatibility folding | Search keys, identifiers |
| NFKD | Decompose plus compatibility | Aggressive normalisation |
The K forms also fold compatibility characters: "fi" becomes "fi", "①" becomes "1", and full-width Latin becomes ASCII. Useful for search and dangerous for round-tripping, since the transformation is lossy.
NFC is the default choice. Normalise on input, store normalised, and comparisons stop being surprising.
Three different notions of "length"
| Concept | Example: "👨👩👧" (family emoji) | Count |
|---|
| Bytes (UTF-8) | Raw octets | 18 |
| Code points | len() in Python | 5 |
| Grapheme clusters | What a person sees | 1 |
That family emoji is three people joined by zero-width joiners — five code points that render as one glyph. len() reports 5, and no user would agree.
Python's standard library has no grapheme support. The regex module does, via \X:
import regex
len(regex.findall(r"\X", "👨👩👧")) # 1
Which of the three you want depends entirely on the question being asked, and picking the right one is the actual skill:
Bytes for storage limits, network payloads, Content-Length, and database column sizes. Code points for most programming — slicing, iteration, regular expressions. Graphemes for anything a person perceives: cursor movement, truncating a display string, counting characters in a tweet.
Three different counts
There are three things you could mean by "length", and they coincide only for ASCII:
Code points — what len(s) returns. A Python 3 str is a sequence of code points, and indexing gives you one.
Bytes — len(s.encode('utf-8')). Depends entirely on the encoding, which is why the encoding has to be named. This is the number that matters for a database column or a network frame.
Grapheme clusters — what a human counts. Python has no built-in for this; it needs a library.
Where it goes wrong
Two strings that render identically can have different lengths and compare unequal, because é can be one code point or two (e plus a combining accent). Users produce both, depending on their keyboard and operating system.
The fix is normalisation: unicodedata.normalize('NFC', s) before comparing or storing. Any system that accepts names, usernames or search terms needs this, and most learn it from a bug report.
Emoji are the other reliable surprise. A family emoji or a skin-tone modifier is several code points joined by zero-width joiners, so len says 5 or 7 for something a user counts as one, and slicing it in half produces garbage.
The practical rule
Use len(s) for indexing and slicing your own text. Use len(s.encode('utf-8')) whenever a limit is a storage or transport limit. Normalise on input. And if you are truncating text for display — a preview, a tweet counter — know that code points are an approximation you may have to replace.
Why truncation goes wrong
Cutting text at a byte boundary can split a multi-byte character in half:
The slice ends mid-character. Three ways out, in increasing correctness:
Truncate in str space, then encode — correct for code points:
Truncate by grapheme — correct for what the user sees, and the only version that will not cut an emoji family apart or orphan a combining accent.
The database version of this bug is common: a VARCHAR(10) counts characters in PostgreSQL and bytes in some MySQL configurations, so ten emoji may or may not fit. Knowing which unit a storage layer counts in is part of the schema.
Length in other languages
| Language | length counts |
|---|
| Python 3 | Code points |
| JavaScript | UTF-16 code units — emoji count as 2 |
| Java | UTF-16 code units |
| Go | Bytes — len(s) on a string |
| Rust | Bytes — chars().count() for code points |
| Swift | Grapheme clusters — the only mainstream language that gets this right by default |
| C | Bytes, up to the NUL terminator |
JavaScript's behaviour explains a familiar class of bug: "👍".length is 2, so a naive character limit counts one emoji as two, and slicing at an odd offset produces a lone surrogate that renders as a replacement glyph.
Go's byte-based len is a deliberate choice — strings are byte slices, and for i, r := range s iterates runes. Swift's grapheme-cluster default is the most user-friendly and makes indexing O(n), which is the trade it accepts.
Being able to name these differences is what turns this question from trivia into a demonstration that you have handled text across systems.
Questions people ask
Is len O(1) in Python? Yes. The count is stored on the object, not computed.
Why does the same visible text have different lengths? Precomposed versus decomposed Unicode. Normalise with NFC before comparing.
How do I count what the user sees? Grapheme clusters, via the regex module's \X. The standard library cannot do it.
How do I get a byte count? len(s.encode("utf-8")), and the answer depends on the encoding.
Why is my emoji length 2 in JavaScript and 1 in Python? JavaScript counts UTF-16 code units; characters above U+FFFF need a surrogate pair.
Should I normalise everything on input? For anything compared, stored as a key, or deduplicated — yes, to NFC. It is cheap and it prevents an entire class of bug.
Recap in one screen
len(str) counts code points; len(bytes) counts bytes; neither counts what a person calls a character.- Precomposed and decomposed forms look identical and compare unequal — normalise to NFC first.
- Grapheme clusters are the user-facing unit, and Python needs the
regex module's \X to count them. - Never truncate encoded bytes at an arbitrary offset; truncate the
str and then encode. - JavaScript and Java count UTF-16 units, Go and Rust count bytes, Swift counts graphemes.