Does len() count characters or bytes?

Neither, strictly. len(s) counts code points. For ASCII that happens to equal both the byte count and what a reader would call characters, which is why the distinction stays hidden until an accent or an emoji arrives.

Overview

The answer

For a str, len counts Unicode code points. For bytes, it counts bytes. Neither necessarily matches what a person would call a character.

1Python
Output

That last line is the interesting one, and the reason this question gets asked at all.

StringsConceptualMedium

Step through it

What to watch

  • The second and third rows look the same and are not equal.
  • One is é as a single code point, the other is e plus a combining mark.
  • Normalising is what makes the comparison behave.

Say this out loud

"len counts code points. Bytes is len(s.encode('utf-8')), and neither is what a user calls a character once you have combining accents or emoji."

Does len() count characters or bytes?

Is len(s) the number of characters or the number of bytes?

Run it

The same visible text three ways, with all three counts printed for each. The two strings in the middle look identical in your browser and are not equal.

2Python
Output

The same text, two lengths

"é" has two valid Unicode representations:

FormCode pointslen
Precomposed (NFC)U+00E91
Decomposed (NFD)U+0065 U+0301 — e + combining accent2

Both display identically. Both are "correct". They are not equal:

3Python
Output

This is a real bug source, not a curiosity. macOS filesystems historically store filenames in NFD while Linux stores what it is given, so the same filename read from two systems compares unequal.

The fix is normalisation before comparing:

4Python
Output
FormDoes whatUse for
NFCCompose — shortest formStorage and comparison
NFDDecompose — base plus marksStripping accents
NFKCCompose plus compatibility foldingSearch keys, identifiers
NFKDDecompose plus compatibilityAggressive normalisation

The K forms also fold compatibility characters: "fi" becomes "fi", "①" becomes "1", and full-width Latin becomes ASCII. Useful for search and dangerous for round-tripping, since the transformation is lossy.

NFC is the default choice. Normalise on input, store normalised, and comparisons stop being surprising.

Three different notions of "length"

ConceptExample: "👨‍👩‍👧" (family emoji)Count
Bytes (UTF-8)Raw octets18
Code pointslen() in Python5
Grapheme clustersWhat a person sees1

That family emoji is three people joined by zero-width joiners — five code points that render as one glyph. len() reports 5, and no user would agree.

Python's standard library has no grapheme support. The regex module does, via \X:

import regex
len(regex.findall(r"\X", "👨‍👩‍👧"))    # 1

Which of the three you want depends entirely on the question being asked, and picking the right one is the actual skill:

Bytes for storage limits, network payloads, Content-Length, and database column sizes. Code points for most programming — slicing, iteration, regular expressions. Graphemes for anything a person perceives: cursor movement, truncating a display string, counting characters in a tweet.

Three different counts

There are three things you could mean by "length", and they coincide only for ASCII:

Code points — what len(s) returns. A Python 3 str is a sequence of code points, and indexing gives you one.

Bytes — len(s.encode('utf-8')). Depends entirely on the encoding, which is why the encoding has to be named. This is the number that matters for a database column or a network frame.

Grapheme clusters — what a human counts. Python has no built-in for this; it needs a library.

Where it goes wrong

Two strings that render identically can have different lengths and compare unequal, because é can be one code point or two (e plus a combining accent). Users produce both, depending on their keyboard and operating system.

The fix is normalisation: unicodedata.normalize('NFC', s) before comparing or storing. Any system that accepts names, usernames or search terms needs this, and most learn it from a bug report.

Emoji are the other reliable surprise. A family emoji or a skin-tone modifier is several code points joined by zero-width joiners, so len says 5 or 7 for something a user counts as one, and slicing it in half produces garbage.

The practical rule

Use len(s) for indexing and slicing your own text. Use len(s.encode('utf-8')) whenever a limit is a storage or transport limit. Normalise on input. And if you are truncating text for display — a preview, a tweet counter — know that code points are an approximation you may have to replace.

Why truncation goes wrong

Cutting text at a byte boundary can split a multi-byte character in half:

5Python
Output

The slice ends mid-character. Three ways out, in increasing correctness:

6Python
Output

Truncate in str space, then encode — correct for code points:

7Python
Output

Truncate by grapheme — correct for what the user sees, and the only version that will not cut an emoji family apart or orphan a combining accent.

The database version of this bug is common: a VARCHAR(10) counts characters in PostgreSQL and bytes in some MySQL configurations, so ten emoji may or may not fit. Knowing which unit a storage layer counts in is part of the schema.

Length in other languages

Languagelength counts
Python 3Code points
JavaScriptUTF-16 code units — emoji count as 2
JavaUTF-16 code units
GoBytes — len(s) on a string
RustBytes — chars().count() for code points
SwiftGrapheme clusters — the only mainstream language that gets this right by default
CBytes, up to the NUL terminator

JavaScript's behaviour explains a familiar class of bug: "👍".length is 2, so a naive character limit counts one emoji as two, and slicing at an odd offset produces a lone surrogate that renders as a replacement glyph.

Go's byte-based len is a deliberate choice — strings are byte slices, and for i, r := range s iterates runes. Swift's grapheme-cluster default is the most user-friendly and makes indexing O(n), which is the trade it accepts.

Being able to name these differences is what turns this question from trivia into a demonstration that you have handled text across systems.

Questions people ask

Is len O(1) in Python? Yes. The count is stored on the object, not computed.

Why does the same visible text have different lengths? Precomposed versus decomposed Unicode. Normalise with NFC before comparing.

How do I count what the user sees? Grapheme clusters, via the regex module's \X. The standard library cannot do it.

How do I get a byte count? len(s.encode("utf-8")), and the answer depends on the encoding.

Why is my emoji length 2 in JavaScript and 1 in Python? JavaScript counts UTF-16 code units; characters above U+FFFF need a surrogate pair.

Should I normalise everything on input? For anything compared, stored as a key, or deduplicated — yes, to NFC. It is cheap and it prevents an entire class of bug.

Recap in one screen

  • len(str) counts code points; len(bytes) counts bytes; neither counts what a person calls a character.
  • Precomposed and decomposed forms look identical and compare unequal — normalise to NFC first.
  • Grapheme clusters are the user-facing unit, and Python needs the regex module's \X to count them.
  • Never truncate encoded bytes at an arbitrary offset; truncate the str and then encode.
  • JavaScript and Java count UTF-16 units, Go and Rust count bytes, Swift counts graphemes.

How the code works

The same visible text three ways, with all three counts printed for each. The two strings in the middle look identical in your browser and are not equal.

How the code works

  1. len(s)Counts code points, because a Python 3 str is a sequence of code points. Indexing returns one of them.
  2. len(s.encode('utf-8'))The byte count, and the only one that answers "will this fit in the column?". It depends on the encoding, which is why the encoding must be named rather than assumed.
  3. a == bFalse for two strings that render identically. Equality is over code points, and the two forms are different sequences of them.
  4. unicodedata.normalize("NFC", ...)Puts both into the same canonical form so they compare equal. Any system taking user-entered names or search terms needs this on the way in.

Change one thing

  • Add a family emoji. len reports 7 or more — it is several people joined by zero-width joiners.
  • Encode the emoji as utf-32-le and compare the byte counts. Fixed-width encodings make len and bytes agree again, at four bytes per code point.

Where this runs

Real CPython, compiled to WebAssembly and running on your own machine — nothing is uploaded. The first run takes a few seconds while the interpreter downloads; after that it is immediate. Need more room, or want to paste your own attempt? Use the Python compiler.

Check yourself

0 of 3

Answer without scrolling back up.

  1. len(s) on a Python 3 string counts:

  2. Two strings render identically but compare unequal. The most likely cause is:

  3. You need to enforce a 100-character database limit. You should measure:

Cheat sheet

Does len() count characters or bytes?

Neither, strictly. len(s) counts code points. For ASCII that happens to equal both the byte count and what a reader would call characters, which is why the distinction stays hidden until an accent or an emoji arrives.

INTERVIEW · vizlearn.in/interview/does-len-count-characters-or-bytes.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.