What is the difference between str and bytes?

A str is a sequence of code points and carries no encoding. bytes is a sequence of 8-bit values and carries no meaning until you name one. encode goes str → bytes, decode comes back, and the rule is to do both at the edges of your program and work in str everywhere inside.

Overview

Two different things

str is a sequence of Unicode code points — abstract characters. bytes is a sequence of integers from 0 to 255 — raw octets. Python 3 keeps them strictly separate, and the separation is the feature.

1Python
Output

Two details catch people out immediately. Indexing bytes yields an int, not a one-byte bytes object — slicing does give bytes, so b[0:1] is b'c' while b[0] is 99. And len differs between the two representations of the same text, which is the source of most truncation bugs.

 strbytes
ElementUnicode code pointInteger 0–255
x[0] givesstr of length 1int
Literal"text"b"text"
Convert.encode().decode()
Use forText you manipulateI/O, network, files, crypto
Mutable versionNonebytearray
StringsConceptualMedium

Step through it

What to watch

  • One code point became two bytes — positions stop matching.
  • Slicing bytes can split a character in half; slicing str cannot.
  • The boundary is always I/O: files, sockets, subprocesses.

Say this out loud

"str is text, bytes is data. Encode at the way out, decode at the way in, and never let bytes travel through your logic."

What is the difference between str and bytes?

What is the difference between str and bytes, and when do you need encode and decode?

Run it

The same text as both types, with the ways they refuse to mix, and both Unicode errors triggered on purpose so you have seen the tracebacks before an interviewer describes one.

2Python
Output

Encode and decode, in the right direction

This is the single most confused pair in Python, and there is a reliable way to remember it: encode goes to bytes; decode comes back to text.

3Python
Output

The mental model: encoding is what you do to send text somewhere, because wires and disks carry bytes. Decoding is what you do on receipt, to recover meaning.

UTF-8 is the default for both, and it is nearly always the right choice. It is ASCII-compatible, so English text costs one byte per character, and it represents every Unicode character.

Mixing the two types raises rather than coercing:

4Python
Output

The == case is the dangerous one, because it fails silently. A dictionary keyed on str will never match a bytes lookup, and the result is a mysteriously empty response rather than an exception. This is the most common bug when reading a file in binary mode and comparing against string literals.

The errors that appear in production

5Python
Output

Both mean the same thing: the data does not fit the codec. The errors parameter decides what happens next.

ValueBehaviourUse when
"strict"Raise — the defaultCorrectness matters
"ignore"Drop the offending charactersRarely; silently loses data
"replace"Substitute ? or U+FFFDDisplay-only output
"surrogateescape"Round-trip undecodable bytesFilenames, lossless pass-through
"backslashreplace"Escape as \xNNDebugging and logs

surrogateescape is the one worth knowing about. It maps undecodable bytes into a reserved surrogate range so that decoding then re-encoding reproduces the original bytes exactly. That is how Python handles filenames on systems where the name is not valid UTF-8 — it can still be opened, even though it cannot be meaningfully displayed.

Reaching for errors="ignore" to make an exception go away is almost always the wrong fix; it converts a loud failure into corrupted data.

Two different things that both print nicely

"hi" and b"hi" look almost identical and are not comparable: "hi" == b"hi" is False, and in Python 3 that is deliberate. A str has no encoding — asking for the bytes of a str is meaningless until you say which bytes, which is why encode takes an argument.

Indexing differs too. s[0] on a str gives a one-character str; b[0] on bytes gives an int. That catches people constantly.

The sandwich rule

Decode at the boundary in, encode at the boundary out, and keep the middle in str. Every layer of a well-behaved program works in text; only the outermost layer knows about UTF-8.

Violating it produces the two classic errors. UnicodeDecodeError means you were handed bytes that are not valid in the encoding you claimed — usually the encoding is wrong, not the data. UnicodeEncodeError means you tried to write a character the target encoding cannot represent, which is what happens when something defaults to ASCII or latin-1.

Where it actually shows up

Reading a file with open(path) gives str and applies your platform's default encoding, which differs between machines and is the source of "works on my laptop" bugs. Pass encoding="utf-8" explicitly, always. open(path, "rb") gives bytes and does not guess.

Sockets, subprocess output, hashlib and most binary formats are bytes. hashlib.sha256(s) is a TypeError until you encode — a hash is defined over bytes, so the encoding is part of the answer.

Where the boundary sits in real code

The practical rule: decode at the edges, work in str in the middle, encode on the way out. Text processing should never happen on bytes.

6Python
Output

Always pass encoding= explicitly when opening text files. Without it, Python uses the platform default, which is UTF-8 on Linux and macOS and historically a legacy code page on Windows — so the same code reads different characters on different machines. Python 3.15 makes UTF-8 the default everywhere, and being explicit remains the habit that survives version differences.

Three things must stay in bytes:

Hashing and cryptography. hashlib.sha256("text") raises; it needs b"text". Digests are defined over bytes, so the encoding is part of the hash. Binary formats. Images, protocol buffers, compressed data — there is no text to recover. Anything length- or offset-sensitive at the wire level, such as a Content-Length header, which counts bytes rather than characters.

bytearray and memoryview

7Python
Output

bytearray is the mutable counterpart, useful for building a buffer incrementally without repeated concatenation. memoryview slices without copying, which is what makes parsing large binary payloads efficient — and it has no str equivalent, because a character offset in a variable-width representation is not a fixed byte offset.

Questions people ask

Which way does encode go? str.encode() produces bytes. Text encodes to bytes for transport.

Why does b[0] give a number? bytes is a sequence of integers. Use b[0:1] for a one-byte bytes object.

Why is b"abc" == "abc" False rather than an error? Equality across unrelated types returns False by design; only ordering and concatenation raise. This is why the comparison bug is so easy to miss.

What is the default encoding? UTF-8 for encode/decode. File open used the locale default before Python 3.15, so pass it explicitly.

How do I get the byte length of text? len(s.encode("utf-8")). len(s) counts characters.

Can I put bytes and str keys in one dict? Yes, and they never collide — which is a footgun, not a feature.

What about UTF-16 or Latin-1? Latin-1 is the identity mapping for bytes 0–255, so it never raises on decode — which makes it a tempting and usually wrong way to silence errors.

Recap in one screen

  • str holds Unicode code points; bytes holds integers 0–255. They never mix implicitly.
  • encode goes str → bytes; decode goes back. UTF-8 unless told otherwise.
  • b[0] is an int; b[0:1] is bytes. And len differs between the two forms of the same text.
  • b"abc" == "abc" is silently False, which is the bug that hides in dictionary lookups.
  • Decode at the input edge, work in str, encode on output — and always pass encoding= to open.

How the code works

The same text as both types, with the ways they refuse to mix, and both Unicode errors triggered on purpose so you have seen the tracebacks before an interviewer describes one.

How the code works

  1. s.encode("utf-8")The encoding is an argument because a str genuinely does not have bytes until you choose. There is no default worth relying on.
  2. b[1] is an intIndexing bytes yields the numeric value, not a one-length bytes. It is the single most common surprise when code written for str is pointed at bytes.
  3. b[:2].decode("utf-8")Raises, because the slice split a two-byte character. Any code that chunks a byte stream has to respect character boundaries, which is why streaming decoders exist.
  4. b.decode("latin-1")The dangerous one: no exception, wrong text. latin-1 maps every byte to something, so it never fails and never warns. Mojibake is this, not a corrupted file.

Change one thing

  • Encode as utf-16 and print the bytes. Note the byte-order mark at the front, and that the length roughly doubles.
  • Call b.decode("utf-8", errors="replace") on the broken slice. You get U+FFFD instead of an exception — useful for logs, wrong for data you intend to keep.

Where this runs

Real CPython, compiled to WebAssembly and running on your own machine — nothing is uploaded. The first run takes a few seconds while the interpreter downloads; after that it is immediate. Need more room, or want to paste your own attempt? Use the Python compiler.

Check yourself

0 of 3

Answer without scrolling back up.

  1. b[0] where b is a bytes object gives you:

  2. Decoding UTF-8 data as latin-1 produces:

  3. Where should encode and decode happen in a well-structured program?

Cheat sheet

What is the difference between str and bytes?

A str is a sequence of code points and carries no encoding. bytes is a sequence of 8-bit values and carries no meaning until you name one. encode goes str → bytes, decode comes back, and the rule is to do both at the edges of your program and work in str everywhere inside.

INTERVIEW · vizlearn.in/interview/str-versus-bytes-in-python.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.