Run it
The same text as both types, with the ways they refuse to mix, and both Unicode errors triggered on purpose so you have seen the tracebacks before an interviewer describes one.
Encode and decode, in the right direction
This is the single most confused pair in Python, and there is a reliable way to remember it: encode goes to bytes; decode comes back to text.
The mental model: encoding is what you do to send text somewhere, because wires and disks carry bytes. Decoding is what you do on receipt, to recover meaning.
UTF-8 is the default for both, and it is nearly always the right choice. It is ASCII-compatible, so English text costs one byte per character, and it represents every Unicode character.
Mixing the two types raises rather than coercing:
The == case is the dangerous one, because it fails silently. A dictionary keyed on str will never match a bytes lookup, and the result is a mysteriously empty response rather than an exception. This is the most common bug when reading a file in binary mode and comparing against string literals.
The errors that appear in production
Both mean the same thing: the data does not fit the codec. The errors parameter decides what happens next.
| Value | Behaviour | Use when |
|---|
"strict" | Raise — the default | Correctness matters |
"ignore" | Drop the offending characters | Rarely; silently loses data |
"replace" | Substitute ? or U+FFFD | Display-only output |
"surrogateescape" | Round-trip undecodable bytes | Filenames, lossless pass-through |
"backslashreplace" | Escape as \xNN | Debugging and logs |
surrogateescape is the one worth knowing about. It maps undecodable bytes into a reserved surrogate range so that decoding then re-encoding reproduces the original bytes exactly. That is how Python handles filenames on systems where the name is not valid UTF-8 — it can still be opened, even though it cannot be meaningfully displayed.
Reaching for errors="ignore" to make an exception go away is almost always the wrong fix; it converts a loud failure into corrupted data.
Two different things that both print nicely
"hi" and b"hi" look almost identical and are not comparable: "hi" == b"hi" is False, and in Python 3 that is deliberate. A str has no encoding — asking for the bytes of a str is meaningless until you say which bytes, which is why encode takes an argument.
Indexing differs too. s[0] on a str gives a one-character str; b[0] on bytes gives an int. That catches people constantly.
The sandwich rule
Decode at the boundary in, encode at the boundary out, and keep the middle in str. Every layer of a well-behaved program works in text; only the outermost layer knows about UTF-8.
Violating it produces the two classic errors. UnicodeDecodeError means you were handed bytes that are not valid in the encoding you claimed — usually the encoding is wrong, not the data. UnicodeEncodeError means you tried to write a character the target encoding cannot represent, which is what happens when something defaults to ASCII or latin-1.
Where it actually shows up
Reading a file with open(path) gives str and applies your platform's default encoding, which differs between machines and is the source of "works on my laptop" bugs. Pass encoding="utf-8" explicitly, always. open(path, "rb") gives bytes and does not guess.
Sockets, subprocess output, hashlib and most binary formats are bytes. hashlib.sha256(s) is a TypeError until you encode — a hash is defined over bytes, so the encoding is part of the answer.
Where the boundary sits in real code
The practical rule: decode at the edges, work in str in the middle, encode on the way out. Text processing should never happen on bytes.
Always pass encoding= explicitly when opening text files. Without it, Python uses the platform default, which is UTF-8 on Linux and macOS and historically a legacy code page on Windows — so the same code reads different characters on different machines. Python 3.15 makes UTF-8 the default everywhere, and being explicit remains the habit that survives version differences.
Three things must stay in bytes:
Hashing and cryptography. hashlib.sha256("text") raises; it needs b"text". Digests are defined over bytes, so the encoding is part of the hash. Binary formats. Images, protocol buffers, compressed data — there is no text to recover. Anything length- or offset-sensitive at the wire level, such as a Content-Length header, which counts bytes rather than characters.
bytearray and memoryview
bytearray is the mutable counterpart, useful for building a buffer incrementally without repeated concatenation. memoryview slices without copying, which is what makes parsing large binary payloads efficient — and it has no str equivalent, because a character offset in a variable-width representation is not a fixed byte offset.
Questions people ask
Which way does encode go? str.encode() produces bytes. Text encodes to bytes for transport.
Why does b[0] give a number? bytes is a sequence of integers. Use b[0:1] for a one-byte bytes object.
Why is b"abc" == "abc" False rather than an error? Equality across unrelated types returns False by design; only ordering and concatenation raise. This is why the comparison bug is so easy to miss.
What is the default encoding? UTF-8 for encode/decode. File open used the locale default before Python 3.15, so pass it explicitly.
How do I get the byte length of text? len(s.encode("utf-8")). len(s) counts characters.
Can I put bytes and str keys in one dict? Yes, and they never collide — which is a footgun, not a feature.
What about UTF-16 or Latin-1? Latin-1 is the identity mapping for bytes 0–255, so it never raises on decode — which makes it a tempting and usually wrong way to silence errors.
Recap in one screen
str holds Unicode code points; bytes holds integers 0–255. They never mix implicitly.encode goes str → bytes; decode goes back. UTF-8 unless told otherwise.b[0] is an int; b[0:1] is bytes. And len differs between the two forms of the same text.b"abc" == "abc" is silently False, which is the bug that hides in dictionary lookups.- Decode at the input edge, work in
str, encode on output — and always pass encoding= to open.