Why does `is` sometimes work on strings?

CPython interns short string literals that look like identifiers, so two of them share one object and is happens to be True. Build the same text at runtime and it is False. It is an optimisation, not a guarantee — compare strings with ==, and keep is for None.

Overview

The two operators

== compares values. is compares identity — whether two names refer to the same object in memory.

1Python
Output

That last pair is the whole problem: is on strings gives answers that depend on interpreter internals, so code relying on it works in testing and fails somewhere else.

The rule is short: use == for values, and is only for None, True, False, and genuine identity checks.

StringsConceptualMedium

Step through it

What to watch

  • Both rows hold the same text; only how it was built differs.
  • == is True in every frame. is is not.
  • Behaviour differs between the REPL and a script — which is the point.

Say this out loud

"That's interning - CPython reuses one object for short literals. It's an implementation detail, so I compare with == and only use `is` for None."

Why does `is` sometimes work on strings?

Why does 'a' is 'a' return True, and why should you never rely on it?

Run it

Four ways of producing the same five characters, with identity and equality printed for each. Some of the is results may differ on your interpreter — that is the lesson, not a bug.

2Python
Output

What interning is

CPython keeps a table of strings so that identical text can share one object, saving memory and making comparison fast — two interned strings can be compared by pointer.

Which strings get interned is an implementation detail, and roughly:

StringInterned?
String literals in source codeUsually
Identifier-like text ("abc", "x_1")Usually
Containing spaces or punctuationOften not
Built at runtime by concatenationNo
Read from a file or networkNo
Short: "", single charactersYes
3Python
Output

The dangerous part is not that is returns False. It is that it returns True often enough to look correct. A test with short literals passes, and production data read from a socket fails the same check.

The identical trap exists for small integers: CPython preallocates −5 to 256, so 256 is 256 is True and 257 is 257 may be False. Same lesson, different type.

4Python
Output

Python emits a SyntaxWarning for is against a literal precisely because this is such a reliable source of bugs.

Why it is a trap

The rules are not part of the language. They vary by CPython version, by whether the string was a literal in a compiled block, by whether it contains a space, and by whether the constant folder saw it. The REPL compiles line by line and a script compiles as a unit, so the same code can behave differently in the two.

Code that uses is on strings therefore works in testing, works in the REPL, and fails on the one input that was built by concatenation. That failure looks like a logic error, not a comparison error, which is why it takes so long to find.

The rule, and the exception

Use == for strings, always. Use is only for singletons — None, True, False — where identity is the intended test.

There is one legitimate use of interning: sys.intern(s), called deliberately on a large set of repeated strings such as parsed field names, cuts memory and speeds up dictionary lookups. That is opting in, which is different from relying on a coincidence.

When is is the right operator

Three legitimate uses, and they are the only ones worth defending in review:

Singleton comparison. None, True and False are unique objects, so identity is exactly right — and x is None also cannot be fooled by a class that overrides __eq__:

5Python
Output

Sentinel values, where "not supplied" must be distinguishable from any legal value including None:

6Python
Output

Genuine identity questions — is this the same object, in caching, cycle detection, or checking whether a mutation aliased something:

7Python
Output

Note that is is also the faster operator, since it compares pointers rather than contents. That is never a good enough reason to use it on values — the correctness risk dwarfs the nanoseconds.

== has its own subtleties

ComparisonResultWhy
1 == 1.0TrueNumeric types compare across types
1 == TrueTruebool subclasses int
float("nan") == float("nan")FalseIEEE 754 requires it
[] == ()FalseDifferent types, no cross-comparison
"café" == "café"FalseDifferent Unicode normalisation
Decimal("0.1") == 0.1FalseBinary float is not exactly 0.1

NaN is the pathological case: it is not equal to itself, so x == x can be False. Yet x is x is always True, which means a NaN inside a list is found by in — in checks identity before equality as an optimisation. Use math.isnan(x) to test for it.

Classes may override __eq__ to define whatever equality means for them, and if they do they should also define __hash__ consistently, or instances behave incorrectly in sets and dicts.

Questions people ask

Why does is sometimes work on strings? Interning makes identical literals share one object. It is an optimisation, not a guarantee.

Should I ever use is on strings? Only after sys.intern, and only in performance-critical code where you control every string involved.

Why the SyntaxWarning? Because x is "text" is almost always a bug, and the interpreter can tell.

Is x is None faster than x == None? Yes, and use it because it is correct: == can be redefined by a class, identity cannot.

Do other implementations intern the same strings? No. PyPy and Jython differ, which is another reason not to depend on it.

How do I compare identity in a set? Store id(obj) — and keep a reference to the object, since ids are reused after garbage collection.

What about is not? The negation, and x is not None is the idiomatic form. Not to be confused with not x is None, which parses the same but reads worse.

Recap in one screen

  • == compares values; is compares object identity.
  • CPython interns some strings and caches small integers, so is on values gives inconsistent, version-dependent answers.
  • The failure mode is that is often returns True in testing and False on runtime-built data.
  • Use is only for None, True, False, sentinels, and real identity questions.
  • float("nan") == float("nan") is False by design, and NaN is still found by in because containment checks identity first.

How the code works

Four ways of producing the same five characters, with identity and equality printed for each. Some of the is results may differ on your interpreter — that is the lesson, not a bug.

How the code works

  1. a is bTwo identical literals in the same compiled block share one object. This is the result people generalise from, and it is the narrowest case.
  2. d = "".join(parts)Built while the program runs, so the interner never sees it. Same characters, different object, is is False.
  3. e is gThe space is the difference. A literal with a space does not look like an identifier, so the rules that produced True above do not apply.
  4. sys.intern(d)The legitimate use: deliberately deduplicating a large set of repeated strings to cut memory and speed up dict lookups. Opting in is not the same as relying on an accident.

Change one thing

  • Move the comparisons into a function and call it. Compiling as one unit can change the answers — which is exactly why this is not something to build on.
  • Try 257 is 257 across a function boundary. Integers have the same cache, the same trap and the same rule.

Where this runs

Real CPython, compiled to WebAssembly and running on your own machine — nothing is uploaded. The first run takes a few seconds while the interpreter downloads; after that it is immediate. Need more room, or want to paste your own attempt? Use the Python compiler.

Check yourself

0 of 3

Answer without scrolling back up.

  1. Why is 'hello' is 'hello' often True?

  2. 'hello world' is ('hello' + ' world') is usually False because:

  3. The only safe use of `is` is with:

Cheat sheet

Why does `is` sometimes work on strings?

CPython interns short string literals that look like identifiers, so two of them share one object and is happens to be True. Build the same text at runtime and it is False. It is an optimisation, not a guarantee — compare strings with ==, and keep is for None.

INTERVIEW · vizlearn.in/interview/string-interning-and-the-is-operator.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.