Run it
Four ways of producing the same five characters, with identity and equality printed for each. Some of the is results may differ on your interpreter — that is the lesson, not a bug.
What interning is
CPython keeps a table of strings so that identical text can share one object, saving memory and making comparison fast — two interned strings can be compared by pointer.
Which strings get interned is an implementation detail, and roughly:
| String | Interned? |
|---|
| String literals in source code | Usually |
Identifier-like text ("abc", "x_1") | Usually |
| Containing spaces or punctuation | Often not |
| Built at runtime by concatenation | No |
| Read from a file or network | No |
Short: "", single characters | Yes |
The dangerous part is not that is returns False. It is that it returns True often enough to look correct. A test with short literals passes, and production data read from a socket fails the same check.
The identical trap exists for small integers: CPython preallocates −5 to 256, so 256 is 256 is True and 257 is 257 may be False. Same lesson, different type.
Python emits a SyntaxWarning for is against a literal precisely because this is such a reliable source of bugs.
Why it is a trap
The rules are not part of the language. They vary by CPython version, by whether the string was a literal in a compiled block, by whether it contains a space, and by whether the constant folder saw it. The REPL compiles line by line and a script compiles as a unit, so the same code can behave differently in the two.
Code that uses is on strings therefore works in testing, works in the REPL, and fails on the one input that was built by concatenation. That failure looks like a logic error, not a comparison error, which is why it takes so long to find.
The rule, and the exception
Use == for strings, always. Use is only for singletons — None, True, False — where identity is the intended test.
There is one legitimate use of interning: sys.intern(s), called deliberately on a large set of repeated strings such as parsed field names, cuts memory and speeds up dictionary lookups. That is opting in, which is different from relying on a coincidence.
When is is the right operator
Three legitimate uses, and they are the only ones worth defending in review:
Singleton comparison. None, True and False are unique objects, so identity is exactly right — and x is None also cannot be fooled by a class that overrides __eq__:
Sentinel values, where "not supplied" must be distinguishable from any legal value including None:
Genuine identity questions — is this the same object, in caching, cycle detection, or checking whether a mutation aliased something:
Note that is is also the faster operator, since it compares pointers rather than contents. That is never a good enough reason to use it on values — the correctness risk dwarfs the nanoseconds.
== has its own subtleties
| Comparison | Result | Why |
|---|
1 == 1.0 | True | Numeric types compare across types |
1 == True | True | bool subclasses int |
float("nan") == float("nan") | False | IEEE 754 requires it |
[] == () | False | Different types, no cross-comparison |
"café" == "café" | False | Different Unicode normalisation |
Decimal("0.1") == 0.1 | False | Binary float is not exactly 0.1 |
NaN is the pathological case: it is not equal to itself, so x == x can be False. Yet x is x is always True, which means a NaN inside a list is found by in — in checks identity before equality as an optimisation. Use math.isnan(x) to test for it.
Classes may override __eq__ to define whatever equality means for them, and if they do they should also define __hash__ consistently, or instances behave incorrectly in sets and dicts.
Questions people ask
Why does is sometimes work on strings? Interning makes identical literals share one object. It is an optimisation, not a guarantee.
Should I ever use is on strings? Only after sys.intern, and only in performance-critical code where you control every string involved.
Why the SyntaxWarning? Because x is "text" is almost always a bug, and the interpreter can tell.
Is x is None faster than x == None? Yes, and use it because it is correct: == can be redefined by a class, identity cannot.
Do other implementations intern the same strings? No. PyPy and Jython differ, which is another reason not to depend on it.
How do I compare identity in a set? Store id(obj) — and keep a reference to the object, since ids are reused after garbage collection.
What about is not? The negation, and x is not None is the idiomatic form. Not to be confused with not x is None, which parses the same but reads worse.
Recap in one screen
== compares values; is compares object identity.- CPython interns some strings and caches small integers, so
is on values gives inconsistent, version-dependent answers. - The failure mode is that
is often returns True in testing and False on runtime-built data. - Use
is only for None, True, False, sentinels, and real identity questions. float("nan") == float("nan") is False by design, and NaN is still found by in because containment checks identity first.