Dict comprehensions
{w: len(w) for w in words}
The colon is the whole difference. Everything before it is the key, everything after is the value, and both are ordinary expressions evaluated per item.
Two patterns come up constantly. Inverting:
{v: k for k, v in prices.items()}
and building from parallel lists:
{k: v for k, v in zip(keys, values)}
though dict(zip(keys, values)) is shorter when there is no transformation to do.
Duplicate keys do not complain
{k: v for k, v in [("a", 1), ("a", 3)]}
gives {"a": 3}. The later value overwrites the earlier one, with no error and no warning. If the input might contain duplicates and you care, that is something to check for, not something Python will tell you about.
The empty-braces trap, again
{} is an empty dict — dictionaries claimed the braces long before sets existed. There is no empty-set literal at all; set() is the only way. This is worth repeating because it is the one inconsistency in an otherwise tidy family.
When to use them
Same rule as list comprehensions: when it fits on a line and reads as a sentence. A dict comprehension with a conditional key expression and a filter is a line you will re-read; a loop is fine, and often kinder.
Inverting and filtering a mapping
Two jobs come up so often that the comprehension form is worth memorising as an idiom rather than derived each time.
Inverting swaps keys and values:
{v: k for k, v in original.items()}
Worth pausing on: this is only lossless if the values were unique. Duplicated values collapse, and the last one processed wins, silently. If that matters, group instead of inverting.
Filtering keeps part of a mapping:
{k: v for k, v in config.items() if v is not None}
That is the standard way to drop unset options before merging configuration, and it reads better than building an empty dict and assigning into it.
Set comprehensions, and what they are for
A set comprehension is a comprehension whose result must be unique, and the two common uses are extracting a distinct field and computing a distinct derived value:
{r["city"] for r in rows}
{len(w) for w in words}
Both could be written as set([...]) around a list comprehension. The direct form avoids building the intermediate list, and says up front that uniqueness is the point rather than an afterthought.
When a loop is still better
The same limit applies as to list comprehensions. One transformation, one optional filter, one line: use the comprehension. As soon as the value needs several steps, a try, or a condition on the key as well as the value, the loop version is easier to read and much easier to change later.
A worked example: a lookup table and a grouping
The two jobs that send people to a dict comprehension, side by side — because only one of them is actually a comprehension:
rows = [("ana", "red"), ("bo", "blue"), ("cy", "red")]
team_of = {name: team for name, team in rows}
print(team_of["cy"])
by_team = {}
for name, team in rows:
by_team.setdefault(team, []).append(name)
print(by_team)
red
{'red': ['ana', 'cy'], 'blue': ['bo']}
The first is a comprehension because each row produces exactly one entry, and later rows are allowed to overwrite earlier ones. The second is a loop because each row *adds to* an entry, and a comprehension cannot do that — every iteration produces an independent key and value, with no access to what the dictionary already holds.
That is the dividing line, and it is worth stating as a test: if the value for a key depends only on the current item, a comprehension works. If it depends on the other items too — a count, a list, a running total — you need a loop, a defaultdict, or a Counter.
Scope, and the variable that does not leak
A comprehension has its own scope, which is why the loop variable does not survive it:
n = "outer"
squares = [n * 2 for n in range(3)]
print(n, squares)
outer [0, 2, 4]
In Python 2 this printed 2, because the comprehension shared the enclosing scope and clobbered n. Python 3 gave comprehensions their own, which removed a whole class of accidental overwrites and is the behaviour to rely on.
Two consequences follow. The variable is unavailable afterwards, so a comprehension cannot be used to compute something you then want to inspect — that needs a loop. And the comprehension can still *read* names from the enclosing scope, which is what makes [x for x in items if x > threshold] work.
The one place this surprises people is inside a class body, where the comprehension's scope cannot see the class's other names. A comprehension in a class body that refers to another class attribute raises NameError, and the fix is to compute it outside the class or pass it in through the iterable.
Sets, and the operations that follow
Building a set with a comprehension is usually the first half of a job whose second half is set algebra, and the two together replace a surprising amount of loop-and-flag code.
{r["city"] for r in rows} gives the distinct cities. Once you have two such sets, a - b is "in the first and not the second", a & b is "in both", and a ^ b is "in one but not both". Those three answer most of the questions people write nested loops for: which records are new, which have disappeared, which are shared.
The comprehension form matters here because it does the deduplication as it goes. Writing set([r["city"] for r in rows]) builds the whole list first and then discards the duplicates, which is more memory for the same answer. For a large input the difference is real, and the direct form also states the intent — uniqueness is the point, not a cleanup step afterwards.
What you give up is order and indexing. A set has neither, so if the result feeds something that cares about sequence, sort it on the way out: sorted({...}) gives a list back in a defined order.
Nesting, and the two conditionals
Comprehensions have two places a condition can appear, and they do different things.
A trailing if filters — items that fail it produce nothing at all: {k: v for k, v in d.items() if v} drops falsy values. A conditional expression in the value slot chooses — every item produces an entry, and the condition picks what goes in it: {k: v or 0 for k, v in d.items()} keeps every key and substitutes a default. Reaching for the wrong one gives you either missing keys or unwanted ones, and the symptom is a dictionary of the wrong size.
Nesting a second for is legal and reads in the order written, outermost first: {c for row in grid for c in row} collects every cell of a grid of rows. The rule is that the clauses appear in the same order as the equivalent nested loops, which is the opposite of the order people guess when they meet it in someone else's code.
Both features are worth knowing and neither is worth combining. A comprehension with two for clauses, a filter and a conditional value is a line that has to be decoded rather than read, and the loop it replaces would have been four clear lines.
What is built, and when it matters
The four bracket forms differ in more than the type of the result: they differ in what exists in memory while they run, and occasionally that is the deciding factor.
A list, set or dict comprehension builds the whole result before anything else happens. For a thousand items that is irrelevant. For ten million it is the difference between a program that runs and one that does not, and the generator expression — round brackets — is the version that holds one item at a time.
The rule of thumb is what happens to the result. If it is consumed once, by a sum, a max, a for, or an any, a generator expression does the same job without materialising anything: sum(x * 2 for x in items) never builds a list. If it is indexed, iterated twice, measured with len, or returned to a caller who will do any of those, it has to be a real collection.
There is one case where building the intermediate is a genuine waste and easy to miss: wrapping a list comprehension in a converter. set([f(x) for x in items]) and dict([(k, f(k)) for k in keys]) both build a list and immediately throw it away. Writing the set or dict comprehension directly skips that entirely, which is one reason those forms exist at all.
The reverse mistake is reaching for a generator expression where a list is wanted, and then calling list() on it. That builds the same list through more machinery, and the comprehension says it more directly.
Readability, stated as a limit
Every section here has hinted at a boundary, and it is worth making explicit, because "use a comprehension when it is clear" is advice that gives no help at the moment of writing.
A practical limit: one for, at most one if, and an expression short enough that the whole thing fits on one line without wrapping. That covers the large majority of genuine uses and produces something a reader takes in at a glance.
Past that, the questions to ask are whether a reader could say what the result contains without tracing it, and whether it can be changed without being rewritten. A comprehension that needs a comment above it to explain what it produces has already answered both.
The escape hatch is usually a named function. Pulling the expression out into a function with a meaningful name leaves a comprehension that is one for and a call — readable again, and the complicated part now has a name, a place to put a docstring, and somewhere to test it.
Questions people ask
Why does {} give a dict and not a set? Dictionaries had the braces first. set() is the only way to write an empty set.
Can I build a frozenset with a comprehension? Not directly. frozenset(x for x in items) wraps a generator expression.
What happens to duplicate keys? The last one wins, silently. Nothing warns you.
Is a dict comprehension faster than a loop? Slightly, because the whole thing runs in one bytecode loop without repeated attribute lookups. Not enough to be the reason to choose it.
Can I use zip inside one? Yes, and {k: v for k, v in zip(keys, values)} is common — though dict(zip(keys, values)) is shorter when nothing is transformed.
Does the order of a dict comprehension matter? Yes. The result keeps insertion order, so the order of the input decides the order of the output.
Can I have an else without an if filter? Only in the value slot, as part of a conditional expression. The trailing filter has no else.
Can a comprehension call a function with side effects? It can, and it should not. A comprehension whose result is discarded is a loop written to look like an expression.
Do comprehensions work in older Python? List comprehensions since 2.0, dict and set comprehensions since 2.7 and 3.0. Anything current supports all of them.
Recap in one screen
- The brackets choose the type:
[] list, {} set, {k: v} dict, () generator. Everything else about the syntax is identical. - A comprehension works when each item produces one independent entry; counting and grouping need a loop,
setdefault or Counter. - Duplicate keys collapse silently, and so do duplicate set members — that is the point of a set and a hazard in a dict.
- The loop variable has its own scope and does not leak, except that it cannot see a class body's other names.
- A trailing
if filters items out; a conditional expression in the value slot keeps them and changes the value.