lambda, map and filter

Small anonymous functions, the two builtins that take them, and why a comprehension usually reads better.

Overview

What a lambda is

square = lambda n: n * n

That is the same as a two-line def, minus the name. The body is a single expression and its value is returned automatically — there is no return, and no room for a statement. lambda n: print(n); return n is a syntax error.

Assigning a lambda to a name, as above, is the one usage style guides actively discourage: if it deserves a name it deserves a def, which also gives it a useful name in tracebacks.

lambda.py

lambda.py Python 3
Output

                    

map_filter_lazy.py

map_filter_lazy.py Python 3
Output

                    

Worth knowing

A lambda holds one expression and returns it. No statements, no return.
map and filter are lazy - wrap in list() to see them.
map(int, words) needs no lambda. map(lambda w: int(w), words) adds nothing.
A comprehension usually reads better; sorted(key=...) is where lambdas shine.

lambda, map and filter: A Practical Guide

A lambda is a function written inline as an expression. map and filter are the two builtins that take one. All three are worth knowing, and in most Python code a comprehension is the better choice.

Where a lambda earns its place

sorted(people, key=lambda p: p[1])

As an argument, where the function is tiny, used once, and naming it would add a line without adding meaning. sorted, min, max and groupby are where you will actually write them.

map and filter

map(f, items)     # apply f to each
filter(f, items)  # keep those where f is true

Both are lazy: they return iterators, not lists. Printing one shows <map object>, and consuming it twice gives nothing the second time — the page demonstrates exactly that, because it catches people.

The comprehension equivalents are usually clearer:

[n * n for n in nums]       # instead of map
[n for n in nums if n % 2 == 0] # instead of filter

They read left to right, they do not need list() around them, and they do not need a lambda at all.

When map is genuinely better

Two cases. First, when the function already exists:

map(int, words)

No lambda, no comprehension variable — just "convert each of these". Writing map(lambda w: int(w), words) instead is pure ceremony.

Second, when consuming several sequences in parallel:

map(lambda x, y: x + y, a, b)

though [x + y for x, y in zip(a, b)] says the same thing and most readers will parse it faster.

The judgement

Know all three, because you will read them. Write comprehensions by default, map(existing_function, items) when it is that shape exactly, and lambdas mainly as key= arguments.

Why the style guide discourages naming a lambda

square = lambda n: n * n works and is discouraged, for a concrete reason rather than an aesthetic one. A def gives the function a name in its own metadata, so a traceback says square instead of <lambda>. When something raises three frames deep, that difference is the gap between reading the traceback and guessing.

def also allows a docstring, type annotations, multiple statements and a default argument - all of which you will eventually want, and adding any of them means converting anyway.

The rule that follows: if it needs a name, it needs a def.

Closures in a loop

A classic trap, and it looks like a lambda problem while actually being a scope one:

funcs = [lambda: i for i in range(3)]
print([f() for f in funcs])      # [2, 2, 2]

Every lambda captured the *variable* i, not its value at the time, and by the time they run the loop has finished with i at 2. Binding the value explicitly as a default argument fixes it:

funcs = [lambda i=i: i for i in range(3)]   # [0, 1, 2]

This is worth recognising because it appears whenever functions are built in a loop - event handlers, callbacks, partial applications - and the symptom is always the same: they all behave like the last one.

functools.partial

When the goal is to fix some arguments of an existing function, partial says it more directly than a lambda:

from functools import partial
to_int = partial(int, base=16)

It also produces something introspectable and picklable, which a lambda is not - relevant the moment multiprocessing is involved, because lambdas cannot be sent to worker processes.

Reading functional code

map and filter are worth knowing well even if you write comprehensions, because a great deal of existing Python uses them, and because they compose with the rest of itertools. Being able to read filter(None, values) - which drops falsy items, using None as "no function" - is more useful than an opinion about whether you would have written it that way.

operator, the module that replaces most lambdas

A large share of the lambdas people write do one of two things: pull out an item, or pull out an attribute. The standard library has both, and they are faster and more readable than the lambda:

from operator import itemgetter

people = [("ana", 91), ("bo", 78), ("cy", 91)]

print(sorted(people, key=itemgetter(1), reverse=True))
print(sorted(people, key=itemgetter(1, 0)))
[('ana', 91), ('cy', 91), ('bo', 78)]
[('bo', 78), ('ana', 91), ('cy', 91)]

itemgetter(1) is lambda p: p[1] with a name that says what it does, and itemgetter(1, 0) returns a tuple, which is how you sort by score and then by name without writing the tuple out. attrgetter is the same for objects, and methodcaller("lower") for calling a method on each item.

Note the first result: ana comes before cy even though they tie, because sorted is stable and ana was first in the input. That guarantee is what makes sorting twice for a two-level sort work at all.

reduce, and why it is not a builtin any more

map and filter survived the move to Python 3 as builtins. reduce did not, and lives in functools:

from functools import reduce

nums = [1, 2, 3, 4]
print(reduce(lambda a, b: a * b, nums))
print(sum(nums), max(nums), any(n > 3 for n in nums))
24
10 4 True

The reason for the demotion is on the second line. Nearly every real use of reduce is a sum, a maximum, a minimum, an any or an all, and each of those has a builtin that says what it means at a glance. What is left — a running product, folding a custom combine function over a sequence — is rare enough to be worth an import and a moment's thought from the reader.

If you do reach for it, pass the initial value: reduce(f, items, 0) returns the initial value for an empty sequence instead of raising TypeError.

Where map and filter lead: itertools

map and filter are the first two of a family. itertools holds the rest, and they share the same property of producing items on demand rather than building lists:

from itertools import islice, chain, takewhile

nums = range(1, 11)

print(list(islice(nums, 3)))
print(list(takewhile(lambda n: n < 4, nums)))
print(list(chain([1, 2], [3, 4])))
[1, 2, 3]
[1, 2, 3]
[1, 2, 3, 4]

islice takes the first few of anything, including an infinite generator. takewhile stops at the first item that fails the test, which is different from filter — filter would carry on and return 1, 2, 3 from the whole range while takewhile stops looking at 4. chain walks several iterables as one without building a combined list.

The laziness is the point of all of them. A pipeline of map, filter and islice over a large file reads only as far as it needs to.

The three-line version of a common job

Take rows, keep the ones that qualify, transform them, and summarise. Written with comprehensions, which is what most Python uses:

rows = [("ana", 91), ("bo", 55), ("cy", 78), ("di", 43)]

passed = [name for name, score in rows if score >= 60]
best = max(rows, key=lambda r: r[1])

print(passed)
print(best)
print(f"{len(passed)}/{len(rows)} passed")
['ana', 'cy']
('ana', 91)
2/4 passed

One lambda appears, as a key= argument, which is exactly the place the earlier section said it belongs. The filtering and the transforming are done by the comprehension rather than by filter and map, and the result reads in the order it happens.

What "functional" means here, and what it does not

lambda, map and filter arrive with a reputation attached, and it is worth separating the useful part from the folklore.

The useful part is that functions are ordinary values in Python. A function can be stored in a list, passed as an argument, returned from another function, and kept in a dictionary. That is what makes key= arguments, decorators, callbacks and dispatch tables possible, and it is the single idea behind everything on this page. A lambda is not a special kind of function; it is the same object a def produces, written inline and left unnamed.

The folklore is that using them makes code functional, and that functional code is better. Python is not a functional language and does not try to be. It has no tail-call optimisation, its lambdas are deliberately limited to one expression, and reduce was moved out of the builtins precisely to discourage the style. Guido van Rossum's stated preference was that a comprehension or a plain loop says the same thing more clearly to more readers.

What actually transfers from functional programming, and is worth adopting, is smaller and less glamorous: prefer functions that return a new value over functions that modify their arguments; avoid depending on state that is not visible in the signature; and keep functions small enough to test on their own. None of that requires a lambda. A page of def statements can follow all three, and a chain of nested map and filter calls can violate them all while looking the part.

When a comprehension stops being the clearer choice

The advice throughout this page is to reach for a comprehension first, so it is worth being precise about where that stops holding.

A comprehension is doing too much when it has more than one for and a condition, when the expression at the front needs its own explanation, or when the whole thing no longer fits on one line without awkward wrapping. At that point it has become a loop that is harder to read than the loop would have been, and the honest move is to write the loop — or to name the inner step as a function and call it from a simple comprehension.

The other case is when you need something a comprehension cannot do: a break, an early return, a try/except around one item, or anything that has to happen for its side effect rather than its value. A comprehension built purely for side effects, with its result thrown away, is a loop wearing the wrong clothes; write for and let the reader see what is happening.

Nesting deserves its own warning. A comprehension inside another comprehension's expression is read outer-first and reasoned about inner-first, which is exactly the sort of thing that is fine while you are writing it and unpleasant three months later. Two statements are almost always better than one clever one.

Naming things the reader will meet

Two words appear constantly in discussions of this material and are rarely defined, which makes a lot of otherwise good explanations hard to follow.

A higher-order function is a function that takes a function as an argument or returns one. sorted is higher-order because of key=; so are map, filter, and every decorator you will write. There is nothing more to the term than that.

A closure is a function that remembers a value from the scope where it was created, and keeps working after that scope has finished. The loop trap earlier on this page is a closure behaving exactly as specified and not as expected: the lambdas remembered the variable rather than the value it held at the time.

Questions people ask

Can a lambda have more than one statement? No. The body is a single expression. If you need a statement, you need a def.

Can a lambda have default arguments? Yes — lambda x, n=2: x ** n — and that is also the trick for capturing a loop variable by value.

Why does printing map(...) show an object? Because it is a lazy iterator. Wrap it in list() to see the items.

Is a comprehension faster than map? They are close. map(int, items) with an existing function is usually slightly faster; map(lambda ..., items) is usually slightly slower. Neither difference is a reason to choose one.

What does filter(None, items) do? Drops every falsy item. None in the function slot means "use the item's own truthiness".

Why can I not pickle a lambda? Pickling stores a reference by name and a lambda has none. That is why multiprocessing rejects them, and where functools.partial earns its place.

Do map and filter work on dictionaries? They work on anything iterable, and iterating a dictionary gives its keys. Use d.items() when you want pairs.

Recap in one screen

  • A lambda is a single-expression function; use one as a key= argument and write a def whenever it deserves a name.
  • map and filter are lazy iterators — consume them once, wrap in list() to keep them.
  • Prefer a comprehension by default; map(existing_function, items) when it is exactly that shape.
  • operator.itemgetter and attrgetter replace the most common lambdas and read better.
  • reduce lives in functools because sum, max, any and all already cover almost every use of it.

Check yourself

0 of 3

Answer without scrolling back up.

  1. What can a lambda body contain?

  2. `list(m)` twice on a map object gives what the second time?

  3. Which is preferred: `map(lambda w: int(w), ws)` or `map(int, ws)`?

Cheat sheet

lambda, map and filter

A lambda is a function written inline as an expression. map and filter are the two builtins that take one. All three are worth knowing, and in most Python code a comprehension is the better choice.

PYTHON · vizlearn.in/python/lambda_map_filter.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.