Shallow vs Deep Copying

Why a copy of a list of lists still shares its inner lists, and when you need copy.deepcopy.

Overview

What a shallow copy actually does

original = [[1, 2], [3, 4]]
shallow = original[:]
shallow[0][0] = 99
print(original)  # [[99, 2], [3, 4]]

The outer list is new — original is shallow is False, and appending to one does not affect the other. But the two inner lists were not copied; both outer lists point at the same two inner lists. Change something one level down and both see it.

original[0] is shallow[0] is True, which is the whole story in one line.

shallow.py

shallow.py Python 3
Output

                    

copy_when.py

copy_when.py Python 3
Output

                    

Worth knowing

A shallow copy duplicates the outer container and reuses everything inside it.
[:], list(x) and x.copy() are all shallow.
Nested mutable data is when you need copy.deepcopy.
If everything inside is immutable, shallow is enough - there is nothing to share.

Shallow vs Deep Copying: A Practical Guide

Copying a list gives you a new list. It does not give you new copies of the things inside it, and that distinction is where the bugs live.

The three shallow copies

nums[:]    list(nums)    nums.copy()

All equivalent. dict.copy() and set.copy() behave the same way, and dict(d) is the dict equivalent of list(l).

When shallow is enough

If everything inside is immutable — numbers, strings, tuples of those — a shallow copy is a complete copy in every way that matters. There is nothing shared that can change, so the distinction disappears.

That covers most everyday copying, which is why [:] is so common and why the problem stays hidden until the day your data has a list inside a list.

When you need deep

import copy
deep = copy.deepcopy(original)

deepcopy walks the whole structure and rebuilds every mutable object it finds. Nested config dictionaries, lists of records, anything parsed from JSON — these are the cases.

The dict version of the trap

config = {"limits": {"max": 10}}
shallow = config.copy()
shallow["limits"]["max"] = 999

The original now reads 999 too. This is the same rule and it bites harder with configuration, because the nesting is the point of the structure.

The rule

Ask what is inside. Flat and immutable: use a slice or .copy(). Nested and mutable: use deepcopy, or restructure so you are not copying a mutable tree at all.

Copying an object, not just a container

copy.copy works on your own classes too, and it copies the same way: a new object whose attributes point at the same values. If an attribute is a list, both copies share it.

Classes can control this. Defining __copy__ and __deepcopy__ lets an object say how it should be duplicated, which matters when it holds something that should not be copied at all - an open file, a database connection, a lock. In practice you rarely write these, but knowing they exist explains why copying some library objects behaves in ways a plain attribute copy would not.

Why deepcopy is slower than it looks

deepcopy walks the entire object graph, duplicates every mutable thing it finds, and keeps a record of what it has already copied so that shared references stay shared and cycles do not become infinite recursion.

That bookkeeping is why it is slow on large structures, and it is also why it is correct in cases a hand-written recursive copy usually gets wrong. If you have ever written a recursive copy helper and hit RecursionError on data that referred back to itself, that is the problem deepcopy already solved.

Cheaper alternatives

Often the goal is not a copy at all, but a modified version of something, without disturbing the original. Several types offer that directly:

new_config = {**defaults, **overrides}    # merged, both untouched
new_items = [*items, extra]               # extended, original untouched
new_point = point._replace(x=5)           # namedtuple

Each of these produces a new outer object and leaves the inputs alone, at the cost of a shallow copy rather than a deep one - which is exactly right when the contents are immutable.

If the contents are mutable and you find yourself needing deepcopy frequently, that is often a signal that the structure would be better as immutable data: tuples, frozensets, or small classes that return new instances rather than mutating themselves.

A practical rule for configuration

Nested dictionaries loaded from JSON or YAML are the single most common place this bites, because the nesting is the whole point and .copy() looks like it did the job.

If a function takes a config dict and changes anything inside it, either it should be documented as doing so, or it should deepcopy first. Silently editing a nested value in a structure the caller still holds is the sort of bug that gets diagnosed as "the settings randomly change", which is a long way from where the mutation actually happened.

Watching the difference with is

The rule is one line of output away from being obvious:

import copy

original = {"a": [1, 2]}
shallow = original.copy()
deep = copy.deepcopy(original)

print(original["a"] is shallow["a"])
print(original["a"] is deep["a"])

shallow["a"].append(3)
print(original, deep)
True
False
{'a': [1, 2, 3]} {'a': [1, 2]}

The first two lines are the whole distinction. After a shallow copy, the inner list is *the same object*; after a deep copy it is a new one. The last line is the consequence: appending through the shallow copy changed the original, and the deep copy was untouched.

Note that the outer dictionaries are different objects in both cases. A shallow copy is a real copy — adding a new key to shallow does not affect original. It is only one level deep, and every problem on this page comes from the level below.

Where this actually comes up

The trap is rarely met head-on. It arrives through four ordinary situations.

Configuration loaded from JSON or YAML. These are nested dictionaries by nature, .copy() looks like it did the job, and a function that adjusts one value has quietly adjusted it for everybody.

A default that is a nested structure. A module-level DEFAULTS dict copied shallowly into each new object means every object shares the same inner dictionaries. The objects look independent and are not.

A list of records. rows[:] gives a new list of the same dictionaries, so sorting or filtering the copy is safe and editing a record through it is not. This is a common one because the copy was made specifically to be safe.

Undo, or "keep the original for comparison". Snapshotting state before a change is the exact case where a shallow copy fails: the snapshot shares the mutable parts with the thing that is about to change, so by the time you compare, both have moved.

The pattern across all four is that the copy was made for safety, and shallow copying provides that safety only for the top level. If the reason for copying is "so that changes over here do not affect over there", the question to ask is how deep the changes go.

Equality survives copying; identity does not

After any copy, shallow or deep, copy == original is True and copy is original is False. That is the intended behaviour and it is worth being explicit about, because it is how you check that a copy did what you meant.

For your own classes it holds only if you defined __eq__. Without one, Python compares by identity, so a copy of your object is *not* equal to the original — which surprises people who have just written a careful __deepcopy__ and find their tests failing on the comparison rather than the copy. Defining __eq__ on a class whose instances get copied around is effectively required, and a dataclass writes it for you.

deepcopy also preserves structure that a naive copy would flatten. If two entries in a dictionary point at the same list, the deep copy has two entries pointing at one *new* list — the sharing is reproduced rather than duplicated. That is usually what you want, and it is one more thing a hand-written recursive copy gets wrong.

Immutability removes the question

The most reliable fix for a copying bug is to not need the copy.

If the structure is built from tuples, frozensets, strings and numbers, then nothing inside it can change, so sharing it is safe and a shallow copy is a complete one. The distinction that this whole page is about simply does not arise. That is why nobody worries about copying a tuple of strings, and why deepcopy on such a structure is wasted work.

The practical version is to make the things that travel immutable. A record passed between functions is better as a frozen dataclass or a NamedTuple than as a dictionary: it cannot be edited by a function that received it, so no defensive copy is needed at any boundary. Where a modified version is required, _replace or dataclasses.replace builds a new one and leaves the original alone.

Keep the mutable structures local, where every name pointing at them is visible in the same function. Then copying becomes a decision you make occasionally rather than a defence you have to remember everywhere.

Deciding, in four questions

When you are about to copy something and are not sure which kind you need, the answer follows from four questions asked in order.

Does anything inside it change? If every item is a number, a string or a tuple of those, any shallow copy is a complete one. Stop here; this covers most copying.

Am I copying so that changes over here do not affect over there? If the answer is no — you are copying to get a list you can sort, or to add one key — then a shallow copy of the outer container is exactly what you want and nothing more is needed.

How deep do the changes go? If the code that follows only touches the top level, shallow is enough regardless of what is nested underneath. If it reaches into nested structures, every level it reaches has to have been copied.

Is the structure large or does it hold something uncopyable? deepcopy on a big object graph costs real time, and on an object holding a connection, a file or a lock it either fails or duplicates something that should be unique. Both are signals to restructure rather than to copy harder — usually by making the shared parts immutable, or by rebuilding the small part you need to change rather than duplicating the whole.

Most real answers land on shallow, and the ones that do not are usually nested configuration or a snapshot taken for comparison. Knowing which of those you are in is more useful than a rule about which function to call.

Questions people ask

Is list(x) a deep copy? No. It is one of the shallow ones, along with x[:] and x.copy().

Does deepcopy handle cycles? Yes. It remembers what it has already copied, so a structure that refers to itself does not cause infinite recursion.

Is deepcopy slow? Relative to a shallow copy, considerably. It walks the whole structure and keeps a record as it goes. For small structures it does not matter.

How do I copy only two levels? There is no built-in for that. A comprehension that shallow-copies each item — {k: v.copy() for k, v in d.items()} — is the usual answer.

Can I stop an attribute being deep-copied? Yes, by defining __deepcopy__ on the class, which is how objects holding a connection or a lock avoid duplicating it.

Does copying a string do anything? No. Strings are immutable, so copy returns the same object.

What about copy.copy versus .copy()? The same thing for builtins. copy.copy also works on objects that have no .copy() method.

Does deepcopy copy class definitions too? No. Classes, functions and modules are treated as atomic and shared rather than duplicated, which is what you want.

Is pickling and unpickling a valid deep copy? It usually produces one, and it is slower and fails on anything unpicklable. Use deepcopy, which is what it is for.

Why does copying a nested list "sometimes" work? Because it works whenever nothing reaches past the top level afterwards. The copy was always shallow; the code simply had not yet touched the shared part.

Should a function copy the arguments it is given? Only if it stores them or mutates them. A function that reads a list and returns a result has no reason to copy anything, and copying defensively everywhere costs more than the bugs it prevents.

What is the fastest way to copy a flat list? x.copy() and x[:] are equivalent and both are fast. list(x) is the one that also accepts any iterable, which is occasionally why it is chosen.

Recap in one screen

  • A shallow copy duplicates the outer container and shares everything inside it; original[0] is shallow[0] is the test.
  • Shallow is a complete copy when the contents are immutable, which is why the problem stays hidden for so long.
  • copy.deepcopy rebuilds the whole structure, preserves shared references and survives cycles — and costs proportionally.
  • The bugs arrive through config dicts, shared defaults, lists of records and snapshots, not through code that obviously copies.
  • Making the data immutable removes the question entirely.

Check yourself

0 of 3

Answer without scrolling back up.

  1. After a shallow copy of `[[1, 2]]`, changing `copy[0][0]`:

  2. When is a shallow copy sufficient?

  3. Which of these is NOT a shallow copy of a list?

Cheat sheet

Shallow vs Deep Copying

The outer list is new — original is shallow is False, and appending to one does not affect the other. But the two inner lists were not copied; both outer lists point at the same two inner lists. Change something one level down and both see it.

PYTHON · vizlearn.in/python/shallow_and_deep_copy.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.