split and join
parts = "a,b,c".split(",") # ['a', 'b', 'c']
",".join(parts) # 'a,b,c'
join is called on the separator, not on the list, which reads backwards until you have seen it a few times. Think of it as "put this between them".
join requires strings. A list of numbers raises TypeError, so convert first: ",".join(str(n) for n in nums).
split() with no argument is a different function in practice: it splits on any run of whitespace and discards empties, which is what you want for scruffy text. split(" ") splits on each single space and will hand you empty strings between doubled spaces.
strip removes characters, not a suffix
"banana".strip("ab")
removes any leading or trailing a or b — it does not remove the string "ab". The argument is a set of characters. This trips people who write filename.strip(".csv") and find it also ate a trailing s or v.
For that job:
"report.csv".removesuffix(".csv")
removeprefix and removesuffix were added in Python 3.9 precisely because the strip misuse was so common.
Tests that read as English
startswith, endswith, isdigit, isalpha all return booleans and read naturally in a condition. endswith accepts a tuple, so name.endswith((".jpg", ".png")) is one call rather than two comparisons.
Case methods and comparison
upper, lower and title return new strings. For case-insensitive comparison, lower() both sides — or casefold(), which handles a few non-English cases lower does not.
Text arriving from a person or a file usually needs the same three steps, and the order is worth doing deliberately.
Strip first, because trailing whitespace and newlines break comparisons in ways that are invisible on screen. Normalise case next, when the comparison should ignore it. Only then split or compare.
line.strip().lower().split(",")
Chaining works because each method returns a new string, which is the same immutability that makes forgetting to assign such a common mistake elsewhere.
For case-insensitive comparison across languages, casefold is stricter than lower and handles a handful of cases lower does not. For English-only text the difference never shows up.
Searching without exceptions
Three ways to ask whether something appears in a string, and they differ in what they give back:
"cat" in text # True or False
text.find("cat") # index, or -1
text.index("cat") # index, or raises ValueError
Use in when you only want to know. Use find when you want the position and absence is expected. Use index when absence is a bug and you want it to say so. Choosing find and then forgetting that -1 is falsy in a different way from 0 is a small classic: if text.find("cat"): is true when the match is at position one and false when it is at position zero.
Splitting text into lines and fields
splitlines() handles the newline variants that split("\n") gets wrong on files written on another operating system. partition splits once and always returns three pieces, which avoids the length check that split with a maxsplit requires:
key, sep, value = line.partition("=")
If the separator was absent, sep is empty and value is empty, and the unpacking still works - no IndexError, no branch.
Building strings in a loop
Repeated concatenation inside a loop creates a new string every time, because strings cannot be modified in place. For a handful of pieces that is irrelevant. For thousands it is genuinely slow, and the fix is to collect and join once:
parts = []
for row in rows:
parts.append(format(row))
text = "\n".join(parts)
This is the same reason join exists as a method on the separator: it is building the whole result in one pass, which it can only do if it is given all the pieces at once.
replace, and the count nobody passes
replace swaps every occurrence and, like everything else here, returns a new string:
text = "a-b-c-d"
print(text.replace("-", " "))
print(text.replace("-", " ", 2))
print(text)
a b c d
a b c-d
a-b-c-d
The third argument caps how many are replaced, counting from the left, and it is the one people forget exists — it saves a regular expression surprisingly often. The last line is the reminder that text itself never changed.
count answers the related question without doing the work: text.count("-") is 3 here, and returns 0 rather than raising when there are none.
A worked example: cleaning one messy line
Real text arrives with the wrong spacing, the wrong case, and a newline on the end. The methods on this page compose into a single readable pass:
raw = " Ana , 91 , Physics\n"
name, score, subject = [f.strip() for f in raw.strip().split(",")]
print(repr(name), int(score), repr(subject))
'Ana' 91 'Physics'
Two strips are doing different jobs and both are needed. The outer raw.strip() removes the leading spaces and the trailing newline from the line as a whole. The one inside the comprehension removes the spaces around each field, which the outer one could not reach. Splitting first and stripping each piece is the general shape; stripping only the whole line leaves " 91 ", which int() happens to forgive and a string comparison does not.
repr() in the output is there deliberately. Printing a string that still has a stray space looks identical to one that does not, and repr is how you see the difference.
Padding and aligning without counting spaces
rows = [("ana", 91), ("bo", 7), ("caroline", 143)]
for name, n in rows:
print(name.ljust(10, ".") + str(n).rjust(4))
ana....... 91
bo........ 7
caroline.. 143
ljust, rjust and center take a width and an optional fill character. zfill is the special case for numbers: "7".zfill(3) gives "007", and it handles a leading minus sign correctly, which rjust(3, "0") does not.
The f-string spellings are f"{name:<10}", f"{n:>4}" and f"{n:^4}", and they are usually the better choice because the value and its formatting stay together. The methods are worth knowing for when the width is computed rather than literal.
The is-something tests, and the one that lies
print("42".isdigit(), "4.2".isdigit(), "-4".isdigit())
print("42".isdecimal(), "\u00b2".isdigit(), "\u00b2".isdecimal())
True False False
True True False
isdigit is not "would int() accept this". It rejects a decimal point and a minus sign, and it accepts superscripts and other numeric characters that int() refuses. isdecimal is the stricter one and the closer match to what people usually mean.
The reliable test for "is this a number" is to try the conversion:
def is_int(text):
try:
int(text)
except ValueError:
return False
return True
That handles the minus sign, the surrounding whitespace and the underscores Python allows in numeric literals, all of which the is methods get wrong in one direction or the other. isalpha, isspace and isupper have no such problem and are safe to use directly.
Why strings are immutable at all
Beginners meet immutability as an inconvenience — the reason name.upper() seems not to work. It is worth knowing what the language buys with it, because the same property explains behaviour in several other places.
A string that cannot change can be shared without anyone worrying about who else holds it. Python takes advantage of that constantly: identical short strings in your source are often the same object in memory, string constants can be stored once and pointed at from many places, and passing a string to a function costs nothing regardless of its length, because nothing is copied.
It also makes strings hashable, and therefore usable as dictionary keys and set members. A hash is computed from the contents; if the contents could change, the key would end up filed under a hash that no longer matches it, and lookups would silently fail. This is exactly why lists cannot be keys and tuples can. Every dictionary you index by name relies on strings being immutable.
The cost is real and narrow: building a string by repeated concatenation is quadratic, because each step copies everything so far. That is the one place where the design bites, and join exists to cover it. Everywhere else, the guarantee that a string you were handed is the string you still have is worth considerably more than the ability to edit one in place.
Unicode, and what a character actually is
A Python string is a sequence of Unicode code points, not bytes. For English text the distinction never comes up, and outside English it comes up immediately.
len("café") is 4, which is what you would hope. But the same text can be written two ways in Unicode — as an é code point, or as e followed by a combining accent — and in the second form len reports 5 while the text looks identical on screen. unicodedata.normalize("NFC", text) collapses the second form into the first, and normalising before comparing is the fix for "these two strings look the same and are not equal".
Emoji push this further: many are several code points joined together, so slicing a string by index can cut one in half and produce something that will not render. If you are truncating text that might contain them, count grapheme clusters with a library rather than characters.
Encoding is the separate question of how those code points become bytes for a file or a network. text.encode("utf-8") produces bytes; data.decode("utf-8") goes back. The errors people hit — UnicodeDecodeError on reading a file, mojibake like é where é should be — are almost always a file written in one encoding and read in another. UTF-8 is the right default everywhere, and specifying it explicitly when opening files is worth the few extra characters, because the default still varies by platform.
When to stop using string methods
The methods on this page handle a great deal, and there is a point past which they stop being the right tool. The signal is a chain of split, strip, find and slicing that is getting longer each time the input surprises you.
If the pattern is genuinely variable — optional whitespace, alternatives, repetition, groups you want to capture — that is what regular expressions are for, and re.search with a named group will be shorter and easier to fix than eight lines of index arithmetic. The trade is that a regular expression is harder to read for anyone who does not write them often, so it earns its place only when the string methods have genuinely run out.
If the text is a known format, do not parse it by hand at all. JSON, CSV, INI files, dates, URLs and email addresses all have a module in the standard library that has already handled the edge cases you have not thought of yet: quoted commas inside a CSV field, escaped characters in JSON, timezone suffixes, percent-encoding. Hand-rolled parsing of a standard format is the most reliable way to produce a program that works on your test file and fails on real data.
The rule that follows is a simple one. String methods for cleaning and simple splitting, a parser for anything with a specification, and regular expressions for the genuinely irregular middle ground.
Questions people ask
Why does strip() not remove a suffix? Because its argument is a set of characters to remove from each end, not a string to match. Use removesuffix for the other job.
Is + on strings actually slow? Not for a few pieces. It is slow inside a loop over thousands, because each + builds a whole new string. Collect into a list and join once.
What is the difference between find and index? find returns -1 when absent, index raises ValueError. Pick based on whether absence is expected or a bug.
How do I split on more than one separator? re.split from the re module — str.split takes one separator only.
Why does "".join(numbers) fail? join requires strings. Convert first: "".join(str(n) for n in numbers).
Does lower() work for every language? Mostly. casefold() is the more aggressive version intended for caseless comparison, and it is what to use when the text is not guaranteed to be English.
How do I check a string against several endings? endswith accepts a tuple: name.endswith((".jpg", ".png")).
Recap in one screen
- Every string method returns a new string; if you do not assign the result, nothing happened.
join is called on the separator and needs strings; split() with no argument is the forgiving version.strip takes characters, not a suffix — removeprefix and removesuffix exist for that.find returns -1, index raises; in is the one to use when you only need a yes or no.isdigit is not "is a number" — try the conversion instead.