TF-IDF

Two numbers multiplied. Term frequency rewards a word for appearing often in this document; inverse document frequency punishes it for appearing in every document. The product is high only for words that are frequent here and rare elsewhere — which is a workable definition of "what this document is about".

Overview

The two halves

Term frequency is how often a term occurs in a document, usually damped — a word appearing twenty times is not twenty times as relevant, so implementations take a logarithm or normalise by document length.

Inverse document frequency is log(N / df), where df is the number of documents containing the term. When a term is in every document that ratio is 1 and the log is 0, so the term contributes nothing at all.

That zero is the elegant part: stopwords are removed by the arithmetic rather than by a list you have to maintain per language.

Parameters

Visualisation

—

Readout

What to watch

  • A term in every document scores zero — the log sees to that.
  • One occurrence of a rare word beats two of a common one.
  • No stopword list is needed; the maths removes them.

TF-IDF: A Practical Guide

How does TF-IDF decide which words in a document actually matter?

Where it falls short

No length normalisation by default. A long document contains more of everything, so raw TF favours it. Dividing by length overcorrects and favours very short ones. BM25's b parameter exists to tune between the two.

Unbounded term frequency. TF keeps growing with repetition. BM25 saturates it, so the tenth occurrence adds far less than the second — which matches how relevance actually behaves.

No understanding of meaning. "car" and "automobile" are unrelated terms. That is what dense embeddings fix, and why modern retrieval runs both and fuses the results — see hybrid search.

Why it still matters in a RAG system

Exact terms still win on identifiers, error codes, product names and rare jargon — precisely the queries where an embedding model has seen too little to place the token meaningfully. A vector-only pipeline reliably fails on "error TS2345".

TF-IDF is also the honest baseline. If a dense retriever cannot beat it on your evaluation set, the problem is the embedding model or the chunking, not the ranking.

Working one score out by hand

Take a four-document corpus where cat appears in two documents and the in all four. For the query cat:

idf(cat) = log(4 / 2) = 0.69
idf(the) = log(4 / 4) = 0.00

A document containing cat twice scores (1 + log 2) × 0.69 = 1.17. A document containing it once scores 0.69. And a document containing the ten times still scores zero for that term, because anything multiplied by zero is zero.

That last line is the one worth internalising. The stopword is not filtered, thresholded or special-cased anywhere. It is removed because the logarithm of one is zero, which is a much more satisfying reason than a hand-maintained list, and it adapts automatically to a corpus where patient or invoice is effectively a stopword.

The variants you will actually meet

Sublinear tf. 1 + log(count) rather than the raw count, which is what the visualisation's damping toggle switches. Almost always on.

Smoothed idf. log(1 + N/df) or log((N+1)/(df+1)) + 1, which scikit-learn uses by default. It avoids a zero weight and a division by zero for unseen terms, and it means scikit-learn's numbers will not match a textbook's.

L2 normalisation. Each document vector scaled to unit length, so cosine similarity between documents is a dot product and long documents stop winning by having more of everything.

When someone says two TF-IDF implementations disagree, it is nearly always one of these three, not a bug.

Where it sits in a modern stack

TF-IDF is rarely the ranker any more — BM25 supersedes it by saturating term frequency and normalising length properly — but it is still the thing to reach for in three situations.

As a baseline. If a dense retriever cannot beat TF-IDF on your evaluation set, the problem is the embedding model, the chunking or the evaluation set itself. It costs minutes to run and it has saved a great many people from tuning the wrong thing.

As a feature. TF-IDF vectors feed classical classifiers — spam filtering, topic labelling, near-duplicate detection — where a sparse interpretable representation beats a dense one and trains in seconds.

As the explanation. Every score decomposes into per-term contributions, so you can say exactly why a document ranked where it did. No dense retriever can do that, and in regulated settings it is sometimes the deciding factor.

Weighting words by how informative they are

Counting words treats "the" and "encephalopathy" as equally interesting. TF-IDF corrects that by multiplying two terms:

TF-IDF(t, d) = TF(t, d) × log(N / DF(t))

Term frequency — how often the term appears in this document. More mentions, more relevance.

Inverse document frequency — the logarithm of the total number of documents divided by how many contain the term. Rare terms score high; ubiquitous terms score near zero.

Worked through with 1,000 documents:

TermDocuments containing itIDF
"the"1,000log(1) = 0
"policy"200log(5) = 1.61
"parental"30log(33) = 3.50
"encephalopathy"2log(500) = 6.21

A term in every document contributes exactly nothing, whatever its frequency. That is automatic stop-word removal, derived from the corpus rather than from a list — and it is the single most useful property of the method.

Why the logarithm

Without it, a term in 1 document out of 1,000 would get a weight of 1,000 and a term in 10 would get 100 — a tenfold difference that vastly overstates how much more informative the first is.

The logarithm compresses that: 6.9 against 4.6. Rare terms still win, and not by an absurd margin.

The same reasoning motivates sublinear term frequency: using 1 + log(TF) instead of the raw count, because a term appearing 50 times is not 50 times more relevant than one appearing once. Scikit-learn exposes this as sublinear_tf=True, and it usually helps.

Two more refinements that matter in practice:

L2 normalisation of each document's vector, so long documents do not dominate simply by containing more words. Applied by default in most implementations.

Smoothing the IDF denominator (log(1 + N/DF)) so a term appearing in no documents does not divide by zero.

Building a working baseline

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline

model = make_pipeline(
    TfidfVectorizer(
        ngram_range=(1, 2),      # unigrams and bigrams
        min_df=3,                # drop terms in fewer than 3 documents
        max_features=50_000,
        sublinear_tf=True,
        stop_words=None,         # IDF already suppresses common words
    ),
    LogisticRegression(max_iter=1000),
)
model.fit(train_texts, train_labels)

That is a strong text classifier in eight lines. It trains in seconds on tens of thousands of documents, runs on a CPU, and is fully interpretable — the model's coefficients map directly to words, so you can read what it learned.

min_df is doing the heavy lifting. Dropping terms that appear in fewer than three documents typically removes most of the vocabulary and almost none of the signal, because the long tail is mostly typos and one-off tokens.

ngram_range=(1, 2) adds bigrams, which is what lets the model represent "not good" — the single most valuable addition for sentiment tasks.

Counting words, then discounting the common ones

TF-IDF is two ideas multiplied together, and both of them are visible in a handful of documents. This builds the matrix by hand, shows what each half contributes, and finds the case where the whole thing collapses.

example_01.pyNumPy
Output

Things to try

  1. Turn idf off and query the. Every document scores, and the ranking is driven by a word that distinguishes nothing.
  2. Query the cat with idf on. Only cat contributes — the stopword cancels itself without a stopword list.
  3. Turn tf damping off and watch the document that repeats a term run away with the score. That runaway is what BM25's saturation fixes.

What to remember

TF-IDF weights a term by how often it appears here times how rare it is everywhere. A term in every document has an idf of zero and drops out by arithmetic rather than by a list. It knows nothing about meaning, which is why a synonym scores zero and why dense retrieval runs alongside it.

What it cannot do

No sense of similar words. "Excellent" and "superb" are unrelated columns. A model trained on documents containing one learns nothing about documents containing the other. This is the fundamental limitation, and it is what embeddings fix.

No word order beyond n-grams. Bigrams capture adjacent pairs; anything longer-range is invisible.

No handling of unseen vocabulary. A term absent from the training corpus has no column, so it contributes nothing at prediction time.

Very high dimensionality. Fifty thousand features, almost all zero per document. Manageable with sparse matrices and sparse-aware models, and it rules out anything that requires dense input.

 TF-IDFEmbeddings
Similar words relatedNoYes
DimensionalityTens of thousands, sparseHundreds, dense
InterpretableYes — coefficients are wordsNo
Needs a model to computeNoYes
Exact term matchingExcellentWeak
Handles paraphraseNoYes

That table is the argument for using both, which is exactly what hybrid retrieval does.

Where it is still the right tool

As the first baseline for any text classification task. It is fast enough to try in five minutes, and it frequently comes within a few points of a fine-tuned transformer on topic classification with adequate labels. Not measuring it means not knowing whether the transformer was worth it.

Keyword retrieval, usually in its refined form BM25, which adds term-frequency saturation and document-length normalisation. Every search engine is built on this.

Keyness analysis — comparing two corpora to find the terms that distinguish them. Far more informative than a word cloud.

Interpretable models in regulated settings, where "the model weighted these words" is a required explanation.

Low-resource environments — no GPU, no model download, no inference latency.

BM25, briefly

BM25 is TF-IDF with two corrections, and it is what you should use for retrieval rather than plain TF-IDF:

Term-frequency saturation. The contribution of repeated occurrences flattens, controlled by k₁ (around 1.2). A document mentioning a term twenty times is not twenty times more relevant.

Document-length normalisation relative to the corpus average, controlled by b (around 0.75). A long document will contain most query terms by chance.

Both corrections address ways raw TF-IDF over-rewards documents. For classification features TF-IDF is fine; for ranking search results, BM25 is meaningfully better.

Questions people ask

Do I need stop-word removal? Not really — IDF drives ubiquitous terms to zero. Removing them shrinks the vocabulary, which is the only benefit.

Should I stem or lemmatise? For TF-IDF, yes — it pools evidence across word forms and shrinks the vocabulary. Never for transformers.

What min_df should I use? 2–5 for most corpora. It removes the long tail of typos and one-offs.

Is TF-IDF outdated? No. It is a strong baseline, the basis of keyword retrieval, and interpretable in ways embeddings are not.

Should I use it with embeddings? Yes — hybrid retrieval combining BM25 with dense search beats either alone on most real corpora.

Why is my matrix enormous? Bigrams without min_df or max_features. Cap both.

Recap in one screen

  • TF-IDF multiplies how often a term appears here by how rare it is across the corpus.
  • A term in every document gets an IDF of zero — automatic, corpus-derived stop-word removal.
  • The logarithm stops rare terms getting absurd weights; sublinear TF does the same for repetition.
  • TF-IDF plus logistic regression is a fast, interpretable baseline worth measuring before anything heavier.
  • It cannot relate similar words, which is what embeddings add — and BM25 is the version to use for retrieval.

Recall check

0 of 3

Say the answer out loud before you reveal it — recalling it is what makes it stick, and rereading it is not.

  1. Without scrolling back — what is the one-line takeaway from this module?

  2. What does this module say about “The two halves”?

  3. What does this module say about “Where it falls short”?

Cheat sheet

TF-IDF

Two numbers multiplied. Term frequency rewards a word for appearing often in this document; inverse document frequency punishes it for appearing in every document. The product is high only for words that are frequent here and rare elsewhere — which is a workable definition of "what this document is about".

GEN AI · vizlearn.in/gen_ai/tf_idf.html

About the author

Ashish Jangra builds and maintains VizLearn. Every module here is written and the visualisation behind it hand-built, so the numbers in a readout come from the same code that draws the picture. Corrections are genuinely welcome and get priority over everything else — if a page states something wrong, or an animation misrepresents what the algorithm does, get in touch.