Skip to content
The term-document matrix

Three short reviews go in and a table of numbers comes out. Each cleaning switch on the right rebuilds that table, and some of them change which two reviews the arithmetic calls closest. The reviews, and the figures at the opening settings, are the ones worked through on this page.

Closest pair on raw counts
R1 & R2
cosine 0.40
Matrix and sparsity
7 × 3
11 of 21 cells filled, 10 empty (48%)
Rows are terms, columns are reviews, cells hold counts. Rows are ordered by how many reviews hold the term, then by total count.
TermR1R2R3df
battery2102
life1102
poor1012
screen0112
great0201
died1001
fast1001
Cosine similarity on raw counts, track fixed from 0 to 1R1 & R20.401stclosestR1 & R30.253rdR2 & R30.272nd

R1 & R2: 3 ÷ (2.828 × 2.646) = 0.401

Terms (rows) 7Filled cells 11 of 21Dot product, R1 and R2 3
Cleaning stages
Cell weighting
Try this
The three reviews, and what a reader says
R1. The battery died fast. Battery life is poor.
reader calls it: complaint · 6 tokens kept
R2. Great battery life, great screen.
reader calls it: praise · 5 tokens kept
R3. The screen is poor.
reader calls it: complaint · 2 tokens kept

Those stance labels are a reader's judgement. Nothing in the matrix produced them, and no setting on this widget can produce them.

On raw counts the matrix puts reviews 1 and 2 closest at 0.40, from a dot product of 3 over 2 terms that contribute to it. A reader groups reviews 1 and 3 together because both are complaints while review 2 is praise. The arithmetic disagrees here. Bag of words measures shared vocabulary, and shared vocabulary is not shared stance. These are the opening settings, so the 7 rows, the 11 filled cells of 21 and both headline cosines are the figures printed on this page. Switching the weighting moves the closest pair: raw counts pick reviews 1 and 2 and tf-idf picks reviews 1 and 3. The stop list removed 4 of 17 tokens.
Everything above is recomputed from the three reviews on every change. Nothing is stored. The stop list here is the two words the worked example uses, not a full standard list.
The text mining pipelineOpen in Dr Phil's Quant Lab ↗