Cosine Similarity

What it means

  • Each item (movie, product, song, paper, etc.) is represented by a feature vector: genres/tags, text keywords, image or audio embeddings, price/brand indicators, etc.
  • Cosine similarity measures how aligned two item vectors are—i.e., how similar their feature profiles are—independent of scale (length).

The math

For item vectors $x, y \in \mathbb{R}^d$:

$\text{cosine\_sim}(x, y)=\frac{x\cdot y}{\lVert x\rVert_2\,\lVert y\rVert_2} =\frac{\sum_{k=1}^d x_k y_k}{\sqrt{\sum_{k=1}^d x_k^2}\,\sqrt{\sum_{k=1}^d y_k^2}}$

  • Range is $[-1, 1]$ in general; with non‑negative features (e.g., one‑hot, TF‑IDF), it’s typically $[0,1]$.
  • Cosine distance (often used in diversity): $1 – \text{cosine\_sim}(x,y)$.

Why cosine?

  • Scale invariance: If you double all features for an item, direction stays the same, so similarity doesn’t change. That’s ideal when only relative composition matters (e.g., genre mix, term proportions).
  • Works well with sparse, high‑dimensional vectors (bags of words, tags).
  • Fast to compute and index; compatible with approximate nearest‑neighbor search (ANN).

Building item feature vectors

Common ways to encode items:

  1. One‑hot / multi‑hot categories (genres, brands, tags).
  2. TF‑IDF from text fields (title/description → token weights).
  3. Learned embeddings (e.g., sentence transformers for text, CNN embeddings for images, audio embeddings).
  4. Hybrid vectors (concatenate normalized content features + numeric attributes; consider scaling/standardization first).

Worked mini‑example

Suppose a simple genre vector: $[{\small \text{Action},\ \text{Romance},\ \text{Comedy}}]$

  • Movie A: $x=[1,0,1]$ (Action & Comedy)
  • Movie B: $y=[1,0,0]$ (Action only)

$x\cdot y = 1,\quad \lVert x\rVert=\sqrt{2},\ \lVert y\rVert=1 \Rightarrow \text{cosine\_sim}(x,y)=\frac{1}{\sqrt{2}}\approx 0.707$

They’re fairly similar because both are Action; A has extra Comedy signal.

How it’s used in recommender systems

  • Item‑to‑item similarity: “Because you liked X, here are items with high cosine similarity to X’s features.”
  • Cold‑start content‑based recs: When a new item has no interactions yet, use its content features to find neighbors via cosine.
  • Diversity (ILD): Define $d(i,j)=1-\text{cosine\_sim}(x_i,x_j)$; then Intra‑List Diversity is the average pairwise $d(i,j)$ across a user’s list.
  • Re‑ranking: Start from a relevance shortlist, then re‑rank to maximize relevance + (1 − cosine) to avoid near‑duplicates.

Practical tips

  • L2‑normalize feature vectors before computing dot products (or use a library’s cosine function, which effectively normalizes).
  • Feature scaling matters: When mixing modalities (text + categories + numbers), standardize/weight sub-blocks so one source doesn’t dominate.
  • Zero vectors: Handle items with no features (avoid divide‑by‑zero; drop or impute).
  • Dimensionality & sparsity: For huge catalogs, precompute top‑K neighbors or use ANN (e.g., FAISS, ScaNN, HNSW).
  • Domain choice of features: Cosine is only as good as the representation; prefer semantically meaningful embeddings/encodings.

Common pitfalls & variants

  • Popularity leakage: If features include popularity counts, they can overshadow content; consider log‑scaling or excluding them from cosine.
  • Negative entries: Some embeddings are mean‑centered → cosine can be negative; that’s fine, just interpret negative as “anti‑similar.”
  • Adjusted cosine (ratings CF): Different concept: compute cosine between item rating vectors after subtracting user means; that’s collaborative, not content‑feature cosine.
  • Angular distance: $\theta=\arccos(\text{cosine\_sim})$ is a true metric on the unit sphere; sometimes used in ANN.

Discover more from Insightful Data Lab

Subscribe to get the latest posts sent to your email.

Similar Posts

Questions, corrections, or additional insights?

This site uses Akismet to reduce spam. Learn how your comment data is processed.