Most recommendation engines running in production today were built before 2018. They work — until they don’t. Cold start kills conversions, the black box frustrates product teams, and behavior-based signals miss what content actually means.
This post walks through migrating from a traditional collaborative filtering setup to a RAG-powered semantic recommendation system. Same goal, completely different engine under the hood.
Part 1 — The Old Way (Collaborative Filtering)#
The classic approach: users who liked X also liked Y. You build a user-item matrix and find similarity by behavior patterns.
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity
# User-item matrix: rows = users, cols = items, values = ratings
user_item_matrix = np.array([
[5, 3, 0, 1],
[4, 0, 4, 1],
[1, 1, 0, 5],
[1, 0, 4, 4],
])
# Find items similar to item 0 (column-wise similarity)
item_similarity = cosine_similarity(user_item_matrix.T)
def get_similar_items(item_id, top_n=3):
scores = list(enumerate(item_similarity[item_id]))
scores = sorted(scores, key=lambda x: x[1], reverse=True)
return [i for i, _ in scores[1:top_n+1]]
print(get_similar_items(0)) # Items most similar to item 0What breaks:
- Cold start — new item with no behavior data gets zero recommendations
- No content understanding — a cat video and a dog video look identical if they have the same watch patterns
- Black box — you cannot explain why an item was recommended
- Data hungry — needs massive user interaction history to work well
Part 2 — The New Way (RAG-Powered Semantic Search)#
Instead of learning from behavior, we embed the content itself into a vector space. Similarity becomes meaning-based, not pattern-based.
from sentence_transformers import SentenceTransformer
import chromadb
# Sample content catalog
items = [
{"id": "1", "title": "Funny cat playing piano", "description": "A cat pressing piano keys in a surprisingly musical way"},
{"id": "2", "title": "Dog learns to skateboard", "description": "Golden retriever successfully rides a skateboard at the park"},
{"id": "3", "title": "Piano masterclass Chopin", "description": "Professional pianist performs Chopin nocturne with commentary"},
{"id": "4", "title": "Cat knocks things off table", "description": "Classic cat behavior compilation, knocking objects off surfaces"},
{"id": "5", "title": "Skateboarding tricks tutorial", "description": "Step by step guide to learning kickflips and ollies"},
]
# Embed content
model = SentenceTransformer("all-MiniLM-L6-v2")
# Store in ChromaDB vector store
client = chromadb.Client()
collection = client.create_collection("content")
for item in items:
embedding = model.encode(item["description"]).tolist()
collection.add(
ids=[item["id"]],
embeddings=[embedding],
documents=[item["description"]],
metadatas=[{"title": item["title"]}]
)
def get_recommendations(query_text, top_n=3):
query_embedding = model.encode(query_text).tolist()
results = collection.query(
query_embeddings=[query_embedding],
n_results=top_n
)
return results["metadatas"][0]
# Find content similar to a cat video
recs = get_recommendations("cat playing musical instrument")
for r in recs:
print(r["title"])
# Output:
# Funny cat playing piano ← semantic match: cat + music
# Piano masterclass Chopin ← semantic match: piano/music
# Cat knocks things off table ← semantic match: cat behaviorWhat improves:
- Cold start solved — new item gets embedded immediately, no behavior needed
- Content understanding — the vector captures meaning, not just patterns
- Works on small datasets — 10 items or 10 million, same system
- Multimodal ready — swap the encoder for images, audio, or video features
Part 3 — The Migration#
Switching from collaborative filtering to semantic search is a 3-step process.
Step 1: Embed your existing content catalog#
def embed_catalog(items: list[dict]) -> None:
model = SentenceTransformer("all-MiniLM-L6-v2")
client = chromadb.PersistentClient(path="./rec_db")
collection = client.get_or_create_collection("catalog")
embeddings = model.encode([item["description"] for item in items]).tolist()
collection.add(
ids=[item["id"] for item in items],
embeddings=embeddings,
documents=[item["description"] for item in items],
metadatas=[{"title": item["title"], "category": item.get("category", "")} for item in items]
)Step 2: Replace item-item similarity with vector lookup#
# BEFORE (collaborative filtering)
def old_recommend(item_id):
return get_similar_items(item_id) # behavior matrix lookup
# AFTER (semantic search)
def new_recommend(item_description, top_n=5):
model = SentenceTransformer("all-MiniLM-L6-v2")
client = chromadb.PersistentClient(path="./rec_db")
collection = client.get_collection("catalog")
query_embedding = model.encode(item_description).tolist()
results = collection.query(query_embeddings=[query_embedding], n_results=top_n)
return results["metadatas"][0]Step 3: Add LLM explainability layer#
This is where Rec Engine 2.0 goes beyond what the old system could ever do — explaining why something was recommended.
import anthropic
def recommend_with_explanation(query: str, top_n: int = 3) -> dict:
recs = new_recommend(query, top_n)
client = anthropic.Anthropic()
rec_titles = [r["title"] for r in recs]
message = client.messages.create(
model="claude-opus-4-5",
max_tokens=256,
messages=[{
"role": "user",
"content": f"A user is watching: '{query}'\n\nWe are recommending: {rec_titles}\n\nIn one sentence each, explain why each recommendation is relevant."
}]
)
return {
"recommendations": recs,
"explanation": message.content[0].text
}
result = recommend_with_explanation("cat playing piano")
print(result["explanation"])Part 4 — Real Results#
| Metric | Collaborative Filtering | RAG 2.0 |
|---|---|---|
| Cold start | Fails | Works immediately |
| Min data required | Thousands of interactions | Zero interactions |
| Explainable | No | Yes |
| Content understanding | No | Yes |
| Multimodal | Hard | Native |
| Latency | Fast (matrix lookup) | Fast (vector search) |
Closing#
If your recommendation engine was built before LLMs were mainstream, it’s leaving relevance on the table. The migration is lighter than you think — embed your catalog, swap the lookup, add an explanation layer.
The cat video finds the piano video not because other users watched both, but because they share meaning. That’s Rec Engine 2.0.
Building something with this or need help migrating your stack? Reach out. +++