Skip to content
Navigation
Dashboard
🎬Video•25 min

Semantic Similarity & Distance

Measure and interpret semantic similarity between texts.

Semantic Similarity

Measuring Similarity

Once you have embeddings, compare them:

Cosine Similarity: Measures angle between vectors. Range: [-1, 1] 1 = identical meaning 0 = unrelated -1 = opposite meaning

Implementation: cos_sim = np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

Interpretation

What scores mean:

  • 0.9+: Very similar, likely same topic
  • 0.7-0.9: Related content
  • 0.5-0.7: Some connection
  • <0.5: Different topics
  • Applications

    Duplicate Detection: Flag documents with similarity > 0.95

    Clustering: Group documents by similarity

    Search Ranking: Rank by relevance score

    Recommendation: "Similar to what you read"

    Challenges

  • Domain-specific vocabulary
  • Negation handling
  • Context length limits
  • Cross-lingual similarity
  • 🎯 Key Takeaways

    • ✓Cosine similarity measures semantic closeness
    • ✓Threshold tuning is task-dependent
    • ✓High similarity indicates related content
    • ✓Consider domain and context

    📚 Additional Resources