Skip to content
Navigation
Dashboard
📖Reading•20 min

Choosing the Right Model

Select the best embedding model for your specific use case.

Choosing Embedding Models

Key Factors

1. Task Type

  • Retrieval: BGE, E5, Instructor
  • Classification: Sentence-BERT
  • Clustering: Most work well
  • 2. Domain

  • General: OpenAI, BGE
  • Code: CodeBERT, StarCoder embeddings
  • Legal/Medical: Domain-specific models
  • 3. Languages

  • English only: More options
  • Multilingual: mE5, multilingual-e5-large
  • 4. Performance vs Cost

  • API: Easy but ongoing cost
  • Local: Free but requires setup
  • Evaluation Approach

    1. Select candidate models 2. Create test dataset 3. Measure retrieval accuracy 4. Compare latency and cost 5. Test edge cases

    Recommended Defaults

    Use Case | Model |
    |----------|-------|
    General | OpenAI text-embedding-3-small |
    Open Source | BGE-large-en-v1.5 |
    Multilingual | multilingual-e5-large |
    Instruction-following | Instructor-large |

    🎯 Key Takeaways

    • ✓Match model to your task type
    • ✓Test on your actual data
    • ✓Consider cost and latency
    • ✓Domain-specific models often perform better

    📚 Additional Resources