Embeddings: What Matters in Practice
It's not about the model. It's about data density.
1. Embeddings Are Geometric Meaning
An embedding is a list of numbers (e.g., [0.1, -0.5, 0.8...]).
These numbers represent the "coordinates" of a piece of text in a multidimensional meaning-space.
Text with similar meaning ends up close together (Cosine Similarity).
But here is the catch: Embeddings compress meaning. They squeeze 1000 words into 1536 numbers. Loss is inevitable.
2. Representation > Model
Developers obsess over "OpenAI vs Cohere vs Llama." But in production, Chunking dominates quality.
If you embed a chunk that contains 3 distinct topics, the vector will be the average of those 3 topics. Result: The vector matches none of them strongly. It lands in "no man's land" in vector space.
Rule: A chunk should express one clear semantic idea.
3. Dimensionality & Curse of Dimensions
- Small Dimensions (384): Fast, low storage. Good for simple sentences.
- Large Dimensions (3072): Slow, expensive. Captures nuance.
But "bigger is better" is false. High-dim vectors need more data to fill the space.
4. Summary
Don't blame the model if your retrieval fails. Blame the data density. If your input is vague, your vector is vague.
Key Intuition: "Garbage in, Average out."