Embedding 4 billion tokens through OpenAI's text-embedding-3-small costs $80*. That's a corpus of 2 million documents at 2,000 tokens each, and the invoice is smaller than a team lunch.
So you pick an API, embed once, and stop thinking about it. The trouble starts on the second pass, and on the vector database bill that arrives every month after the first.
What the first pass costs
List prices per million tokens, from TokenCost as of July 2026:

Run the same 4-billion-token corpus through the pricier models and you get $480 on Voyage 4 large, $520 on text-embedding-3-large, and $800 on Gemini Embedding 2. If someone tells you to self-host embeddings to save $80, check what they sell. (We sell something. Keep reading anyway.)
The pass you'll run ten times
You can't mix vectors from two models in one index. Change the model, the chunk size, the overlap, or the header text you prepend to each chunk, and you re-embed everything.
If you care about retrieval quality, you'll change those settings a lot. Ten chunking experiments on that corpus means ten full passes, and the per-token price gets multiplied by every one of them.
Vendors also retire models. OpenAI shut down its first-generation embedding models in January 2024, and every team that had built on them re-embedded on OpenAI's schedule rather than their own.
Open weights don't come with a deprecation email. Once the files are on your own disk, the model stays until you decide to replace it.
The bill that comes every month
Split 4 billion tokens into 500-token chunks and you have about 8 million vectors. Stored as float32, that's 131 GB at 4,096 dimensions, 49 GB at 1,536, and 33 GB at 1,024.
You pay for that memory in your vector database every month, once per replica. The dimension count usually shows up as a quality setting in a model card, and it ends up as a line on your infrastructure invoice.
Then there's where the text goes. Embedding a contract means sending the whole contract to whoever runs the model, and a RAG index over tickets, medical notes, or source code sends all of it in bulk. For a regulated team, that settles the question before anyone looks at the price table.
Which open models you can actually ship
"Open" covers a lot of licenses. For a commercial product, you want Apache 2.0 or MIT, and you want to read that on the model card itself.
Here's a starting shortlist. Every model below is Apache 2.0 or MIT, and all of them run on kibbu as well as on common open runtimes like vLLM, Text Embeddings Inference, and sentence-transformers:

Cheap, fast indexing of short chunks
Qwen3 Embedding is the family most teams test first, and for good reason: the 8B model ranked first on the MTEB multilingual leaderboard at release in June 2025, with a score of 70.58. The smaller models on the list often win anyway, because a 568M model needs about 14 times less compute per re-index than an 8B one.
Two popular models to skip for commercial use: Jina Embeddings v4 is CC-BY-NC, and EmbeddingGemma ships under Google's own Gemma terms.
Leaderboards get you a shortlist. Fifty of your own users' queries, each with the documents that should come back, will pick the winner.
Four ways teams break their own index
Skipping the query prefix. Most of these models expect queries and documents to be marked differently, and each one has its own convention. Qwen3 Embedding and multilingual-e5-instruct want Instruct: {task}\nQuery: {query} on queries only, Nomic wants search_query: and search_document:, and Arctic Embed wants query: on queries. The Qwen team measured a 1% to 5% retrieval gain from the instruction, and it's the step people skip most.
Truncating dimensions and stopping there. Qwen3 Embedding, Arctic Embed, and Nomic all support Matryoshka truncation, so you can keep, for example, the first 1,024 of Qwen3 8B's 4,096 dimensions. That's the 131 GB to 33 GB cut from earlier, but you have to L2-normalize again after truncating or your cosine scores drift. Pick the dimension before you build the index, because changing it later is another full pass.
Mixing builds of the same model. A quantized GGUF build and a full-precision build produce slightly different vectors. Embed documents with one and queries with the other, and recall slips without a single error in your logs. Pin the weights, the quantization, and the runtime version, and write them into the index metadata.
Running queries and indexing on one path. A query is 20 tokens that need to come back in milliseconds. A re-index is billions of tokens that need to be done by morning. Serve queries from a small, always-on deployment and send indexing wherever compute is cheapest.
Whichever path you pick, put an OpenAI-compatible client in front of it. Kibbu, vLLM, TEI, and Ollama all serve /v1/embeddings, so a model swap is a base URL and a model name.
Where the re-index should run
Indexing is the easiest job in inference. Each chunk is a single forward pass with no token-by-token decoding, chunks don't depend on each other, and a failed one just gets retried. Every sub-1B model on the list fits in the memory of any recent laptop.
Here's the rough arithmetic, not a benchmark. A forward pass costs about 2 FLOPs per parameter per token, so a 0.6B model (Qwen3 0.6B, or roughly BGE-M3 and Arctic Embed) needs 1.2 GFLOPs per token and the full corpus needs about 4.8 × 10^18 FLOPs. At an assumed [5 TFLOPS effective] per laptop, that's [about 11 days] on one machine, or [5 to 6 hours] across 50 machines sitting idle overnight.
Measure your own hardware before you trust my 5 TFLOPS. Even if the real number is half that, a full re-index fits inside one night on laptops your company already paid for, and the tenth chunking experiment costs electricity.
None of this pays off if you embed once and never touch the index again. For that team, $80 to an API is the right call.
* 4 billion tokens × $0.02 per million tokens, OpenAI's list price for text-embedding-3-small (OpenAI). The 2-million-document corpus is an illustrative example; price your own by multiplying your token count by the rate.
At Kibbu, we run any Apache 2.0 or MIT embedding models, from Qwen3 Embedding to BGE-M3, behind an OpenAI-compatible API on compute companies already own. It's built for jobs like the re-index above that can wait a few hours. If you're running this math on your own corpus, try it with us.



