ct smith

docs goblin

I tried Cohere's embedding models

October 2, 2026

I did some light embedding model benchmarking on my new project's docs. I wanted to see how Cohere's new embedding models fare against my faves: the Voyage 4 models. I also wanted to figure out how to tune my hybrid retrieval technique so I figured trying a few different models and hammering my data with questions would help flush out issues with my technique (it did).

Anyway -- I'm not an AI scientist or a researcher, this is straight up folk science. I compared five Voyage and Cohere embedding models on the same 128-page documentation set using 200 search questions, including 179 with documented answers. With vector-only search, Voyage 4 Lite found an expected answer within the first three pages for 95.0% of answerable questions, followed by Voyage 4 Large at 94.4%, Voyage 4 and Cohere Pro at 93.3%, and Cohere Fast at 92.7%. Cohere returned embeddings faster in our timing samples, while adding word matching through our current hybrid ranking lowered every model’s top-three success rate. The accuracy differences were small, so the strongest takeaway is that I need to keep tweaking hybrid search (and that it might not actually even matter for plain search, aka not-the-chatbot).

As far as the actual mechanics of the experiment: I used a script to ask a bunch of questions using Shoalrun's existing search tools. Here's what it did:

  1. Loaded a fixed copy of my docs: 128 pages split into 1,097 passages.
  2. Used Voyage and Cohere’s APIs to turn the passages and questions into embeddings. It reused previously saved embeddings when the model, settings, and input text matched.
  3. Ran each question through Shoalrun’s actual search code using word matching, vector search, and hybrid search.
  4. Compared the returned pages against the expected pre-labeled pages.
  5. Calculated the scores and saved each question’s results in the JSON.

No LLM/chatbot judged the results. The script checked whether the expected pages appeared near the top. Of the 200 questions, 179 had documented answers and contributed to the success scores; the other 21 were retained as no-answer cases.

Also, this only tells me how good the models performed with my specific docs, I'm not claiming benchmark authority or anything here. Again, this is CT's Folk Science Power Hour, not an AI research lab.


Run: 2026-10-02 Dimensions: 1024 float32. Corpus: 128 pages, 1097 passages. Queries: 200 (179 answerable).

All models use identical passage text, chunking, queries, filters, and retrieval code. Document and query input types are set explicitly; truncation is disabled. No reranker is used. Hit rate requires a hit from every expected group of alternative pages; MRR uses the rank where all groups have been found. Metrics use the top 60 retrieved passages, collapsed to distinct pages. No-answer queries are reported separately and excluded from success metrics.

Lexical search matches words using BM25, which rewards distinctive query words and accounts for passage length, plus title, heading, phrase, and synonym bonuses. Vector search compares numerical embeddings to find similar meaning even when wording differs. Hybrid search combines the lexical and vector rankings with equal-weight reciprocal rank fusion. Hybrid can help exact identifiers, but can also move a good semantic result down.

How to read this benchmark

This benchmark asks: when someone searches our documentation, does the right page appear near the top? It compares search results, not the quality of a chatbot's written answers. Higher Hit rates and MRR are better; lower request times are faster.

Read the vector rows to compare embedding models by themselves. Read the hybrid rows to compare those models with Shoalrun's current word-matching and ranking combination. A lower hybrid score means that combination moved useful results down; it does not mean the embedding model changed.

The Documents column names the model that converted the documentation into embeddings. The Queries column names the model that converted the search questions. These can differ when the provider explicitly supports compatible models. For example, voyage-4-large / voyage-4-lite means Large embeds the docs and Lite embeds the questions. none / none means word matching without an embedding model. Voyage and Cohere are the providers; Large, Lite, Pro, and Fast are model names, not benchmark scores.

Text and search terms
Term Plain-language meaning
Benchmark A repeatable test that compares systems on the same inputs and scoring rules.
Corpus; frozen corpus The collection of documentation being searched. Frozen means its contents stay fixed between comparisons.
Document; page One published documentation page. Stable document IDs identify pages even if their URLs change.
Passage; chunk A smaller section of a page. Search first ranks passages, then this benchmark scores distinct pages.
Chunking Splitting pages into passages. Keeping this the same helps make model comparisons fair.
Query The words or question someone enters into search. A query here is a labeled test case, not a chatbot response.
Retrieval Finding and ordering existing documentation that may answer a query.
Index Prepared data that makes the documentation searchable. This test uses a passage index and model-specific embeddings.
Embedding; vector A list of numbers that an embedding model creates from text. Similar text tends to produce vectors that are close together. The numbers are not a written answer or an explicit list of topics.
Embedding model A model that converts text into vectors so search can compare its meaning.
Dimensions The number of values in each vector: 1,024 in this test. More dimensions do not automatically mean better results.
float32 A number format using 32 bits for each value. All models use the same format here.
Normalized vector A vector scaled to have length one. This makes the similarity calculation consistent.
Cosine similarity A comparison of the directions of two vectors. More similar directions receive higher scores. It is not a probability that a page is correct.
Shared embedding space A provider-supported arrangement in which different models produce compatible vectors. Matching dimensions alone does not make models compatible; we do not mix Voyage and Cohere vectors.
Document/query input types API settings telling the provider whether text is documentation to store or a question to search with. These can affect the resulting vectors.
Lexical search Word matching, useful for exact names, commands, configuration keys, and error messages.
BM25 “Best Matching 25,” a standard word-matching scoring formula. It rewards query words that are uncommon across the corpus, considers their frequency in a passage with diminishing returns, and adjusts for passage length. Shoalrun adds title, heading, phrase, and synonym bonuses.
Semantic; vector search Searching by similarity of meaning using embeddings. It can match “publish my site” with documentation that says “deploy your documentation.”
Hybrid search Combining word-matching and vector results. It can help exact terms, but a poorly balanced combination can lower useful results.
Reciprocal rank fusion (RRF) Combining rankings using positions rather than their raw scores. Here each passage receives 1 / (60 + position) from each list in which it appears, with positions starting at one. The contributions are added with equal weight. The number 60 softens the advantage of the first few positions.
Retrieval depth; page collapse How many passages search considers, and how multiple passages from the same page become one page result. Hybrid considers up to 240 passages from each search and keeps 60 after fusion; vector keeps 60 directly. Pages keep the order of their first appearance.
Reranker An extra model or scoring step that reorders retrieved results after the initial search. This benchmark uses no reranker.
Publication scope; reader filters Settings that restrict eligible results to an edition, language, or selected reader view. The same settings apply to all models. Reader views control visibility, not authorization.
Truncation Cutting off text to fit a limit. Provider input truncation is disabled here so models receive the complete submitted text.
Labels and success scores
Term Plain-language meaning
Label; expected page The page or pages chosen in advance as acceptable answers to a test query. Labels can miss other relevant pages.
Answerable; no-answer (none) Answerable means the frozen docs contain an expected answer. No-answer cases ask for information those docs do not contain. They remain in the JSON but are excluded from Hit rates and MRR; this report does not score whether search correctly refuses an unanswerable question.
Alternative pages Multiple acceptable pages for one question. Finding any one satisfies that group.
Expected groups; synthesis A question needing several pieces of information can have multiple groups of acceptable pages. A hit requires a page from every group.
Rank A page's position in the search results, starting at one. For multiple expected groups, we use the position where all groups have been found. JSON rank 0 means at least one group was not found within the retrieved results.
Hit@1 / Hit@3 / Hit@5 / Hit@10 The percentage of answerable queries whose expected pages were found within the first 1, 3, 5, or 10 distinct pages. For example, Hit@3 of 95% means success within the first three pages for about 95 out of 100 answerable questions.
MRR Mean reciprocal rank: average 1 / rank across answerable questions, with a miss contributing zero. First place contributes 1, second 0.5, third about 0.333. The score ranges from 0 to 1 and rewards finding answers early. For synthesis questions this test uses the rank where all required groups are satisfied.
Why Hit@3 and MRR can disagree A model can find more answers within three pages yet put fewer of them first. Hit@3 rewards top-three success equally; MRR rewards first place more than second or third.
Split; original; additional Subsets of the query set. The original fixture contains 94 cases, including 83 answerable ones. The expanded benchmark adds 106 cases, including 96 answerable ones, for 200 total and 179 answerable.
Fixture; CI A fixture is saved test data. CI means continuous integration: automated checks run as code changes. The original fixture is also used by those checks.
Query kinds exact: names or identifiers; task: how to accomplish something; paraphrase: different wording for a topic; concept: explanations; mixed: combinations of wording or intent; typo: misspelled queries; synthesis: multiple required pieces; none: no documented answer. These are author-assigned categories, not model judgments.
Held-out evaluation A test set kept separate from examples used to develop or tune the system. These author-labeled queries are not an independently judged held-out evaluation, so the results do not establish which model is best for all search tasks.
Comparisons and uncertainty
Term Plain-language meaning
Baseline; candidate The reference setup and the setup compared against it. Here the baseline uses Voyage Large for documents and Voyage Lite for queries. Comparisons use the same search mode on both sides.
Paired comparison Comparing both setups on the same questions, rather than comparing unrelated sets of questions.
Hit@3 difference; delta Candidate Hit@3 minus baseline Hit@3, expressed in percentage points. Going from 90% to 92% is +2 percentage points. A positive difference favors the candidate; a negative one favors the baseline.
Wins; losses; ties A win means the candidate succeeds within three pages and the baseline misses. A loss is the reverse. A tie means both succeed or both miss; it does not require identical rankings.
Bootstrap; seeded sampling Repeatedly drawing questions from this set, allowing repeats, to estimate how much the paired difference varies. The test uses 2,000 samples. A fixed random seed makes the calculation repeatable.
95% interval The middle 95% of those bootstrap differences. An interval crossing zero includes both an advantage and a disadvantage. It is not a 95% probability that a model is universally better. Related questions can make these intervals look more certain than an independent test would.
Independent queries Questions whose outcomes do not strongly depend on the same topic or wording. Many questions here cover related pages, so their outcomes can be linked.
Timing and reproducibility
Term Plain-language meaning
API; request The provider's interface, and one call to it asking for embeddings. API stands for application programming interface.
Batch Several texts embedded in one request, up to 32 here. One request does not necessarily mean one passage or query.
Live / cached Live means an API call was made; cached means a previously saved result was reused. For example, 0/35 means zero live requests and 35 reused batches.
Live document ms Total measured time for live document-embedding requests in this run, in milliseconds. Zero means every document batch was cached, not that embedding the corpus originally took no time.
Single-query samples Fresh API calls that each embed one query, bypassing the cache. Their times estimate embedding request latency. Samples include repeated query texts and are separate from the full set of quality-test questions.
Latency; p50; p95 Latency is elapsed request time. p50 is the median: about half the measured calls were at or below it. p95 is the value at or below which about 95% fell. Lower is faster; 100 ms is one tenth of a second.
Network time; retries Travel time to the provider and extra attempts after a failed request. Both are included in measured live request time.
Rate-limit pacing; wall time Pacing deliberately waits to stay under the provider's request/token limits. These waits are excluded from the timing samples. Wall time is the elapsed time for the whole run, including waits and other work. The timing table is not total wall time or end-to-end search latency.
Token; billed tokens A token is a piece of text counted by a model's tokenizer. Providers may report billed input-token counts. Recorded usage covers the live quality-embedding batches, not cached requests or the separate latency samples, so it is not a complete billing statement.
JSON; schema version JSON is the structured data file behind the readable report. A schema version identifies the report's field layout.
Digest; fingerprint; SHA-256 A reproducible hash of input data used to check that the corpus or questions have not changed. SHA-256 is the hashing algorithm. These hashes are not API keys or encryption passwords.
Reproducibility Keeping the inputs and settings available so another run can be compared fairly. Provider updates and changing network conditions can still affect results.
Guide to the JSON fields
Fields What they contain
schemaVersion, createdAt, dimension Report layout version, UTC run timestamp, and vector length.
corpusDigest, queryDigest Fingerprints of the passage index and labeled query file.
pages, passages, queries, answerable Counts of distinct docs, text chunks, test cases, and cases included in success scores. queries is also used inside a model record for query-request statistics.
models, name, provider, status, error Embedding model records, provider names, completion status, and any sanitized error message.
documents, queries, liveRequests, cachedRequests, liveMs, billedTokens, tokensReported Per-model document/query batch statistics. tokensReported says whether all live batches in that category supplied usage counts. These records exclude preflight access checks and fresh latency samples.
latency, samples, query, repeat, ms, p50, p95 Fresh single-query timing observations, their query text, repetition number starting at zero, milliseconds, and summary percentiles.
runs, documentModel, queryModel, scores Each tested document/query model pairing and its lexical, vector, or hybrid scores.
overall, byKind, bySplit, count, hitRate, mrr Scores for all answerable questions, by category, or by original/additional subset. count is the answerable denominator; hitRate keys 1, 3, 5, 10 hold fractions from 0 to 1.
rows, kind, query, split Individual labeled test cases, their category, text, and subset.
expect, expectGroups Acceptable document IDs, or groups of alternatives that must each be satisfied. expectGroups takes precedence when present.
variant, readerSelection, focus Optional publication scope and reader-view selections. A variant identifies channel, version, and language; focus is one reader-filter dimension.
rank, top, bestScore Expected-answer rank, first ten distinct result IDs, and the first passage's raw retrieval score. Raw scores use different scales across search modes and should not be compared as probabilities or accuracy percentages.
comparisons, candidate, mode, delta, ci95, wins, losses, ties Paired comparisons against the baseline. delta and ci95 use fractions in JSON: 0.02 means +2 percentage points.

Retrieval results

Documents Queries Mode Hit@1 Hit@3 Hit@5 Hit@10 MRR
none none lexical 52.5% 73.7% 81.0% 89.9% 0.656
voyage-4-large voyage-4-large vector 77.7% 94.4% 97.2% 98.3% 0.862
voyage-4-large voyage-4-large hybrid 69.3% 89.4% 93.9% 97.8% 0.799
voyage-4 voyage-4 vector 70.4% 93.3% 98.3% 99.4% 0.818
voyage-4 voyage-4 hybrid 66.5% 91.1% 93.9% 98.3% 0.787
voyage-4-lite voyage-4-lite vector 67.6% 95.0% 97.2% 98.9% 0.810
voyage-4-lite voyage-4-lite hybrid 68.2% 91.6% 93.9% 97.2% 0.800
embed-v5.0-pro embed-v5.0-pro vector 73.7% 93.3% 97.2% 98.9% 0.838
embed-v5.0-pro embed-v5.0-pro hybrid 67.0% 91.6% 93.9% 97.2% 0.794
embed-v5.0-fast embed-v5.0-fast vector 69.8% 92.7% 97.2% 98.9% 0.812
embed-v5.0-fast embed-v5.0-fast hybrid 66.5% 91.1% 95.0% 97.2% 0.791
voyage-4-large voyage-4-lite vector 69.3% 94.4% 97.2% 98.9% 0.817
voyage-4-large voyage-4-lite hybrid 69.8% 89.9% 93.9% 97.8% 0.808
embed-v5.0-pro embed-v5.0-fast vector 75.4% 92.7% 97.2% 98.9% 0.846
embed-v5.0-pro embed-v5.0-fast hybrid 69.3% 91.6% 94.4% 97.2% 0.806

Original and added queries

Success metrics exclude no-answer cases. Original means the unchanged 94-row CI fixture; additional means the 106 new cases labeled against the same frozen docs before running the models. These are author-labeled cases, not an independently judged held-out set.

Documents Queries Mode Split Answerable Hit@3 MRR
none none lexical original 83 72.3% 0.650
none none lexical additional 96 75.0% 0.661
voyage-4-large voyage-4-large vector original 83 92.8% 0.878
voyage-4-large voyage-4-large vector additional 96 95.8% 0.848
voyage-4-large voyage-4-large hybrid original 83 86.7% 0.817
voyage-4-large voyage-4-large hybrid additional 96 91.7% 0.783
voyage-4 voyage-4 vector original 83 92.8% 0.865
voyage-4 voyage-4 vector additional 96 93.8% 0.778
voyage-4 voyage-4 hybrid original 83 90.4% 0.808
voyage-4 voyage-4 hybrid additional 96 91.7% 0.770
voyage-4-lite voyage-4-lite vector original 83 95.2% 0.838
voyage-4-lite voyage-4-lite vector additional 96 94.8% 0.786
voyage-4-lite voyage-4-lite hybrid original 83 91.6% 0.813
voyage-4-lite voyage-4-lite hybrid additional 96 91.7% 0.788
embed-v5.0-pro embed-v5.0-pro vector original 83 95.2% 0.877
embed-v5.0-pro embed-v5.0-pro vector additional 96 91.7% 0.805
embed-v5.0-pro embed-v5.0-pro hybrid original 83 91.6% 0.800
embed-v5.0-pro embed-v5.0-pro hybrid additional 96 91.7% 0.788
embed-v5.0-fast embed-v5.0-fast vector original 83 95.2% 0.847
embed-v5.0-fast embed-v5.0-fast vector additional 96 90.6% 0.782
embed-v5.0-fast embed-v5.0-fast hybrid original 83 90.4% 0.785
embed-v5.0-fast embed-v5.0-fast hybrid additional 96 91.7% 0.796
voyage-4-large voyage-4-lite vector original 83 95.2% 0.839
voyage-4-large voyage-4-lite vector additional 96 93.8% 0.799
voyage-4-large voyage-4-lite hybrid original 83 89.2% 0.819
voyage-4-large voyage-4-lite hybrid additional 96 90.6% 0.799
embed-v5.0-pro embed-v5.0-fast vector original 83 94.0% 0.882
embed-v5.0-pro embed-v5.0-fast vector additional 96 91.7% 0.814
embed-v5.0-pro embed-v5.0-fast hybrid original 83 91.6% 0.820
embed-v5.0-pro embed-v5.0-fast hybrid additional 96 91.7% 0.793

Paired comparisons

Differences are candidate minus the current Voyage Large/Lite baseline. The 95% intervals use 2,000 seeded paired bootstrap samples of answerable query rows. Related queries are not independent; these intervals describe this labeled set and do not establish general model superiority.

Candidate Mode Hit@3 difference (percentage points) 95% interval (percentage points) Wins Losses
voyage-4-large / voyage-4-large vector 0.0 -2.8 to 2.8 3 3
voyage-4-large / voyage-4-large hybrid -0.6 -2.8 to 1.1 1 2
voyage-4 / voyage-4 vector -1.1 -5.0 to 2.2 4 6
voyage-4 / voyage-4 hybrid 1.1 0.0 to 2.8 2 0
voyage-4-lite / voyage-4-lite vector 0.6 -1.7 to 2.8 3 2
voyage-4-lite / voyage-4-lite hybrid 1.7 0.0 to 3.9 3 0
embed-v5.0-pro / embed-v5.0-pro vector -1.1 -4.5 to 2.2 4 6
embed-v5.0-pro / embed-v5.0-pro hybrid 1.7 -1.7 to 5.0 6 3
embed-v5.0-fast / embed-v5.0-fast vector -1.7 -6.1 to 2.2 6 9
embed-v5.0-fast / embed-v5.0-fast hybrid 1.1 -1.7 to 4.5 5 3
embed-v5.0-pro / embed-v5.0-fast vector -1.7 -5.0 to 1.7 4 7
embed-v5.0-pro / embed-v5.0-fast hybrid 1.7 -1.7 to 5.0 6 3

Live embedding timing

Document and quality-query requests use batches of at most 32. Batched timings measure request time, not interactive query latency or total indexing wall time. Fresh single-query samples bypass the cache; they include API/network time and any provider retries, but exclude retrieval time and intentional client rate-limit pacing. Models run sequentially, so service conditions and request order can affect timings.

Model Document requests, live/cached Live document ms Query requests, live/cached Single-query samples p50 ms p95 ms
voyage-4-large 0/35 0 3/4 24 174 311
voyage-4 0/35 0 3/4 24 149 183
voyage-4-lite 0/35 0 5/2 24 148 162
embed-v5.0-pro 0/35 0 5/2 24 110 120
embed-v5.0-fast 0/35 0 5/2 24 97 101

Limits and reproducibility

Corpus and query fingerprints are retained in the detailed JSON report. Each passage and query has a separate vector for each model; cross-model retrieval is tested only within a vendor’s documented shared embedding space.

Some questions already appear in the project’s automated search tests; others were added for this experiment. These results tell me how the models performed on my docs, not which model is best everywhere. The full results JSON includes each question’s rankings and the detailed measurements. Running the benchmark didn’t change the site’s search.

Model references: Cohere Embed 5, Cohere Embed API, Voyage embeddings.