Reranking with a classifier instead of a chat model

TL;DR

A search in Stuga returns 24 candidate passages, and a judge reorders them so that the model reads the best eight. We compared 23 judges on two public benchmarks whose questions and answers were marked by people. Every judge reordered the same 24 candidates and saw each one’s title, heading path and 1,200 characters of its text. The main measure is Hit@1: how often the first passage is one that answers the question.

Questions that ask for facts. MIRACL has questions over Wikipedia in six languages; we took 240 and reordered the 24 candidates MMTEB publishes for each. Without a judge, a relevant passage came first for 36% of questions. TypeSafe’s Jev, a System One classifier, raised that to 77%, in 131 ms for $0.37 per 1,000 searches, and put a relevant passage among the eight the model reads for 99% of questions, against 80%. The best LLM judge, Claude Opus 5.5, reached 81%. The 4.2 points between them are within this test’s margin of error (−2.9 to 10.8 points, adjusted for the 23 comparisons with Jev), so the data say neither that one is better nor that they are equal. Claude Opus 5.5 took 27 times as long and cost 100 times as much. Among the newest models tested, Claude Sonnet 5.5 reached 80% in 2.0 s, at $20 per 1,000 searches; GPT-6.1 Sol reached 77% in 2.2 s, at $14.

Questions that need reasoning. BRIGHT’s 117 StackOverflow questions are answered by documentation pages that may share few words with the question. Search found such a page among the 24 candidates for only 52% of questions, and its own order put one first for 19% of all 117. The passages are long, typically 4,000 characters, and of the judges reading the first 1,200 of each, as Stuga’s did until this study, none clearly beat search’s order. Reading instead the 1,200 characters that best match the question, 21 of 24 judges rose, 7 of them clearly: over all 117 questions, Claude Opus 5.5 then put a gold passage first for 32% and Jev for 25%. We designed that reading after seeing the first results and measured it on the same questions, so it is exploratory; Stuga’s judges now read that part (Section 4.3).

In this setup, Jev was a fast, cheap judge for questions that ask for facts, and no judge we tested clearly beat it. On long passages, what a judge reads of each mattered as much as which judge it is.

We ran this test to decide what Stuga’s Ask and its tools for agents show an AI before it answers: which judge orders the candidates, and which part of each passage it reads. Section 6 shows the result in a workspace.

60%65%70%75%80%85%90%↑ Right passage first100 ms300 ms1.0 s3.0 s10 s30 sMedian time per search (log scale) →Amazon Rerank 1.0 Reranker, Bedrock us-west-2 Hit@1 77% (95% CI 71%–82%) Median 435 msCohere Rerank 3.5 Reranker, Bedrock us-west-2 Hit@1 78% (95% CI 73%–83%) Median 163 ms · $2.0 per 1,000 searchesAmazon Nova 2 Lite LLM judge, Bedrock us-west-2 Hit@1 62% (95% CI 56%–68%) Median 1.3 s · $2.2 per 1,000 searchesMistral Large 3 LLM judge, Bedrock us-west-2 Hit@1 66% (95% CI 60%–72%) Median 7.0 s · $3.8 per 1,000 searchesMiniMax M2.5 LLM judge, Bedrock us-west-2 Hit@1 71% (95% CI 66%–77%) Median 8.6 s · $4.4 per 1,000 searchesDeepSeek V3.2 LLM judge, Bedrock us-west-2 Hit@1 73% (95% CI 67%–78%) Median 5.9 s · $3.2 per 1,000 searchesLlama 4 Maverick LLM judge, Bedrock us-west-2 Hit@1 73% (95% CI 67%–78%) Median 945 ms · $1.2 per 1,000 searchesGPT-6 Luna LLM judge, OpenAI API Hit@1 73% (95% CI 67%–78%) Median 1.5 s · $0.52 per 1,000 searchesgpt-oss-120b LLM judge, Bedrock us-west-2 Hit@1 74% (95% CI 68%–79%) Median 9.7 s · $1.2 per 1,000 searchesGPT-6 Sol LLM judge, OpenAI API Hit@1 75% (95% CI 70%–80%) Median 2.3 s · $11 per 1,000 searchesQwen3 235B A22B 2507 LLM judge, Bedrock us-west-2 Hit@1 75% (95% CI 69%–80%) Median 3.6 s · $1.2 per 1,000 searchesGrok 4.6 LLM judge, Bedrock us-west-2 Hit@1 76% (95% CI 70%–81%) Median 12 s · $20 per 1,000 searchesGPT-6.1 Sol LLM judge, Bedrock us-west-2 Hit@1 77% (95% CI 72%–82%) Median 2.2 s · $14 per 1,000 searchesGLM-5 LLM judge, Bedrock us-west-2 Hit@1 77% (95% CI 72%–83%) Median 2.7 s · $6.1 per 1,000 searchesClaude Haiku 4.5 LLM judge, Bedrock us-west-2 Hit@1 79% (95% CI 74%–84%) Median 1.9 s · $9.3 per 1,000 searchesClaude Sonnet 5.5 LLM judge, Bedrock us-west-2 Hit@1 80% (95% CI 75%–85%) Median 2.0 s · $20 per 1,000 searchesClaude Sonnet 5 LLM judge, Bedrock us-west-2 Hit@1 80% (95% CI 75%–85%) Median 3.7 s · $20 per 1,000 searchesClaude Opus 5.5 LLM judge, Bedrock us-west-2 Hit@1 81% (95% CI 76%–86%) Median 3.5 s · $37 per 1,000 searchesJev System One classifier, TypeSafe API Hit@1 77% (95% CI 72%–82%) Median 131 ms · $0.37 per 1,000 searchesJev, one request per passage System One classifier, TypeSafe API Hit@1 78% (95% CI 73%–83%) Median 244 ms · $0.64 per 1,000 searchesJevClaude Sonnet 5.5Claude Opus 5.5Cohere Rerank 3.5Claude Sonnet 5GPT-6 SolGPT-6.1 SolGrok 4.6Amazon Rerank 1.0Llama 4 MaverickAmazon Nova 2 LiteGPT-6 LunaMistral Large 3Claude Haiku 4.5
60%65%70%75%80%85%90%↑ Right passage first100 ms1.0 s10 sMedian time per search (log scale) →Amazon Rerank 1.0 Reranker, Bedrock us-west-2 Hit@1 77% (95% CI 71%–82%) Median 435 msCohere Rerank 3.5 Reranker, Bedrock us-west-2 Hit@1 78% (95% CI 73%–83%) Median 163 ms · $2.0 per 1,000 searchesAmazon Nova 2 Lite LLM judge, Bedrock us-west-2 Hit@1 62% (95% CI 56%–68%) Median 1.3 s · $2.2 per 1,000 searchesMistral Large 3 LLM judge, Bedrock us-west-2 Hit@1 66% (95% CI 60%–72%) Median 7.0 s · $3.8 per 1,000 searchesMiniMax M2.5 LLM judge, Bedrock us-west-2 Hit@1 71% (95% CI 66%–77%) Median 8.6 s · $4.4 per 1,000 searchesDeepSeek V3.2 LLM judge, Bedrock us-west-2 Hit@1 73% (95% CI 67%–78%) Median 5.9 s · $3.2 per 1,000 searchesLlama 4 Maverick LLM judge, Bedrock us-west-2 Hit@1 73% (95% CI 67%–78%) Median 945 ms · $1.2 per 1,000 searchesGPT-6 Luna LLM judge, OpenAI API Hit@1 73% (95% CI 67%–78%) Median 1.5 s · $0.52 per 1,000 searchesgpt-oss-120b LLM judge, Bedrock us-west-2 Hit@1 74% (95% CI 68%–79%) Median 9.7 s · $1.2 per 1,000 searchesGPT-6 Sol LLM judge, OpenAI API Hit@1 75% (95% CI 70%–80%) Median 2.3 s · $11 per 1,000 searchesQwen3 235B A22B 2507 LLM judge, Bedrock us-west-2 Hit@1 75% (95% CI 69%–80%) Median 3.6 s · $1.2 per 1,000 searchesGrok 4.6 LLM judge, Bedrock us-west-2 Hit@1 76% (95% CI 70%–81%) Median 12 s · $20 per 1,000 searchesGPT-6.1 Sol LLM judge, Bedrock us-west-2 Hit@1 77% (95% CI 72%–82%) Median 2.2 s · $14 per 1,000 searchesGLM-5 LLM judge, Bedrock us-west-2 Hit@1 77% (95% CI 72%–83%) Median 2.7 s · $6.1 per 1,000 searchesClaude Haiku 4.5 LLM judge, Bedrock us-west-2 Hit@1 79% (95% CI 74%–84%) Median 1.9 s · $9.3 per 1,000 searchesClaude Sonnet 5.5 LLM judge, Bedrock us-west-2 Hit@1 80% (95% CI 75%–85%) Median 2.0 s · $20 per 1,000 searchesClaude Sonnet 5 LLM judge, Bedrock us-west-2 Hit@1 80% (95% CI 75%–85%) Median 3.7 s · $20 per 1,000 searchesClaude Opus 5.5 LLM judge, Bedrock us-west-2 Hit@1 81% (95% CI 76%–86%) Median 3.5 s · $37 per 1,000 searchesJev System One classifier, TypeSafe API Hit@1 77% (95% CI 72%–82%) Median 131 ms · $0.37 per 1,000 searchesJev, one request per passage System One classifier, TypeSafe API Hit@1 78% (95% CI 73%–83%) Median 244 ms · $0.64 per 1,000 searchesJevSonnet 5.5Claude Opus 5.5CohereSonnet 5GPT-6 SolGrok 4.6Llama 4Nova 2 LiteMistralHaiku 4.5
Figure 1. Hosted rankers on MMTEB's MIRACL candidate lists: share of questions with a relevant passage ranked first, against median latency per search. First-stage order: 36%.

1. Introduction

Stuga answers questions from a workspace’s documents (Ask) and serves the same retrieval to agents over MCP. Retrieval returns 24 candidate passages; a judge reorders them, and the model reads the first eight. The judge decides what the model sees, so we measured the options on public benchmarks, in the shape Stuga runs them. The primary comparison is Jev, which Stuga’s Reranking setting uses, against the LLM judges Stuga falls back to. The contribution is an empirical comparison through the interfaces a product calls, not a new ranker:

Five questions guided the study:

  1. How often does each judge put a relevant passage first, and among the eight?
  2. At what latency and cost?
  3. How do LLM judges fail?
  4. Can the judge run on the machine that runs Stuga?
  5. What should a judge read of a long passage?

2. Background

2.1 Retrieval in Stuga

Stuga splits each document at its headings: every section is a passage, split further only past 12,000 characters, and a document without headings is cut into pieces of about 6,000. Each passage gets one vector, from its heading path and text. A question runs two searches in one SQL statement over these passages: BM25 [5] through pg_search and vector similarity through pgvector, 96 passages each. Reciprocal rank fusion [6] (k = 60) merges the lists and keeps 24 (Figure 2).

Question from Ask or an agent Keywords (BM25) pg_search · top 96 Meaning (vectors) pgvector · top 96 Fusion (RRF) 24 candidates Judge reorders all 24 what this post measures ≤ 3 per document 8 passages Model Question from Ask or an agent Keywords (BM25) pg_search · top 96 Meaning (vectors) pgvector · top 96 Fusion (RRF) 24 candidates Judge reorders all 24 what this post measures ≤ 3 per document 8 passages Model writes the answer
Figure 2. From question to the eight passages the model reads.

Three terms recur below. The first stage is the search that picks the 24 candidates; its order, with no judge, is the baseline every judge is measured against. A relevant passage is one the benchmark’s annotators marked as answering the question; BRIGHT calls it gold. Hit@1 is the share of questions whose first passage is relevant.

A first stage finds the neighbourhood, not the house: a passage about the question’s subject can outrank the one that answers it. On MIRACL’s candidate lists a relevant passage came first for 36% of questions and was among the first eight for 80%. Figure 3 follows three of them.

“When was the town of Simcoe, Ontario founded?”

No judge: first-stage order Answer #22 of 24

  1. 1 Barrie (electoral district) Prior to the 2015 election Barrie was a federal electoral district in Ontario, Canada, that has been represented in the House of Commons of Canada since
  2. 2 Simcoe, Ontario A cultural club for people of Croatian descent operates in this town; the formal name given to this organization is the 531st branch of the Croatian Fraternal Union
  3. 3 Simcoe, Ontario Simcoe is an unincorporated community and former town in Southwestern Ontario, Canada near Lake Erie. It is the county seat and largest community of Norfolk County.
  4. 4 Elizabeth Simcoe Elizabeth Simcoe left a diary that provides a valuable impression of life in colonial Ontario. First published in 1934, there was a subsequent transcription publis
  5. 5 History of Markham, Ontario When Upper and Lower Canada were established in 1791, Colonel John Graves Simcoe was appointed the first Lieutenant-Governor of Upper Canada. Simcoe nam
  6. 6 Simcoe, Ontario Simcoe was incorporated as a town in 1878 and had its own town council and mayor until December 31, 2000. In 2001, the town and all other municipalities within the
  7. 7 Simcoe, Ontario Rural Canadian towns similar to Simcoe are close to dying due to economic and transportation issues that prevent people from holding meaningful employment and being
  8. 8 177th Battalion (Simcoe Foresters), CEF The 177th (Simcoe Foresters) Battalion, CEF was a unit in the Canadian Expeditionary Force raised during the First World War by the 35th Sim

The answer is #22: the model never reads it.

Jev Answer #1 of 24 · 105 ms

  1. 1 Simcoe, Ontario Simcoe was founded in 1795 by Lieutenant Governor John Graves Simcoe. Initially, the settlement consisted of two distinct areas, Birdtown, named by William Bird who Answer0.98
  2. 2 Communities in Norfolk County, Ontario The population in 1850 was about 1600; in that year, Simcoe became the County seat of Norfolk County. Simcoe was incorporated as a town in 18 0.45
  3. 3 Simcoe, Ontario Simcoe was incorporated as a town in 1878 and had its own town council and mayor until December 31, 2000. In 2001, the town and all other municipalities within the 0.42
  4. 4 History of Markham, Ontario When Upper and Lower Canada were established in 1791, Colonel John Graves Simcoe was appointed the first Lieutenant-Governor of Upper Canada. Simcoe nam 0.07
  5. 5 Simcoe, Ontario Simcoe is an unincorporated community and former town in Southwestern Ontario, Canada near Lake Erie. It is the county seat and largest community of Norfolk County. 0.06
  6. 6 Elizabeth Simcoe Elizabeth Simcoe left a diary that provides a valuable impression of life in colonial Ontario. First published in 1934, there was a subsequent transcription publis 0.05
  7. 7 Communities in Norfolk County, Ontario Simcoe is the administrative centre of Norfolk County, with a population of 16,000 making it Norfolk's largest community. Simcoe is located a 0.05
  8. 8 Henry Dundas, 1st Viscount Melville He was friends with John Graves Simcoe, Lieutenant Governor of Upper Canada. Simcoe named the town of Dundas, Ontario, in southern Ontario after 0.05
Figure 3. The first eight of 24 candidates for three MIRACL questions, without a judge (left) and after one (right). Scores are each judge's own: a probability, a 0–10 score, a relevance score.

2.2 Three kinds of judge

Table 1. The judges compared.

LLM judge Reranker System One classifier
Tested GPT-6, Claude, Grok, Qwen, Llama and others Cohere Rerank, Amazon Rerank, Qwen3-Reranker Jev, Kev, Laya
Reads all 24 passages in one prompt the question with one passage at a time the passages as a state, one question each
Returns text: a JSON list of 0–10 scores a relevance score a probability
Built on a generative LLM classically a BERT-style encoder [2, 3]; Qwen3-Reranker on the Qwen3 LLM Jev: not published. Kev: Qwen3.5 with a LoRA and a decision head. Laya: ModernBERT [7] and mmBERT

An LLM judge scores passages by generating text, the approach RankGPT [4] popularised. A reranker is a cross-encoder in the tradition of monoBERT [3]: the question and one passage pass through the model together and a head outputs a score. Cohere and Amazon do not publish their architectures. A System One classifier answers typed questions about a state with probabilities in a single call. The name follows Kahneman’s fast, intuitive System 1 [1]; a model that reasons before answering is the deliberate System 2.

Stuga’s LLM judge sends this system prompt, with each candidate’s title, heading path and the 1,200 characters of its text that best match the question (Section 4.3):

You are a search relevance judge. Given a user query and numbered
document snippets, rate how well EACH snippet helps answer or act on the query.
Score each 0-10 (10 = directly answers it; 0 = irrelevant). Judge only relevance,
not writing quality. Respond with ONLY a JSON array of {"i":<number>,"score":<0-10>}
for every snippet, no prose.

Its System One judge sends the same 24 passages to Jev [8] in one request:

{
"model": "jev-latest",
"state": { "query": "…", "passages": { "p0": { "title": "…", "section": "…", "text": "…" }, "…": {} } },
"questions": {
"p0": {
"type": "noul",
"instructions": "Does passage p0 help answer the query?",
"criteria": {
"true": "The passage states the information the query asks for, or information needed to answer it",
"false": "The passage is only on a related topic, or is irrelevant to the query"
}
}
}
}

3. Method

3.1 Test sets

We used two public benchmarks whose questions and relevance judgments were made by people, so that neither the questions nor the labels came from us or from a model under test. Both are openly licensed and published with the benchmark.

3.2 Candidates

Every ranker reordered the same 24 candidates per question. For MIRACL they are the first 24 of MMTEB’s published list for the question, in its order. Those lists hold a relevant passage in their first 24 for 79% of the six languages’ dev questions, and we drew ours from those.

For BRIGHT we retrieved the candidates as Stuga does, from the question alone, over the subset’s 107,081 passages, as BRIGHT splits them, in Stuga’s two configurations. With embeddings, BM25 and vector search each return 96 passages, fused by reciprocal rank (k = 60), with one vector per passage from OpenAI’s text-embedding-3-large at 1,024 dimensions; without them, BM25’s first 24 stand. A gold passage was among the 24 for 61 of the 117 questions (52%) with embeddings and for 47 (40%) with BM25 alone. The rest are retrieval misses that no judge can fix, so BRIGHT’s scores below count only the questions with a gold passage among the candidates, unless they say all 117. Over all 117, no ranker put a gold passage first for more than 22%.

3.3 Judges

We compared current models across the listed families. OpenAI judges used OpenAI’s API or Bedrock’s OpenAI-compatible endpoint, TypeSafe its own API; the other hosted models ran on Bedrock in us-west-2, on US or global inference profiles. LLM judges received Stuga’s prompt verbatim, at the lowest reasoning each model offers (low for GPT-6.1 Sol, none for GPT-6 Sol) and with up to 8,192 output tokens, as Stuga’s judge sends them. Jev received Stuga’s single request and, separately, one request per passage. Kev and Laya, which serve the same System One API on local hardware, received one passage per request: Kev’s README notes training on states of up to 384 tokens, and Laya, an encoder, reads its state as text, so it received plain text and chose its English or multilingual checkpoint itself. Rerankers scored each question and passage pair.

Every judge saw each candidate’s title, heading path and 1,200 characters of its text: the first 1,200, as Stuga’s judge read them until this study, and on BRIGHT also the 1,200 that best match the question, as it reads them now (Section 4.3). On MIRACL’s short passages the two differ for 1% of candidates, so its results stand for both.

3.4 Measures

3.5 Environment

The benchmark client and every local model ran on one Mac: Apple M6, 32 GB, macOS 27. Qwen3-Reranker ran as Q8_0 GGUF in llama.cpp 0.5.0, Kev-4B on MLX 0.32 in bfloat16, and Laya 0.3.20 on PyTorch 2.14, one model at a time.

4. Results

Table 2. All rankers. Hit@1 on MIRACL and on BRIGHT with embeddings, BRIGHT's judges reading the 1,200 characters of each passage that best match the question (Section 4.3), and on MIRACL the difference from Jev on the same questions in points with its unadjusted 95% interval; latency, cost and failures on MIRACL. Where: Cloud or Local. Released: see Section 3.4.
Claude Opus 5.5 LLM judge Cloud Sept 2026 81% (76%–86%) +4.2 (0.0 to +8.8) 61% (48%–72%) 3.5 s $37 0
Qwen3-Reranker 4B Reranker Local Jun 2025 81% (75%–86%) +3.8 (−2.5 to +10.0) 23% (13%–34%) 5.3 s local 0
Claude Sonnet 5.5 LLM judge Cloud Sept 2026 80% (75%–85%) +2.5 (−2.1 to +7.1) 54% (41%–67%) 2.0 s $20 0
Claude Sonnet 5 LLM judge Cloud Jun 2026 80% (75%–85%) +2.5 (−2.1 to +7.1) 56% (43%–69%) 3.7 s $20 0
Claude Haiku 4.5 LLM judge Cloud Oct 2025 79% (74%–84%) +2.1 (−2.9 to +6.7) 39% (26%–52%) 1.9 s $9.3 0
Cohere Rerank 3.5 Reranker Cloud Dec 2024 78% (73%–83%) +1.3 (−4.6 to +6.7) 23% (13%–33%) 163 ms $2.0 0
Jev, one request per passage Classifier Cloud Sept 2026 78% (73%–83%) +1.3 (−3.3 to +5.4) 46% (33%–59%) 244 ms $0.64 0
GPT-6.1 Sol LLM judge Cloud Sept 2026 77% (72%–82%) 0.0 (−5.0 to +4.6) 52% (39%–66%) 2.2 s $14 0
Jev Classifier Cloud Sept 2026 77% (72%–82%) — 48% (34%–61%) 131 ms $0.37 0
GLM-5 LLM judge Cloud Feb 2026 77% (72%–83%) 0.0 (−4.2 to +3.8) 51% (39%–64%) 2.7 s $6.1 0
Amazon Rerank 1.0 Reranker Cloud Dec 2024 77% (71%–82%) −0.4 (−6.3 to +5.4) 20% (10%–30%) 435 ms not listed 0
Grok 4.6 LLM judge Cloud Aug 2026 76% (70%–81%) −1.3 (−6.3 to +3.3) 56% (43%–67%) 12 s $20 0
GPT-6 Sol LLM judge Cloud Sept 2026 75% (70%–80%) −2.1 (−7.1 to +2.9) 49% (36%–62%) 2.3 s $11 0
Qwen3 235B A22B 2507 LLM judge Cloud Jul 2025 75% (69%–80%) −2.1 (−7.5 to +2.5) 44% (31%–56%) 3.6 s $1.2 0
Qwen3-Reranker 0.6B Reranker Local May 2025 75% (69%–80%) −2.5 (−9.6 to +4.2) 28% (18%–39%) 987 ms local 0
gpt-oss-120b LLM judge Cloud Aug 2025 74% (68%–79%) −3.3 (−7.5 to +0.8) 51% (38%–64%) 9.7 s $1.2 0
Llama 4 Maverick LLM judge Cloud Apr 2025 73% (67%–78%) −4.2 (−9.2 to +1.3) 41% (30%–54%) 945 ms $1.2 0
GPT-6 Luna LLM judge Cloud Sept 2026 73% (67%–78%) −4.2 (−9.6 to +1.3) 51% (38%–64%) 1.5 s $0.52 11%
DeepSeek V3.2 LLM judge Cloud Dec 2025 73% (67%–78%) −4.6 (−10.0 to +0.8) 54% (41%–67%) 5.9 s $3.2 3%
MiniMax M2.5 LLM judge Cloud Feb 2026 71% (66%–77%) −5.8 (−10.8 to −0.4) 44% (31%–57%) 8.6 s $4.4 10%
Mistral Large 3 LLM judge Cloud Dec 2025 66% (60%–72%) −11.3 (−17.1 to −5.4) 44% (31%–57%) 7.0 s $3.8 9%
Kev-4B Classifier Local Sept 2026 64% (58%–70%) −12.9 (−19.2 to −6.3) 30% (18%–41%) 5.4 s local 0
Amazon Nova 2 Lite LLM judge Cloud Dec 2025 62% (56%–68%) −15.4 (−21.3 to −9.6) 52% (39%–64%) 1.3 s $2.2 0
Laya Classifier Local Sept 2026 49% (43%–55%) −27.9 (−34.6 to −20.4) 16% (8%–26%) 646 ms local 0
No reranker (first-stage order) — — — 36% (30%–43%) −40.8 (−48.3 to −33.8) 36% (25%–48%) — — 0

4.1 MIRACL: questions that ask for facts

The candidate lists here are MMTEB’s, so this section measures reranking a fixed list; it says nothing about how Stuga’s own retrieval would have ranked MIRACL. Reranking lifted Hit@1 from 36% to at most 81%, Claude Opus 5.5’s. Table 2 gives each ranker’s difference from Jev on the same 240 questions. Claude Opus 5.5 was 4.2 points ahead (0.0 to 8.8, or −2.9 to 10.8 adjusted for the 23 comparisons with Jev). Another 17 rankers, of all three kinds, scored between 73% and 81%, each within the margin of error of Jev even before adjustment. Adjusted, Mistral Large 3, Amazon Nova 2 Lite, Kev-4B and Laya remain clearly below Jev. What separates the leaders is latency and cost (Figure 1): Jev and Cohere Rerank 3.5 answer in 131 ms and 163 ms, the LLM judges among them 945 ms to 12 s. With any of these rankers a relevant passage is among the eight the model reads for 93%–100% of questions, against 80% without a judge.

4.2 By language

Table 3. MIRACL Hit@1 by language, 40 questions each; the best in each row in bold.
Language n JevClaude Opus 5.5GPT-6 SolCohere Rerank 3.5Qwen3-Reranker 4B (local)
English 40 70%75%70%73%85%
Chinese 40 70%88%78%75%78%
Japanese 40 80%75%75%80%80%
Korean 40 75%75%73%70%75%
Arabic 40 83%90%78%88%88%
Thai 40 85%85%78%85%80%

With 40 questions per language, a difference of one question is 2.5 points, and no ranker is best in every language.

4.3 BRIGHT: questions that need reasoning

0% 10% 20% 30% 40% 50% 60% 70% No reranker 36% Hit@1, share of questions with a gold passage first Claude Opus 5.5: 39% reading the first 1,200 characters, 61% reading the 1,200 that best match (+21.3 points, 95% CI +11.5 to +32.8) Claude Opus 5.5 39% → 61% Claude Sonnet 5: 34% reading the first 1,200 characters, 56% reading the 1,200 that best match (+21.3 points, 95% CI +8.2 to +36.1) Claude Sonnet 5 34% → 56% Grok 4.6: 31% reading the first 1,200 characters, 56% reading the 1,200 that best match (+24.6 points, 95% CI +13.1 to +37.7) Grok 4.6 31% → 56% Claude Sonnet 5.5: 38% reading the first 1,200 characters, 54% reading the 1,200 that best match (+16.4 points, 95% CI +3.3 to +29.5) Claude Sonnet 5.5 38% → 54% DeepSeek V3.2: 25% reading the first 1,200 characters, 54% reading the 1,200 that best match (+29.5 points, 95% CI +18.0 to +42.6) DeepSeek V3.2 25% → 54% Amazon Nova 2 Lite: 36% reading the first 1,200 characters, 52% reading the 1,200 that best match (+16.4 points, 95% CI +3.3 to +31.1) Amazon Nova 2 Lite 36% → 52% GPT-6.1 Sol: 33% reading the first 1,200 characters, 52% reading the 1,200 that best match (+19.7 points, 95% CI +9.8 to +31.1) GPT-6.1 Sol 33% → 52% GPT-6 Luna: 43% reading the first 1,200 characters, 51% reading the 1,200 that best match (+8.2 points, 95% CI −6.6 to +23.0) GPT-6 Luna 43% → 51% GLM-5: 36% reading the first 1,200 characters, 51% reading the 1,200 that best match (+14.8 points, 95% CI +1.6 to +27.9) GLM-5 36% → 51% gpt-oss-120b: 25% reading the first 1,200 characters, 51% reading the 1,200 that best match (+26.2 points, 95% CI +13.1 to +41.0) gpt-oss-120b 25% → 51% GPT-6 Sol: 28% reading the first 1,200 characters, 49% reading the 1,200 that best match (+21.3 points, 95% CI +9.8 to +32.8) GPT-6 Sol 28% → 49% Jev: 31% reading the first 1,200 characters, 48% reading the 1,200 that best match (+16.4 points, 95% CI +6.6 to +27.9) Jev 31% → 48% Jev, one request per passage: 31% reading the first 1,200 characters, 46% reading the 1,200 that best match (+14.8 points, 95% CI +4.9 to +26.2) Jev, one request per passage 31% → 46% MiniMax M2.5: 34% reading the first 1,200 characters, 44% reading the 1,200 that best match (+9.8 points, 95% CI −6.6 to +26.2) MiniMax M2.5 34% → 44% Qwen3 235B A22B 2507: 33% reading the first 1,200 characters, 44% reading the 1,200 that best match (+11.5 points, 95% CI −1.6 to +24.6) Qwen3 235B A22B 2507 33% → 44% Mistral Large 3: 30% reading the first 1,200 characters, 44% reading the 1,200 that best match (+14.8 points, 95% CI +0.0 to +27.9) Mistral Large 3 30% → 44% Llama 4 Maverick: 36% reading the first 1,200 characters, 41% reading the 1,200 that best match (+4.9 points, 95% CI −6.6 to +16.4) Llama 4 Maverick 36% → 41% Claude Haiku 4.5: 33% reading the first 1,200 characters, 39% reading the 1,200 that best match (+6.6 points, 95% CI −3.3 to +18.0) Claude Haiku 4.5 33% → 39% Kev-4B (local): 31% reading the first 1,200 characters, 30% reading the 1,200 that best match (−1.6 points, 95% CI −13.1 to +8.2) Kev-4B 31% → 30% Qwen3-Reranker 0.6B (local): 21% reading the first 1,200 characters, 28% reading the 1,200 that best match (+6.6 points, 95% CI −3.3 to +16.4) Qwen3-Reranker 0.6B 21% → 28% Qwen3-Reranker 4B (local): 23% reading the first 1,200 characters, 23% reading the 1,200 that best match (+0.0 points, 95% CI −11.5 to +11.5) Qwen3-Reranker 4B 23% → 23% Cohere Rerank 3.5: 13% reading the first 1,200 characters, 23% reading the 1,200 that best match (+9.8 points, 95% CI +0.0 to +19.7) Cohere Rerank 3.5 13% → 23% Amazon Rerank 1.0: 16% reading the first 1,200 characters, 20% reading the 1,200 that best match (+3.3 points, 95% CI −4.9 to +13.1) Amazon Rerank 1.0 16% → 20% Laya (local): 23% reading the first 1,200 characters, 16% reading the 1,200 that best match (−6.6 points, 95% CI −14.8 to +1.6) Laya 23% → 16% 0% 20% 40% 60% No reranker 36% Hit@1, share of questions with a gold passage first Claude Opus 5.5: 39% reading the first 1,200 characters, 61% reading the 1,200 that best match (+21.3 points, 95% CI +11.5 to +32.8) Claude Opus 5.5 39% → 61% Claude Sonnet 5: 34% reading the first 1,200 characters, 56% reading the 1,200 that best match (+21.3 points, 95% CI +8.2 to +36.1) Claude Sonnet 5 34% → 56% Grok 4.6: 31% reading the first 1,200 characters, 56% reading the 1,200 that best match (+24.6 points, 95% CI +13.1 to +37.7) Grok 4.6 31% → 56% Claude Sonnet 5.5: 38% reading the first 1,200 characters, 54% reading the 1,200 that best match (+16.4 points, 95% CI +3.3 to +29.5) Claude Sonnet 5.5 38% → 54% DeepSeek V3.2: 25% reading the first 1,200 characters, 54% reading the 1,200 that best match (+29.5 points, 95% CI +18.0 to +42.6) DeepSeek V3.2 25% → 54% Amazon Nova 2 Lite: 36% reading the first 1,200 characters, 52% reading the 1,200 that best match (+16.4 points, 95% CI +3.3 to +31.1) Amazon Nova 2 Lite 36% → 52% GPT-6.1 Sol: 33% reading the first 1,200 characters, 52% reading the 1,200 that best match (+19.7 points, 95% CI +9.8 to +31.1) GPT-6.1 Sol 33% → 52% GPT-6 Luna: 43% reading the first 1,200 characters, 51% reading the 1,200 that best match (+8.2 points, 95% CI −6.6 to +23.0) GPT-6 Luna 43% → 51% GLM-5: 36% reading the first 1,200 characters, 51% reading the 1,200 that best match (+14.8 points, 95% CI +1.6 to +27.9) GLM-5 36% → 51% gpt-oss-120b: 25% reading the first 1,200 characters, 51% reading the 1,200 that best match (+26.2 points, 95% CI +13.1 to +41.0) gpt-oss-120b 25% → 51% GPT-6 Sol: 28% reading the first 1,200 characters, 49% reading the 1,200 that best match (+21.3 points, 95% CI +9.8 to +32.8) GPT-6 Sol 28% → 49% Jev: 31% reading the first 1,200 characters, 48% reading the 1,200 that best match (+16.4 points, 95% CI +6.6 to +27.9) Jev 31% → 48% Jev, one request per passage: 31% reading the first 1,200 characters, 46% reading the 1,200 that best match (+14.8 points, 95% CI +4.9 to +26.2) Jev, one request per passage 31% → 46% MiniMax M2.5: 34% reading the first 1,200 characters, 44% reading the 1,200 that best match (+9.8 points, 95% CI −6.6 to +26.2) MiniMax M2.5 34% → 44% Qwen3 235B A22B 2507: 33% reading the first 1,200 characters, 44% reading the 1,200 that best match (+11.5 points, 95% CI −1.6 to +24.6) Qwen3 235B A22B 2507 33% → 44% Mistral Large 3: 30% reading the first 1,200 characters, 44% reading the 1,200 that best match (+14.8 points, 95% CI +0.0 to +27.9) Mistral Large 3 30% → 44% Llama 4 Maverick: 36% reading the first 1,200 characters, 41% reading the 1,200 that best match (+4.9 points, 95% CI −6.6 to +16.4) Llama 4 Maverick 36% → 41% Claude Haiku 4.5: 33% reading the first 1,200 characters, 39% reading the 1,200 that best match (+6.6 points, 95% CI −3.3 to +18.0) Claude Haiku 4.5 33% → 39% Kev-4B (local): 31% reading the first 1,200 characters, 30% reading the 1,200 that best match (−1.6 points, 95% CI −13.1 to +8.2) Kev-4B 31% → 30% Qwen3-Reranker 0.6B (local): 21% reading the first 1,200 characters, 28% reading the 1,200 that best match (+6.6 points, 95% CI −3.3 to +16.4) Qwen3-Reranker 0.6B 21% → 28% Qwen3-Reranker 4B (local): 23% reading the first 1,200 characters, 23% reading the 1,200 that best match (+0.0 points, 95% CI −11.5 to +11.5) Qwen3-Reranker 4B 23% → 23% Cohere Rerank 3.5: 13% reading the first 1,200 characters, 23% reading the 1,200 that best match (+9.8 points, 95% CI +0.0 to +19.7) Cohere Rerank 3.5 13% → 23% Amazon Rerank 1.0: 16% reading the first 1,200 characters, 20% reading the 1,200 that best match (+3.3 points, 95% CI −4.9 to +13.1) Amazon Rerank 1.0 16% → 20% Laya (local): 23% reading the first 1,200 characters, 16% reading the 1,200 that best match (−6.6 points, 95% CI −14.8 to +1.6) Laya 23% → 16%
Figure 4. Hit@1 on BRIGHT StackOverflow with embeddings, 61 questions: each judge reading the first 1,200 characters of each passage (hollow) and the 1,200 that best match the question (filled). The dashed line is search's own order.

Reading each passage’s first 1,200 characters. With embeddings, search’s own order put a gold passage first for 36% of the 61 questions it found one for (Figure 4, Table 4). No ranker clearly beat that. The best LLM judge, GPT-6 Luna, was 6.6 points ahead (−6.6 to 19.7), and Jev 4.9 points behind (−21.3 to 11.5). Laya, Cohere Rerank 3.5 and Amazon Rerank 1.0 fell below it, though within the margin of error once adjusted for the 24 comparisons.

A judge that does not beat the first stage still changes its order, gaining some questions and losing others. GPT-6 Luna put a gold passage first for 11 questions where search had not, and moved one out of first place for 7; Jev did so for 12 and 15, Cohere Rerank 3.5 for 5 and 19.

Why search’s order held up. BRIGHT’s passages are long. The median is 4,000 characters, against 291 on MIRACL, and 99% of the gold passages run past the 1,200 characters a judge sees. Search scores each passage whole. Ordering the same 24 candidates by BM25 on whole passages put a gold passage first for 34% of questions, close to search’s own 36%; on their first 1,200 characters alone, for 13%. Most of the words that led BM25 to the gold passage lie beyond what a judge reads. Stuga’s own passages, whole sections of up to 12,000 characters, can be longer still.

Reading the 1,200 characters that best match. We designed this after the results above and measured it on the same questions, so it is exploratory until tested on others. We gave every judge, for each passage longer than 1,200 characters, the 1,200 that hold the most of the question’s words, the rarer ones counting for more, instead of the first 1,200, within the same 1,200-character limit. 21 of 24 judges rose, 7 of them clearly once adjusted for the 24 comparisons (Figure 4). Claude Opus 5.5 reached 61%, from 39%, and Jev 48%, from 31%; only Claude Opus 5.5 clearly beat search’s own order. Cohere Rerank 3.5 and Amazon Rerank 1.0 stayed below it, at 23% and 20%. Kev-4B, Qwen3-Reranker 4B and Laya, all on this machine, did not rise. Over all 117 questions, a retrieval miss counting as a miss, Claude Opus 5.5 put a gold passage first for 32% and Jev for 25%, against search’s 19%. With Jev a gold passage was among the eight the model reads for 93% of the 61 questions, against 72% without a judge. Stuga’s judges now read the best-matching 1,200 characters.

With and without embeddings. Table 4’s two Hit@1 columns count different questions: 36% of the 61 that search with embeddings found a gold passage for, and 40% of the 47 that BM25 alone found one for. With embeddings, search found a gold passage for every question BM25 alone did and for 14 more, but put it first for only 2 of those 14: they raise the count and lower the share. Over all 117 questions, search put a gold passage first for 19% with embeddings and 16% without. On BM25’s candidates the judges gained more. Reading the first 1,200 characters, Claude Opus 5.5 was 17.0 points ahead of search (−2.1 to 36.2); reading the best-matching 1,200, Claude Opus 5.5, Claude Sonnet 5.5 and MiniMax M2.5 clearly beat search’s 40%, at 72%, 70% and 68%.

Table 4. BRIGHT StackOverflow under Stuga's two first stages, each judge reading the first 1,200 characters of each passage and the 1,200 that best match the question. Hit@1 counts the questions whose candidates hold a gold passage, a different number under each, except in the second row, which counts all 117. The difference from the first-stage order, for the best-matching 1,200, is in points with its unadjusted 95% interval; bold where the interval adjusted for the 24 comparisons excludes zero.
Ranker BM25 + embeddings61 questionsBM25 only47 questions
First 1,200Best 1,200vs first stageFirst 1,200Best 1,200vs first stage
No reranker (first-stage order) 36%—40%—
No reranker, over all 117 questions 19%—16%—
Claude Opus 5.5 39% 61% +24.6 (+9.8 to +39.3) 57% 72% +31.9 (+14.9 to +48.9)
Claude Sonnet 5 34% 56% +19.7 (+4.9 to +34.5) 53% 66% +25.5 (+8.5 to +42.6)
Grok 4.6 31% 56% +19.7 (+1.6 to +36.1) 47% 66% +25.5 (+6.4 to +44.7)
Claude Sonnet 5.5 38% 54% +18.0 (+1.6 to +34.4) 51% 70% +29.8 (+12.8 to +46.8)
DeepSeek V3.2 25% 54% +18.0 (+3.3 to +32.8) 38% 60% +19.1 (+2.1 to +36.2)
GPT-6.1 Sol 33% 52% +16.4 (0.0 to +32.8) 51% 64% +23.4 (+6.4 to +42.6)
Amazon Nova 2 Lite 36% 52% +16.4 (+4.9 to +27.9) 47% 51% +10.6 (−2.1 to +23.4)
GPT-6 Luna 43% 51% +14.8 (0.0 to +29.5) 55% 64% +23.4 (+6.4 to +40.4)
GLM-5 36% 51% +14.8 (0.0 to +29.5) 49% 68% +27.7 (+8.5 to +46.8)
gpt-oss-120b 25% 51% +14.8 (−1.6 to +31.1) 40% 57% +17.0 (0.0 to +34.0)
GPT-6 Sol 28% 49% +13.1 (−1.6 to +27.9) 49% 64% +23.4 (+4.3 to +42.6)
Jev 31% 48% +11.5 (−4.9 to +27.9) 40% 60% +19.1 (0.0 to +38.3)
Jev, one request per passage 31% 46% +9.8 (−6.6 to +24.6) 40% 57% +17.0 (0.0 to +34.0)
MiniMax M2.5 34% 44% +8.2 (−8.2 to +26.2) 47% 68% +27.7 (+12.7 to +44.7)
Qwen3 235B A22B 2507 33% 44% +8.2 (−6.6 to +23.0) 51% 60% +19.1 (+2.1 to +38.3)
Mistral Large 3 30% 44% +8.2 (−4.9 to +21.3) 38% 43% +2.1 (−12.8 to +17.0)
Llama 4 Maverick 36% 41% +4.9 (−9.8 to +19.7) 47% 49% +8.5 (−8.5 to +27.7)
Claude Haiku 4.5 33% 39% +3.3 (−13.1 to +18.1) 40% 60% +19.1 (+2.1 to +36.2)
Kev-4B 31% 30% −6.6 (−19.7 to +8.2) 38% 47% +6.4 (−8.5 to +23.4)
Qwen3-Reranker 0.6B 21% 28% −8.2 (−23.0 to +6.6) 36% 45% +4.3 (−10.6 to +19.1)
Qwen3-Reranker 4B 23% 23% −13.1 (−29.5 to +4.9) 40% 47% +6.4 (−8.5 to +21.3)
Cohere Rerank 3.5 13% 23% −13.1 (−27.9 to +1.6) 36% 51% +10.6 (−4.3 to +27.7)
Amazon Rerank 1.0 16% 20% −16.4 (−31.1 to −1.6) 36% 45% +4.3 (−12.8 to +21.3)
Laya 23% 16% −19.7 (−31.1 to −8.2) 28% 23% −17.0 (−29.8 to −4.3)

4.4 How LLM judges fail

Stuga reads the scores wherever an answer states them unambiguously: as the JSON array it asks for, as the same objects one per line or followed by a note, or as [n] score lines. Four failure modes remained, all from producing the ranking as text. An unusable answer leaves the first-stage order.

Reasoning costs time. Grok 4.6 wrote 1,623 output tokens per search on average for a list of 24 scores, where GPT-6 Sol wrote 198, and took 12 s.

4.5 Cost

$0.10 $0.30 $1.0 $3.0 $10 $30 $100 $ per 1,000 searches, list price (log scale) Jev (TypeSafe API): $0.37 per 1,000 searches Jev $0.37 GPT-6 Luna (OpenAI API): $0.52 per 1,000 searches GPT-6 Luna $0.52 Jev, one request per passage (TypeSafe API): $0.64 per 1,000 searches Jev, one request per passage $0.64 Llama 4 Maverick (Bedrock us-west-2): $1.2 per 1,000 searches Llama 4 Maverick $1.2 gpt-oss-120b (Bedrock us-west-2): $1.2 per 1,000 searches gpt-oss-120b $1.2 Qwen3 235B A22B 2507 (Bedrock us-west-2): $1.2 per 1,000 searches Qwen3 235B A22B 2507 $1.2 Cohere Rerank 3.5 (Bedrock us-west-2): $2.0 per 1,000 searches Cohere Rerank 3.5 $2.0 Amazon Nova 2 Lite (Bedrock us-west-2): $2.2 per 1,000 searches Amazon Nova 2 Lite $2.2 DeepSeek V3.2 (Bedrock us-west-2): $3.2 per 1,000 searches DeepSeek V3.2 $3.2 Mistral Large 3 (Bedrock us-west-2): $3.8 per 1,000 searches Mistral Large 3 $3.8 MiniMax M2.5 (Bedrock us-west-2): $4.4 per 1,000 searches MiniMax M2.5 $4.4 GLM-5 (Bedrock us-west-2): $6.1 per 1,000 searches GLM-5 $6.1 Claude Haiku 4.5 (Bedrock us-west-2): $9.3 per 1,000 searches Claude Haiku 4.5 $9.3 GPT-6 Sol (OpenAI API): $11 per 1,000 searches GPT-6 Sol $11 GPT-6.1 Sol (Bedrock us-west-2): $14 per 1,000 searches GPT-6.1 Sol $14 Grok 4.6 (Bedrock us-west-2): $20 per 1,000 searches Grok 4.6 $20 Claude Sonnet 5 (Bedrock us-west-2): $20 per 1,000 searches Claude Sonnet 5 $20 Claude Sonnet 5.5 (Bedrock us-west-2): $20 per 1,000 searches Claude Sonnet 5.5 $20 Claude Opus 5.5 (Bedrock us-west-2): $37 per 1,000 searches Claude Opus 5.5 $37 $0.10 $1.0 $10 $100 $ per 1,000 searches, list price (log scale) Jev (TypeSafe API): $0.37 per 1,000 searches Jev $0.37 GPT-6 Luna (OpenAI API): $0.52 per 1,000 searches GPT-6 Luna $0.52 Jev, one request per passage (TypeSafe API): $0.64 per 1,000 searches Jev, one request per passage $0.64 Llama 4 Maverick (Bedrock us-west-2): $1.2 per 1,000 searches Llama 4 Maverick $1.2 gpt-oss-120b (Bedrock us-west-2): $1.2 per 1,000 searches gpt-oss-120b $1.2 Qwen3 235B A22B 2507 (Bedrock us-west-2): $1.2 per 1,000 searches Qwen3 235B A22B 2507 $1.2 Cohere Rerank 3.5 (Bedrock us-west-2): $2.0 per 1,000 searches Cohere Rerank 3.5 $2.0 Amazon Nova 2 Lite (Bedrock us-west-2): $2.2 per 1,000 searches Amazon Nova 2 Lite $2.2 DeepSeek V3.2 (Bedrock us-west-2): $3.2 per 1,000 searches DeepSeek V3.2 $3.2 Mistral Large 3 (Bedrock us-west-2): $3.8 per 1,000 searches Mistral Large 3 $3.8 MiniMax M2.5 (Bedrock us-west-2): $4.4 per 1,000 searches MiniMax M2.5 $4.4 GLM-5 (Bedrock us-west-2): $6.1 per 1,000 searches GLM-5 $6.1 Claude Haiku 4.5 (Bedrock us-west-2): $9.3 per 1,000 searches Claude Haiku 4.5 $9.3 GPT-6 Sol (OpenAI API): $11 per 1,000 searches GPT-6 Sol $11 GPT-6.1 Sol (Bedrock us-west-2): $14 per 1,000 searches GPT-6.1 Sol $14 Grok 4.6 (Bedrock us-west-2): $20 per 1,000 searches Grok 4.6 $20 Claude Sonnet 5 (Bedrock us-west-2): $20 per 1,000 searches Claude Sonnet 5 $20 Claude Sonnet 5.5 (Bedrock us-west-2): $20 per 1,000 searches Claude Sonnet 5.5 $20 Claude Opus 5.5 (Bedrock us-west-2): $37 per 1,000 searches Claude Opus 5.5 $37
Figure 5. List price per 1,000 MIRACL searches. Local rankers have no per-search price.

Jev counted 8,801 input tokens per MIRACL search and costs input only. Claude Opus 5.5 read 7,093 and wrote 273. Cohere Rerank 3.5 is billed per query, whatever its length.

4.6 Local judges

Qwen3-Reranker 4B reached 81% on MIRACL, level with the hosted leaders, in 5.3 s: a cross-encoder reads the question with each of the 24 passages, and Thai, Japanese and Korean text is long in tokens. The 0.6B model reached 75% in 987 ms.

Two open projects serve the System One API: Kev [11] and Laya [12]. Kev-4B reached 64% in 5.4 s. Laya, whose English checkpoint is ModernBERT-large and whose multilingual one is mmBERT-base, reached 49% in 646 ms; its README places further accuracy in fine-tuning.

40%45%50%55%60%65%70%75%80%85%90%↑ Right passage first100 ms300 ms1.0 s3.0 s10 s30 sMedian time per search (log scale) →Qwen3-Reranker 0.6B (local) Reranker, llama.cpp on this machine Hit@1 75% (95% CI 69%–80%) Median 987 ms · $0.00 per 1,000 searchesQwen3-Reranker 4B (local) Reranker, llama.cpp on this machine Hit@1 81% (95% CI 75%–86%) Median 5.3 s · $0.00 per 1,000 searchesLaya (local) System One classifier, this machine Hit@1 49% (95% CI 43%–55%) Median 646 ms · $0.00 per 1,000 searchesKev-4B (local) System One classifier, this machine Hit@1 64% (95% CI 58%–70%) Median 5.4 s · $0.00 per 1,000 searchesJev System One classifier, TypeSafe API Hit@1 77% (95% CI 72%–82%) Median 131 ms · $0.37 per 1,000 searchesJev (hosted, for reference)Qwen3-Reranker 4BKev-4BQwen3-Reranker 0.6BLaya
40%50%60%70%80%90%↑ Right passage first100 ms1.0 s10 sMedian time per search (log scale) →Qwen3-Reranker 0.6B (local) Reranker, llama.cpp on this machine Hit@1 75% (95% CI 69%–80%) Median 987 ms · $0.00 per 1,000 searchesQwen3-Reranker 4B (local) Reranker, llama.cpp on this machine Hit@1 81% (95% CI 75%–86%) Median 5.3 s · $0.00 per 1,000 searchesLaya (local) System One classifier, this machine Hit@1 49% (95% CI 43%–55%) Median 646 ms · $0.00 per 1,000 searchesKev-4B (local) System One classifier, this machine Hit@1 64% (95% CI 58%–70%) Median 5.4 s · $0.00 per 1,000 searchesJev System One classifier, TypeSafe API Hit@1 77% (95% CI 72%–82%) Median 131 ms · $0.37 per 1,000 searchesJev (hosted)Qwen3 4BKevQwen3 0.6BLaya
Figure 6. Local rankers on MIRACL, with Jev for reference, on Figure 1's time axis.

5. Choosing a first stage and a judge

In this study’s setup, the results suggest the following defaults. None was tested on a team’s own documents, with other chunking or with real users’ questions, and each holds only as far as two benchmarks of this size allow (Section 7).

Table 5. Choosing a first stage and a judge, from the results above.

Situation Suggestion Evidence
By default Add embeddings to BM25 On BRIGHT, search with embeddings found a gold passage for 61 rather than 47 of 117 questions, losing none that BM25 alone found; a judge cannot rank what search missed.
Questions that ask for facts Rank with Jev On MIRACL, 77% in 131 ms at $0.37 per 1,000 searches; no ranker was clearly better.
A few points more, whatever the wait and cost An LLM judge such as Claude Opus 5.5 4.2 points more than Jev on MIRACL, not shown to be better, at 27 times the latency and 100 times the cost; some LLM judges return unusable answers (Section 4.4).
Text that must stay on the machine Qwen3-Reranker 4B, locally 81% on MIRACL, but 5.3 s per search on an Apple M6; the 0.6B model 75% in 987 ms. Kev and Laya reached 64% and 49%.
Long passages Show the judge the part that matches the question On BRIGHT, 21 of 24 judges rose reading the best-matching 1,200 characters instead of the first 1,200, Jev from 31% to 48%; exploratory, measured on the questions it was designed on.
Questions that need reasoning An LLM judge, and better retrieval On BRIGHT, reading the best-matching 1,200 characters, Claude Opus 5.5 reached 61% against search’s 36% and Jev 48%; search still missed a gold passage for 48% of questions.

For Stuga’s defaults this means Semantic search on and Reranking with Jev, reading the part of each passage that best matches the question. A team whose questions are mostly troubleshooting may gain more from an LLM judge, at its latency and cost; rewriting the question before retrieving helped on BRIGHT in its authors’ tests [15], and we did not test it in Stuga.

6. In Stuga

Reranking is set up under Settings → This node → AI providers, beside Built-in AI and Semantic search, with a TypeSafe key, directly or through OpenRouter. Every search from Ask and from agents then goes through it. Without it, the chat model judges while Built-in AI is on; with neither, fusion order stands.

Stuga's AI providers settings: Built-in AI with GPT-6 Sol, Semantic search with text-embedding-3-large, and Reranking with TypeSafe jev-latest.
Figure 7. Reranking in Stuga’s settings.
An Ask answer in Stuga comparing breach-notification deadlines in six privacy laws, each point cited to the law's own text.
Figure 8. An Ask answer across six laws in six languages, from passages Jev ranked.

A hosted judge sees what it ranks: the question and up to 24 passages, 1,200 characters of each, go to TypeSafe or OpenRouter, as they go to the chat provider when it judges. Reranking accepts any System One endpoint, so a server such as Kev on the same machine keeps everything local, at 64% on MIRACL and 5.4 s per search.

Ranking decides what an agent reads. What it writes back to a document or table waits, by default, for a person to accept or reject: how that review works.

7. Limitations

8. Reproducibility

Code, the two test sets, every candidate list and every call’s result are at stuga-dev/rerank-bench (code MIT; data under each source’s licence). The test sets are built from pinned revisions of the public datasets.

Terminal window
git clone https://github.com/stuga-dev/rerank-bench && cd rerank-bench && pnpm install
BENCH_KEYS=keys.json node src/run.ts --set miracl --only jev,gpt-6-sol --limit 5
node src/report.ts results/miracl/<date>

A ranker maps a query and 24 passages to 24 scores; pull requests with new rankers are welcome.

Stuga has no business relationship with TypeSafe or any other vendor named here; all calls were paid at list price.

References

  1. D. Kahneman. Thinking, Fast and Slow. Farrar, Straus and Giroux, 2011.
  2. J. Devlin, M.-W. Chang, K. Lee, K. Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL, 2019.
  3. R. Nogueira, K. Cho. Passage Re-ranking with BERT. arXiv:1901.04085, 2019.
  4. W. Sun et al. Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents. EMNLP, 2023.
  5. S. Robertson, H. Zaragoza. The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval 3(4), 2009.
  6. G. V. Cormack, C. L. A. Clarke, S. Büttcher. Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods. SIGIR, 2009.
  7. B. Warner et al. Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference. arXiv:2412.13663, 2024.
  8. TypeSafe. Introducing System One models and Jev. 15 September 2026.
  9. B. Efron, R. J. Tibshirani. An Introduction to the Bootstrap. Chapman & Hall, 1993.
  10. Amazon Web Services. Amazon Bedrock pricing, retrieved 28 September 2026.
  11. J. Palmer. Kev and kev-4b, 2026.
  12. Convai Innovations. Laya and Laya multilingual (models); laya 0.3.20 (code), 2026.
  13. X. Zhang, N. Thakur, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, X. Li, Q. Liu, M. Rezagholizadeh, J. Lin. MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages. TACL 11, 2023.
  14. K. Enevoldsen et al. MMTEB: Massive Multilingual Text Embedding Benchmark. arXiv:2502.13595, 2025.
  15. H. Su et al. BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval. arXiv:2407.12883, 2024.