Reranking with a classifier instead of a chat model
TL;DR
A search in Stuga returns 24 candidate passages, and a judge reorders them so that the model reads the best eight. We compared 23 judges on two public benchmarks whose questions and answers were marked by people. Every judge reordered the same 24 candidates and saw each one’s title, heading path and 1,200 characters of its text. The main measure is Hit@1: how often the first passage is one that answers the question.
Questions that ask for facts. MIRACL has questions over Wikipedia in six languages; we took 240 and reordered the 24 candidates MMTEB publishes for each. Without a judge, a relevant passage came first for 36% of questions. TypeSafe’s Jev, a System One classifier, raised that to 77%, in 131 ms for $0.37 per 1,000 searches, and put a relevant passage among the eight the model reads for 99% of questions, against 80%. The best LLM judge, Claude Opus 5.5, reached 81%. The 4.2 points between them are within this test’s margin of error (−2.9 to 10.8 points, adjusted for the 23 comparisons with Jev), so the data say neither that one is better nor that they are equal. Claude Opus 5.5 took 27 times as long and cost 100 times as much. Among the newest models tested, Claude Sonnet 5.5 reached 80% in 2.0 s, at $20 per 1,000 searches; GPT-6.1 Sol reached 77% in 2.2 s, at $14.
Questions that need reasoning. BRIGHT’s 117 StackOverflow questions are answered by documentation pages that may share few words with the question. Search found such a page among the 24 candidates for only 52% of questions, and its own order put one first for 19% of all 117. The passages are long, typically 4,000 characters, and of the judges reading the first 1,200 of each, as Stuga’s did until this study, none clearly beat search’s order. Reading instead the 1,200 characters that best match the question, 21 of 24 judges rose, 7 of them clearly: over all 117 questions, Claude Opus 5.5 then put a gold passage first for 32% and Jev for 25%. We designed that reading after seeing the first results and measured it on the same questions, so it is exploratory; Stuga’s judges now read that part (Section 4.3).
In this setup, Jev was a fast, cheap judge for questions that ask for facts, and no judge we tested clearly beat it. On long passages, what a judge reads of each mattered as much as which judge it is.
We ran this test to decide what Stuga’s Ask and its tools for agents show an AI before it answers: which judge orders the candidates, and which part of each passage it reads. Section 6 shows the result in a workspace.
1. Introduction
Stuga answers questions from a workspace’s documents (Ask) and serves the same retrieval to agents over MCP. Retrieval returns 24 candidate passages; a judge reorders them, and the model reads the first eight. The judge decides what the model sees, so we measured the options on public benchmarks, in the shape Stuga runs them. The primary comparison is Jev, which Stuga’s Reranking setting uses, against the LLM judges Stuga falls back to. The contribution is an empirical comparison through the interfaces a product calls, not a new ranker:
- quality, latency and cost for 23 rerankers of three kinds, hosted and local, on the same candidates;
- a fact-seeking benchmark in six languages and a reasoning-heavy one, the second under both of Stuga’s first stages;
- on long passages, what the judge should read: each passage’s first 1,200 characters or the 1,200 that best match the question;
- every call’s result and the code to reproduce or extend it.
Five questions guided the study:
- How often does each judge put a relevant passage first, and among the eight?
- At what latency and cost?
- How do LLM judges fail?
- Can the judge run on the machine that runs Stuga?
- What should a judge read of a long passage?
2. Background
2.1 Retrieval in Stuga
Stuga splits each document at its headings: every section is a passage, split further only past 12,000 characters, and a document without headings is cut into pieces of about 6,000. Each passage gets one vector, from its heading path and text. A question runs two searches in one SQL statement over these passages: BM25 [5] through pg_search and vector similarity through pgvector, 96 passages each. Reciprocal rank fusion [6] (k = 60) merges the lists and keeps 24 (Figure 2).
Three terms recur below. The first stage is the search that picks the 24 candidates; its order, with no judge, is the baseline every judge is measured against. A relevant passage is one the benchmark’s annotators marked as answering the question; BRIGHT calls it gold. Hit@1 is the share of questions whose first passage is relevant.
A first stage finds the neighbourhood, not the house: a passage about the question’s subject can outrank the one that answers it. On MIRACL’s candidate lists a relevant passage came first for 36% of questions and was among the first eight for 80%. Figure 3 follows three of them.
“When was the town of Simcoe, Ontario founded?”
No judge: first-stage order Answer #22 of 24
- 1 Barrie (electoral district) Prior to the 2015 election Barrie was a federal electoral district in Ontario, Canada, that has been represented in the House of Commons of Canada since
- 2 Simcoe, Ontario A cultural club for people of Croatian descent operates in this town; the formal name given to this organization is the 531st branch of the Croatian Fraternal Union
- 3 Simcoe, Ontario Simcoe is an unincorporated community and former town in Southwestern Ontario, Canada near Lake Erie. It is the county seat and largest community of Norfolk County.
- 4 Elizabeth Simcoe Elizabeth Simcoe left a diary that provides a valuable impression of life in colonial Ontario. First published in 1934, there was a subsequent transcription publis
- 5 History of Markham, Ontario When Upper and Lower Canada were established in 1791, Colonel John Graves Simcoe was appointed the first Lieutenant-Governor of Upper Canada. Simcoe nam
- 6 Simcoe, Ontario Simcoe was incorporated as a town in 1878 and had its own town council and mayor until December 31, 2000. In 2001, the town and all other municipalities within the
- 7 Simcoe, Ontario Rural Canadian towns similar to Simcoe are close to dying due to economic and transportation issues that prevent people from holding meaningful employment and being
- 8 177th Battalion (Simcoe Foresters), CEF The 177th (Simcoe Foresters) Battalion, CEF was a unit in the Canadian Expeditionary Force raised during the First World War by the 35th Sim
The answer is #22: the model never reads it.
Jev Answer #1 of 24 · 105 ms
- 1 Simcoe, Ontario Simcoe was founded in 1795 by Lieutenant Governor John Graves Simcoe. Initially, the settlement consisted of two distinct areas, Birdtown, named by William Bird who Answer0.98
- 2 Communities in Norfolk County, Ontario The population in 1850 was about 1600; in that year, Simcoe became the County seat of Norfolk County. Simcoe was incorporated as a town in 18 0.45
- 3 Simcoe, Ontario Simcoe was incorporated as a town in 1878 and had its own town council and mayor until December 31, 2000. In 2001, the town and all other municipalities within the 0.42
- 4 History of Markham, Ontario When Upper and Lower Canada were established in 1791, Colonel John Graves Simcoe was appointed the first Lieutenant-Governor of Upper Canada. Simcoe nam 0.07
- 5 Simcoe, Ontario Simcoe is an unincorporated community and former town in Southwestern Ontario, Canada near Lake Erie. It is the county seat and largest community of Norfolk County. 0.06
- 6 Elizabeth Simcoe Elizabeth Simcoe left a diary that provides a valuable impression of life in colonial Ontario. First published in 1934, there was a subsequent transcription publis 0.05
- 7 Communities in Norfolk County, Ontario Simcoe is the administrative centre of Norfolk County, with a population of 16,000 making it Norfolk's largest community. Simcoe is located a 0.05
- 8 Henry Dundas, 1st Viscount Melville He was friends with John Graves Simcoe, Lieutenant Governor of Upper Canada. Simcoe named the town of Dundas, Ontario, in southern Ontario after 0.05
Claude Opus 5.5 Answer #1 of 24 · 3.2 s
- 1 Simcoe, Ontario Simcoe was founded in 1795 by Lieutenant Governor John Graves Simcoe. Initially, the settlement consisted of two distinct areas, Birdtown, named by William Bird who Answer10/10
- 2 Simcoe, Ontario Simcoe was incorporated as a town in 1878 and had its own town council and mayor until December 31, 2000. In 2001, the town and all other municipalities within the 6/10
- 3 Communities in Norfolk County, Ontario The population in 1850 was about 1600; in that year, Simcoe became the County seat of Norfolk County. Simcoe was incorporated as a town in 18 6/10
- 4 Simcoe, Ontario Simcoe is an unincorporated community and former town in Southwestern Ontario, Canada near Lake Erie. It is the county seat and largest community of Norfolk County. 2/10
- 5 Simcoe, Ontario A cultural club for people of Croatian descent operates in this town; the formal name given to this organization is the 531st branch of the Croatian Fraternal Union 1/10
- 6 Simcoe, Ontario Rural Canadian towns similar to Simcoe are close to dying due to economic and transportation issues that prevent people from holding meaningful employment and being 1/10
- 7 Simcoe, Ontario Norfolk Street tends to be the dividing road between the "haves" and the "have nots" in this community. Most houses west of Norfolk Street are aging dilapidated bui 1/10
- 8 Communities in Norfolk County, Ontario Simcoe is the administrative centre of Norfolk County, with a population of 16,000 making it Norfolk's largest community. Simcoe is located a 1/10
GPT-6 Sol Answer #1 of 24 · 2.2 s
- 1 Simcoe, Ontario Simcoe was founded in 1795 by Lieutenant Governor John Graves Simcoe. Initially, the settlement consisted of two distinct areas, Birdtown, named by William Bird who Answer10/10
- 2 Simcoe, Ontario Simcoe was incorporated as a town in 1878 and had its own town council and mayor until December 31, 2000. In 2001, the town and all other municipalities within the 5/10
- 3 Communities in Norfolk County, Ontario The population in 1850 was about 1600; in that year, Simcoe became the County seat of Norfolk County. Simcoe was incorporated as a town in 18 5/10
- 4 Simcoe, Ontario Simcoe is an unincorporated community and former town in Southwestern Ontario, Canada near Lake Erie. It is the county seat and largest community of Norfolk County. 3/10
- 5 Simcoe, Ontario A cultural club for people of Croatian descent operates in this town; the formal name given to this organization is the 531st branch of the Croatian Fraternal Union 1/10
- 6 Communities in Norfolk County, Ontario Simcoe is the administrative centre of Norfolk County, with a population of 16,000 making it Norfolk's largest community. Simcoe is located a 1/10
- 7 Barrie (electoral district) Prior to the 2015 election Barrie was a federal electoral district in Ontario, Canada, that has been represented in the House of Commons of Canada since 0/10
- 8 Elizabeth Simcoe Elizabeth Simcoe left a diary that provides a valuable impression of life in colonial Ontario. First published in 1934, there was a subsequent transcription publis 0/10
Cohere Rerank 3.5 Answer #1 of 24 · 214 ms
- 1 Simcoe, Ontario Simcoe was founded in 1795 by Lieutenant Governor John Graves Simcoe. Initially, the settlement consisted of two distinct areas, Birdtown, named by William Bird who Answer0.92
- 2 Simcoe, Ontario Simcoe was incorporated as a town in 1878 and had its own town council and mayor until December 31, 2000. In 2001, the town and all other municipalities within the 0.90
- 3 Communities in Norfolk County, Ontario The population in 1850 was about 1600; in that year, Simcoe became the County seat of Norfolk County. Simcoe was incorporated as a town in 18 0.81
- 4 History of Markham, Ontario When Upper and Lower Canada were established in 1791, Colonel John Graves Simcoe was appointed the first Lieutenant-Governor of Upper Canada. Simcoe nam 0.42
- 5 Simcoe, Ontario Simcoe is an unincorporated community and former town in Southwestern Ontario, Canada near Lake Erie. It is the county seat and largest community of Norfolk County. 0.32
- 6 Lynn River Music and Arts Festival The Lynn River Music and Arts Festival (formerly the Simcoe Rotary Friendship Festival and the Simcoe Friendship Festival) is the community's old 0.31
- 7 Henry Dundas, 1st Viscount Melville He was friends with John Graves Simcoe, Lieutenant Governor of Upper Canada. Simcoe named the town of Dundas, Ontario, in southern Ontario after 0.30
- 8 Elizabeth Simcoe Elizabeth Simcoe left a diary that provides a valuable impression of life in colonial Ontario. First published in 1934, there was a subsequent transcription publis 0.29
“Are verbal contracts binding?”
No judge: first-stage order Answer #18 of 24
- 1 Oculus Sacerdotis 5. On marriage engagements and secret marriages. The verbal consent between a man and a woman to marry, even if not betrothed, is enough to make their marriage bi
- 2 Student rights in higher education "Carr v. St. Johns University" (1962) and "Healey v. Larsson" (1971, 1974) established that students and institutions of higher education formed
- 3 Stipulatio Firstly, if you stipulate for slave A and there are two slaves called A, which slave the stipulation is binding for depends on evidence extraneous to the verbal contract
- 4 Muamalat At least one source (a scholar identified as "Barbarti") defines "aqad" (contract) as a “legal relationship created by the conjunction of two declarations, from which flow
- 5 The Puffy Shirt A blog dedicated to the legality of the issues that arise in Seinfeld episodes, Seinfeld Law, discusses whether or not Jerry’s verbal and physical affirmations crea
- 6 NBA salary cap During the moratorium, teams are restricted from commenting on deals. Teams and players can reach verbal agreements, but they are not binding. Contracts can be signe
- 7 Marriage in Islam In Islam, marriage is a legal contract between a man and a woman. Both the groom and the bride are to consent to the marriage of their own free wills. A formal, b
- 8 Student rights in higher education Decision making should not be arbitrary or capricious / random and, thus, interfere with fairness. While this case concerned a private school, "H
The answer is #18: the model never reads it.
Jev Answer #2 of 24 · 116 ms
- 1 Student rights in higher education "Carr v. St. Johns University" (1962) and "Healey v. Larsson" (1971, 1974) established that students and institutions of higher education formed 0.93
- 2 Student rights in higher education Verbal contracts are binding. They must be made in an official capacity, however, to be binding. "Dezick v. Umpqua Community College" (1979) foun Answer0.93
- 3 Marriage in Islam In Islam, marriage is a legal contract between a man and a woman. Both the groom and the bride are to consent to the marriage of their own free wills. A formal, b 0.89
- 4 Oculus Sacerdotis 5. On marriage engagements and secret marriages. The verbal consent between a man and a woman to marry, even if not betrothed, is enough to make their marriage bi 0.87
- 5 Muamalat At least one source (a scholar identified as "Barbarti") defines "aqad" (contract) as a “legal relationship created by the conjunction of two declarations, from which flow 0.79
- 6 NBA salary cap During the moratorium, teams are restricted from commenting on deals. Teams and players can reach verbal agreements, but they are not binding. Contracts can be signe 0.79
- 7 Languages of the Roman Empire Roman law was written in Latin, and the "letter of the law" was tied strictly to the words in which it was expressed. Any language, however, could be 0.77
- 8 Islamic marital practices The marriage contract that binds the marital union is called the "Akad Nikah", a verbal agreement sealed by a financial sum known as the "mas kahwin", and 0.76
Claude Opus 5.5 Answer #1 of 24 · 3.3 s
- 1 Student rights in higher education Verbal contracts are binding. They must be made in an official capacity, however, to be binding. "Dezick v. Umpqua Community College" (1979) foun Answer7/10
- 2 Student rights in higher education "Carr v. St. Johns University" (1962) and "Healey v. Larsson" (1971, 1974) established that students and institutions of higher education formed 6/10
- 3 Muamalat At least one source (a scholar identified as "Barbarti") defines "aqad" (contract) as a “legal relationship created by the conjunction of two declarations, from which flow 5/10
- 4 Oculus Sacerdotis 5. On marriage engagements and secret marriages. The verbal consent between a man and a woman to marry, even if not betrothed, is enough to make their marriage bi 4/10
- 5 NBA salary cap During the moratorium, teams are restricted from commenting on deals. Teams and players can reach verbal agreements, but they are not binding. Contracts can be signe 4/10
- 6 Languages of the Roman Empire Roman law was written in Latin, and the "letter of the law" was tied strictly to the words in which it was expressed. Any language, however, could be 4/10
- 7 Robert Sharpe (railway contractor) In 1871 Paul Wallace Sharpe and William John Sharpe attempted to sue San Paulo Railway to recover additional costs of £617,143 which they had inc 4/10
- 8 Sweet Grass (Cree chief) The legacy of Treaty 6 continues to affect the Cree till the modern day. Issues arise from the mixed interpretations of the Treaty by both the Indigenous g 4/10
GPT-6 Sol Answer #1 of 24 · 2.5 s
- 1 Student rights in higher education Verbal contracts are binding. They must be made in an official capacity, however, to be binding. "Dezick v. Umpqua Community College" (1979) foun Answer9/10
- 2 Student rights in higher education "Carr v. St. Johns University" (1962) and "Healey v. Larsson" (1971, 1974) established that students and institutions of higher education formed 8/10
- 3 Muamalat At least one source (a scholar identified as "Barbarti") defines "aqad" (contract) as a “legal relationship created by the conjunction of two declarations, from which flow 6/10
- 4 NBA salary cap During the moratorium, teams are restricted from commenting on deals. Teams and players can reach verbal agreements, but they are not binding. Contracts can be signe 6/10
- 5 Oculus Sacerdotis 5. On marriage engagements and secret marriages. The verbal consent between a man and a woman to marry, even if not betrothed, is enough to make their marriage bi 5/10
- 6 Marriage in Islam In Islam, marriage is a legal contract between a man and a woman. Both the groom and the bride are to consent to the marriage of their own free wills. A formal, b 5/10
- 7 Languages of the Roman Empire Roman law was written in Latin, and the "letter of the law" was tied strictly to the words in which it was expressed. Any language, however, could be 5/10
- 8 Robert Sharpe (railway contractor) In 1871 Paul Wallace Sharpe and William John Sharpe attempted to sue San Paulo Railway to recover additional costs of £617,143 which they had inc 5/10
Cohere Rerank 3.5 Answer #1 of 24 · 143 ms
- 1 Student rights in higher education Verbal contracts are binding. They must be made in an official capacity, however, to be binding. "Dezick v. Umpqua Community College" (1979) foun Answer0.90
- 2 Student rights in higher education "Carr v. St. Johns University" (1962) and "Healey v. Larsson" (1971, 1974) established that students and institutions of higher education formed 0.84
- 3 Oculus Sacerdotis 5. On marriage engagements and secret marriages. The verbal consent between a man and a woman to marry, even if not betrothed, is enough to make their marriage bi 0.72
- 4 Stipulatio Firstly, if you stipulate for slave A and there are two slaves called A, which slave the stipulation is binding for depends on evidence extraneous to the verbal contract 0.70
- 5 Languages of the Roman Empire Roman law was written in Latin, and the "letter of the law" was tied strictly to the words in which it was expressed. Any language, however, could be 0.64
- 6 NBA salary cap During the moratorium, teams are restricted from commenting on deals. Teams and players can reach verbal agreements, but they are not binding. Contracts can be signe 0.62
- 7 The Puffy Shirt A blog dedicated to the legality of the issues that arise in Seinfeld episodes, Seinfeld Law, discusses whether or not Jerry’s verbal and physical affirmations crea 0.62
- 8 Robert Sharpe (railway contractor) In 1871 Paul Wallace Sharpe and William John Sharpe attempted to sue San Paulo Railway to recover additional costs of £617,143 which they had inc 0.58
“What kind of government does Laos have?”
No judge: first-stage order Answer #19 of 24
- 1 LGBT rights in Laos Laos does not recognize same-sex marriages, nor any other form of same-sex union. There have been no known debates of such unions being legalized in the near fu
- 2 LGBT rights in Laos While homosexuality is legal in Laos, it is very difficult to assess the current state of acceptance and violence that lesbian, gay, bisexual, and transgender (
- 3 Isan language The Lao (Isan) language in Thailand is classified by "Ethnologue" as a ""de facto" language of provincial identity" which is defined as a language that "is the langua
- 4 Iu Mien people Since Iu Mien people had settled into Laos and Thailand, they’ve gained more independence. One group of Iu Mien migrated from Vietnam to Thailand. The other migrated
- 5 Government policies and the subprime mortgage crisis Economist Paul Krugman described the run on the shadow banking system as the "core of what happened" to cause the crisis. "As t
- 6 Subprime mortgage crisis Nobel laureate economist Paul Krugman described the run on the shadow banking system as the "core of what happened" to cause the crisis. As the shadow bank
- 7 Shadow banking system Economist Paul Krugman described the run on the shadow banking system as the "core of what happened" to cause the crisis. "As the shadow banking system expand
- 8 Munich (film) David Edelstein of "Slate" argued that "The Israeli government and many conservative and pro-Israeli commentators have lambasted the film for naiveté, for implying th
The answer is #19: the model never reads it.
Jev Answer #6 of 24 · 104 ms
- 1 Flag of Laos The flag of Laos consists of three horizontal stripes, with the middle stripe in blue being twice the height of the top and bottom red stripes. In the middle is a whit 0.92
- 2 Human rights in Laos However, according to Amnesty International, Human Rights Watch, the Center for Public Policy Analysis, the United League for Democracy in Laos, the Lao Human 0.85
- 3 LGBT rights in Laos While homosexuality is legal in Laos, it is very difficult to assess the current state of acceptance and violence that lesbian, gay, bisexual, and transgender ( 0.77
- 4 Economy of Laos In an attempt to stimulate further international commerce, the PDR Lao government accepted Australian aid to build a bridge across the Mekong River to Thailand. The 0.46
- 5 1962 in the Vietnam War The Geneva Agreement on Laos was signed by 14 countries, including China, the Soviet Union, and the United States. The agreement declared a cease fire betwe 0.41
- 6 Royal Lao Government in Exile The Royal Lao Government in Exile claims that it is an interim democratic government consisting of eighty representatives from Lao political organizat Answer0.39
- 7 History of Laos since 1945 In 1975, the Pathēt Lao forces on the Plain of Jars supported by North Vietnamese heavy artillery and other units began advancing westward. In late April 0.34
- 8 Lao Evangelical Church Decree 92 is the Prime Minister's Decree that allows the Church to exist in Laos. It states how a religion can register with the Government and what the reli 0.14
Claude Opus 5.5 Answer #9 of 24 · 3.2 s
- 1 Flag of Laos The flag of Laos consists of three horizontal stripes, with the middle stripe in blue being twice the height of the top and bottom red stripes. In the middle is a whit 7/10
- 2 Human rights in Laos However, according to Amnesty International, Human Rights Watch, the Center for Public Policy Analysis, the United League for Democracy in Laos, the Lao Human 6/10
- 3 LGBT rights in Laos While homosexuality is legal in Laos, it is very difficult to assess the current state of acceptance and violence that lesbian, gay, bisexual, and transgender ( 4/10
- 4 1962 in the Vietnam War The Geneva Agreement on Laos was signed by 14 countries, including China, the Soviet Union, and the United States. The agreement declared a cease fire betwe 3/10
- 5 History of Laos since 1945 In 1975, the Pathēt Lao forces on the Plain of Jars supported by North Vietnamese heavy artillery and other units began advancing westward. In late April 3/10
- 6 Lao Evangelical Church Decree 92 is the Prime Minister's Decree that allows the Church to exist in Laos. It states how a religion can register with the Government and what the reli 2/10
- 7 Economy of Laos In an attempt to stimulate further international commerce, the PDR Lao government accepted Australian aid to build a bridge across the Mekong River to Thailand. The 2/10
- 8 Christianity in Laos According to the US government and other agencies there have been instances of the Laotian government attempting to make Christians renounce their faith, and h 2/10
The answer is #9: the model never reads it.
GPT-6 Sol Answer #13 of 24 · 2.5 s
- 1 Flag of Laos The flag of Laos consists of three horizontal stripes, with the middle stripe in blue being twice the height of the top and bottom red stripes. In the middle is a whit 8/10
- 2 LGBT rights in Laos While homosexuality is legal in Laos, it is very difficult to assess the current state of acceptance and violence that lesbian, gay, bisexual, and transgender ( 6/10
- 3 Human rights in Laos However, according to Amnesty International, Human Rights Watch, the Center for Public Policy Analysis, the United League for Democracy in Laos, the Lao Human 6/10
- 4 1962 in the Vietnam War The Geneva Agreement on Laos was signed by 14 countries, including China, the Soviet Union, and the United States. The agreement declared a cease fire betwe 3/10
- 5 Economy of Laos In an attempt to stimulate further international commerce, the PDR Lao government accepted Australian aid to build a bridge across the Mekong River to Thailand. The 2/10
- 6 History of Laos since 1945 In 1975, the Pathēt Lao forces on the Plain of Jars supported by North Vietnamese heavy artillery and other units began advancing westward. In late April 2/10
- 7 Laotian Civil War In late April, the Pathēt Lao took the government outpost at Sala Phou Khoum crossroads which opened up Route 13 to a Pathēt Lao advance toward Muang Kassy. For t 2/10
- 8 LGBT rights in Laos Laos does not recognize same-sex marriages, nor any other form of same-sex union. There have been no known debates of such unions being legalized in the near fu 1/10
The answer is #13: the model never reads it.
Cohere Rerank 3.5 Answer #2 of 24 · 156 ms
- 1 Human rights in Laos However, according to Amnesty International, Human Rights Watch, the Center for Public Policy Analysis, the United League for Democracy in Laos, the Lao Human 0.73
- 2 Royal Lao Government in Exile The Royal Lao Government in Exile claims that it is an interim democratic government consisting of eighty representatives from Lao political organizat Answer0.68
- 3 Flag of Laos The flag of Laos consists of three horizontal stripes, with the middle stripe in blue being twice the height of the top and bottom red stripes. In the middle is a whit 0.60
- 4 LGBT rights in Laos While homosexuality is legal in Laos, it is very difficult to assess the current state of acceptance and violence that lesbian, gay, bisexual, and transgender ( 0.55
- 5 1962 in the Vietnam War The Geneva Agreement on Laos was signed by 14 countries, including China, the Soviet Union, and the United States. The agreement declared a cease fire betwe 0.53
- 6 Economy of Laos In an attempt to stimulate further international commerce, the PDR Lao government accepted Australian aid to build a bridge across the Mekong River to Thailand. The 0.44
- 7 Christianity in Laos According to the US government and other agencies there have been instances of the Laotian government attempting to make Christians renounce their faith, and h 0.42
- 8 History of Laos since 1945 In 1975, the Pathēt Lao forces on the Plain of Jars supported by North Vietnamese heavy artillery and other units began advancing westward. In late April 0.35
2.2 Three kinds of judge
Table 1. The judges compared.
| LLM judge | Reranker | System One classifier | |
|---|---|---|---|
| Tested | GPT-6, Claude, Grok, Qwen, Llama and others | Cohere Rerank, Amazon Rerank, Qwen3-Reranker | Jev, Kev, Laya |
| Reads | all 24 passages in one prompt | the question with one passage at a time | the passages as a state, one question each |
| Returns | text: a JSON list of 0–10 scores | a relevance score | a probability |
| Built on | a generative LLM | classically a BERT-style encoder [2, 3]; Qwen3-Reranker on the Qwen3 LLM | Jev: not published. Kev: Qwen3.5 with a LoRA and a decision head. Laya: ModernBERT [7] and mmBERT |
An LLM judge scores passages by generating text, the approach RankGPT [4] popularised. A reranker is a cross-encoder in the tradition of monoBERT [3]: the question and one passage pass through the model together and a head outputs a score. Cohere and Amazon do not publish their architectures. A System One classifier answers typed questions about a state with probabilities in a single call. The name follows Kahneman’s fast, intuitive System 1 [1]; a model that reasons before answering is the deliberate System 2.
Stuga’s LLM judge sends this system prompt, with each candidate’s title, heading path and the 1,200 characters of its text that best match the question (Section 4.3):
You are a search relevance judge. Given a user query and numbereddocument snippets, rate how well EACH snippet helps answer or act on the query.Score each 0-10 (10 = directly answers it; 0 = irrelevant). Judge only relevance,not writing quality. Respond with ONLY a JSON array of {"i":<number>,"score":<0-10>}for every snippet, no prose.Its System One judge sends the same 24 passages to Jev [8] in one request:
{ "model": "jev-latest", "state": { "query": "…", "passages": { "p0": { "title": "…", "section": "…", "text": "…" }, "…": {} } }, "questions": { "p0": { "type": "noul", "instructions": "Does passage p0 help answer the query?", "criteria": { "true": "The passage states the information the query asks for, or information needed to answer it", "false": "The passage is only on a related topic, or is irrelevant to the query" } } }}3. Method
3.1 Test sets
We used two public benchmarks whose questions and relevance judgments were made by people, so that neither the questions nor the labels came from us or from a model under test. Both are openly licensed and published with the benchmark.
- MIRACL [13], in MMTEB’s reranking version [14]: questions over Wikipedia written by native speakers, who also judged which passages answer them. We drew 40 questions in each of English, Chinese, Japanese, Korean, Arabic and Thai, the languages of the privacy laws in Stuga’s sample workspaces, with a fixed seed from those with a relevant passage among their first 24 candidates. Passages are CC BY-SA 4.0; MIRACL’s judgments are Apache-2.0.
- BRIGHT [15], StackOverflow subset: 117 real StackOverflow questions, often with code. The gold passages are the documentation pages that accepted answers link to, confirmed by annotators. Finding them takes reasoning: the question describes a problem, the documentation describes a function. CC BY 4.0.
3.2 Candidates
Every ranker reordered the same 24 candidates per question. For MIRACL they are the first 24 of MMTEB’s published list for the question, in its order. Those lists hold a relevant passage in their first 24 for 79% of the six languages’ dev questions, and we drew ours from those.
For BRIGHT we retrieved the candidates as Stuga does, from the question alone, over the subset’s
107,081 passages, as BRIGHT splits them, in Stuga’s two configurations. With embeddings, BM25 and
vector search each return 96 passages, fused by reciprocal rank (k = 60), with one vector per
passage from OpenAI’s text-embedding-3-large at 1,024 dimensions; without them, BM25’s first 24
stand. A gold passage was among the 24 for 61 of
the 117 questions (52%) with embeddings and for
47 (40%) with BM25 alone. The rest are retrieval misses that no judge can
fix, so BRIGHT’s scores below count only the questions with a gold passage among the candidates,
unless they say all 117. Over all 117, no ranker put a gold passage first for more than
22%.
3.3 Judges
We compared current models across the listed families. OpenAI judges used OpenAI’s API or Bedrock’s
OpenAI-compatible endpoint, TypeSafe its own API; the other hosted models ran on Bedrock in
us-west-2, on US or global inference profiles. LLM judges received Stuga’s prompt verbatim, at the
lowest reasoning each model offers (low for GPT-6.1 Sol, none for GPT-6 Sol)
and with up to 8,192 output tokens, as Stuga’s judge sends them. Jev received
Stuga’s single request and, separately, one request per passage. Kev and Laya, which serve the same
System One API on local hardware, received one passage per request: Kev’s README notes training on
states of up to 384 tokens, and Laya, an encoder, reads its state as text, so it received plain text
and chose its English or multilingual checkpoint itself. Rerankers scored each question and passage
pair.
Every judge saw each candidate’s title, heading path and 1,200 characters of its text: the first 1,200, as Stuga’s judge read them until this study, and on BRIGHT also the 1,200 that best match the question, as it reads them now (Section 4.3). On MIRACL’s short passages the two differ for 1% of candidates, so its results stand for both.
3.4 Measures
- Hit@1, the share of questions with a relevant passage first, the passage the model reads first. Recall@8, the share with one among the eight passages Ask reads. 95% confidence intervals are percentile bootstrap over questions [9]. Rankers are compared on the same questions, by the paired difference and its bootstrap interval; where a ranker is compared with many others, the interval is also Bonferroni-adjusted for their number. A difference whose interval includes zero is within the margin of error: the test cannot tell the two rankers apart, which does not make them equal.
- Latency, the wall-clock time of each search from the client, retries included, reported as median and 95th percentile (for 28 retried GPT-6 Luna calls on MIRACL, the answering attempt only). Each ranker sent one request at a time.
- Cost per 1,000 searches, from measured tokens at list price: the Pi library’s model catalog,
with vendor-price overrides for hosted profiles recorded in
meta.json, TypeSafe’s announcement for Jev ($0.042 per million input tokens, output free) [8], and AWS’s pricing page for Cohere Rerank 3.5 ($2 per 1,000 queries) [10]. AWS lists no price for Amazon Rerank 1.0. An unusable answer counts toward latency and cost, because Stuga waits for it and pays for it. - Release date: the vendor’s own, otherwise the model’s first listing on OpenRouter or Hugging Face, each with its source in the repository.
3.5 Environment
The benchmark client and every local model ran on one Mac: Apple M6, 32 GB, macOS 27. Qwen3-Reranker ran as Q8_0 GGUF in llama.cpp 0.5.0, Kev-4B on MLX 0.32 in bfloat16, and Laya 0.3.20 on PyTorch 2.14, one model at a time.
4. Results
| Claude Opus 5.5 | LLM judge | Cloud | Sept 2026 | 81% (76%–86%) | +4.2 (0.0 to +8.8) | 61% (48%–72%) | 3.5 s | $37 | 0 |
| Qwen3-Reranker 4B | Reranker | Local | Jun 2025 | 81% (75%–86%) | +3.8 (−2.5 to +10.0) | 23% (13%–34%) | 5.3 s | local | 0 |
| Claude Sonnet 5.5 | LLM judge | Cloud | Sept 2026 | 80% (75%–85%) | +2.5 (−2.1 to +7.1) | 54% (41%–67%) | 2.0 s | $20 | 0 |
| Claude Sonnet 5 | LLM judge | Cloud | Jun 2026 | 80% (75%–85%) | +2.5 (−2.1 to +7.1) | 56% (43%–69%) | 3.7 s | $20 | 0 |
| Claude Haiku 4.5 | LLM judge | Cloud | Oct 2025 | 79% (74%–84%) | +2.1 (−2.9 to +6.7) | 39% (26%–52%) | 1.9 s | $9.3 | 0 |
| Cohere Rerank 3.5 | Reranker | Cloud | Dec 2024 | 78% (73%–83%) | +1.3 (−4.6 to +6.7) | 23% (13%–33%) | 163 ms | $2.0 | 0 |
| Jev, one request per passage | Classifier | Cloud | Sept 2026 | 78% (73%–83%) | +1.3 (−3.3 to +5.4) | 46% (33%–59%) | 244 ms | $0.64 | 0 |
| GPT-6.1 Sol | LLM judge | Cloud | Sept 2026 | 77% (72%–82%) | 0.0 (−5.0 to +4.6) | 52% (39%–66%) | 2.2 s | $14 | 0 |
| Jev | Classifier | Cloud | Sept 2026 | 77% (72%–82%) | — | 48% (34%–61%) | 131 ms | $0.37 | 0 |
| GLM-5 | LLM judge | Cloud | Feb 2026 | 77% (72%–83%) | 0.0 (−4.2 to +3.8) | 51% (39%–64%) | 2.7 s | $6.1 | 0 |
| Amazon Rerank 1.0 | Reranker | Cloud | Dec 2024 | 77% (71%–82%) | −0.4 (−6.3 to +5.4) | 20% (10%–30%) | 435 ms | not listed | 0 |
| Grok 4.6 | LLM judge | Cloud | Aug 2026 | 76% (70%–81%) | −1.3 (−6.3 to +3.3) | 56% (43%–67%) | 12 s | $20 | 0 |
| GPT-6 Sol | LLM judge | Cloud | Sept 2026 | 75% (70%–80%) | −2.1 (−7.1 to +2.9) | 49% (36%–62%) | 2.3 s | $11 | 0 |
| Qwen3 235B A22B 2507 | LLM judge | Cloud | Jul 2025 | 75% (69%–80%) | −2.1 (−7.5 to +2.5) | 44% (31%–56%) | 3.6 s | $1.2 | 0 |
| Qwen3-Reranker 0.6B | Reranker | Local | May 2025 | 75% (69%–80%) | −2.5 (−9.6 to +4.2) | 28% (18%–39%) | 987 ms | local | 0 |
| gpt-oss-120b | LLM judge | Cloud | Aug 2025 | 74% (68%–79%) | −3.3 (−7.5 to +0.8) | 51% (38%–64%) | 9.7 s | $1.2 | 0 |
| Llama 4 Maverick | LLM judge | Cloud | Apr 2025 | 73% (67%–78%) | −4.2 (−9.2 to +1.3) | 41% (30%–54%) | 945 ms | $1.2 | 0 |
| GPT-6 Luna | LLM judge | Cloud | Sept 2026 | 73% (67%–78%) | −4.2 (−9.6 to +1.3) | 51% (38%–64%) | 1.5 s | $0.52 | 11% |
| DeepSeek V3.2 | LLM judge | Cloud | Dec 2025 | 73% (67%–78%) | −4.6 (−10.0 to +0.8) | 54% (41%–67%) | 5.9 s | $3.2 | 3% |
| MiniMax M2.5 | LLM judge | Cloud | Feb 2026 | 71% (66%–77%) | −5.8 (−10.8 to −0.4) | 44% (31%–57%) | 8.6 s | $4.4 | 10% |
| Mistral Large 3 | LLM judge | Cloud | Dec 2025 | 66% (60%–72%) | −11.3 (−17.1 to −5.4) | 44% (31%–57%) | 7.0 s | $3.8 | 9% |
| Kev-4B | Classifier | Local | Sept 2026 | 64% (58%–70%) | −12.9 (−19.2 to −6.3) | 30% (18%–41%) | 5.4 s | local | 0 |
| Amazon Nova 2 Lite | LLM judge | Cloud | Dec 2025 | 62% (56%–68%) | −15.4 (−21.3 to −9.6) | 52% (39%–64%) | 1.3 s | $2.2 | 0 |
| Laya | Classifier | Local | Sept 2026 | 49% (43%–55%) | −27.9 (−34.6 to −20.4) | 16% (8%–26%) | 646 ms | local | 0 |
| No reranker (first-stage order) | — | — | — | 36% (30%–43%) | −40.8 (−48.3 to −33.8) | 36% (25%–48%) | — | — | 0 |
4.1 MIRACL: questions that ask for facts
The candidate lists here are MMTEB’s, so this section measures reranking a fixed list; it says nothing about how Stuga’s own retrieval would have ranked MIRACL. Reranking lifted Hit@1 from 36% to at most 81%, Claude Opus 5.5’s. Table 2 gives each ranker’s difference from Jev on the same 240 questions. Claude Opus 5.5 was 4.2 points ahead (0.0 to 8.8, or −2.9 to 10.8 adjusted for the 23 comparisons with Jev). Another 17 rankers, of all three kinds, scored between 73% and 81%, each within the margin of error of Jev even before adjustment. Adjusted, Mistral Large 3, Amazon Nova 2 Lite, Kev-4B and Laya remain clearly below Jev. What separates the leaders is latency and cost (Figure 1): Jev and Cohere Rerank 3.5 answer in 131 ms and 163 ms, the LLM judges among them 945 ms to 12 s. With any of these rankers a relevant passage is among the eight the model reads for 93%–100% of questions, against 80% without a judge.
4.2 By language
| Language | n | Jev | Claude Opus 5.5 | GPT-6 Sol | Cohere Rerank 3.5 | Qwen3-Reranker 4B (local) |
|---|---|---|---|---|---|---|
| English | 40 | 70% | 75% | 70% | 73% | 85% |
| Chinese | 40 | 70% | 88% | 78% | 75% | 78% |
| Japanese | 40 | 80% | 75% | 75% | 80% | 80% |
| Korean | 40 | 75% | 75% | 73% | 70% | 75% |
| Arabic | 40 | 83% | 90% | 78% | 88% | 88% |
| Thai | 40 | 85% | 85% | 78% | 85% | 80% |
With 40 questions per language, a difference of one question is 2.5 points, and no ranker is best in every language.
4.3 BRIGHT: questions that need reasoning
Reading each passage’s first 1,200 characters. With embeddings, search’s own order put a gold passage first for 36% of the 61 questions it found one for (Figure 4, Table 4). No ranker clearly beat that. The best LLM judge, GPT-6 Luna, was 6.6 points ahead (−6.6 to 19.7), and Jev 4.9 points behind (−21.3 to 11.5). Laya, Cohere Rerank 3.5 and Amazon Rerank 1.0 fell below it, though within the margin of error once adjusted for the 24 comparisons.
A judge that does not beat the first stage still changes its order, gaining some questions and losing others. GPT-6 Luna put a gold passage first for 11 questions where search had not, and moved one out of first place for 7; Jev did so for 12 and 15, Cohere Rerank 3.5 for 5 and 19.
Why search’s order held up. BRIGHT’s passages are long. The median is 4,000 characters, against 291 on MIRACL, and 99% of the gold passages run past the 1,200 characters a judge sees. Search scores each passage whole. Ordering the same 24 candidates by BM25 on whole passages put a gold passage first for 34% of questions, close to search’s own 36%; on their first 1,200 characters alone, for 13%. Most of the words that led BM25 to the gold passage lie beyond what a judge reads. Stuga’s own passages, whole sections of up to 12,000 characters, can be longer still.
Reading the 1,200 characters that best match. We designed this after the results above and measured it on the same questions, so it is exploratory until tested on others. We gave every judge, for each passage longer than 1,200 characters, the 1,200 that hold the most of the question’s words, the rarer ones counting for more, instead of the first 1,200, within the same 1,200-character limit. 21 of 24 judges rose, 7 of them clearly once adjusted for the 24 comparisons (Figure 4). Claude Opus 5.5 reached 61%, from 39%, and Jev 48%, from 31%; only Claude Opus 5.5 clearly beat search’s own order. Cohere Rerank 3.5 and Amazon Rerank 1.0 stayed below it, at 23% and 20%. Kev-4B, Qwen3-Reranker 4B and Laya, all on this machine, did not rise. Over all 117 questions, a retrieval miss counting as a miss, Claude Opus 5.5 put a gold passage first for 32% and Jev for 25%, against search’s 19%. With Jev a gold passage was among the eight the model reads for 93% of the 61 questions, against 72% without a judge. Stuga’s judges now read the best-matching 1,200 characters.
With and without embeddings. Table 4’s two Hit@1 columns count different questions: 36% of the 61 that search with embeddings found a gold passage for, and 40% of the 47 that BM25 alone found one for. With embeddings, search found a gold passage for every question BM25 alone did and for 14 more, but put it first for only 2 of those 14: they raise the count and lower the share. Over all 117 questions, search put a gold passage first for 19% with embeddings and 16% without. On BM25’s candidates the judges gained more. Reading the first 1,200 characters, Claude Opus 5.5 was 17.0 points ahead of search (−2.1 to 36.2); reading the best-matching 1,200, Claude Opus 5.5, Claude Sonnet 5.5 and MiniMax M2.5 clearly beat search’s 40%, at 72%, 70% and 68%.
| Ranker | BM25 + embeddings61 questions | BM25 only47 questions | ||||
|---|---|---|---|---|---|---|
| First 1,200 | Best 1,200 | vs first stage | First 1,200 | Best 1,200 | vs first stage | |
| No reranker (first-stage order) | 36% | — | 40% | — | ||
| No reranker, over all 117 questions | 19% | — | 16% | — | ||
| Claude Opus 5.5 | 39% | 61% | +24.6 (+9.8 to +39.3) | 57% | 72% | +31.9 (+14.9 to +48.9) |
| Claude Sonnet 5 | 34% | 56% | +19.7 (+4.9 to +34.5) | 53% | 66% | +25.5 (+8.5 to +42.6) |
| Grok 4.6 | 31% | 56% | +19.7 (+1.6 to +36.1) | 47% | 66% | +25.5 (+6.4 to +44.7) |
| Claude Sonnet 5.5 | 38% | 54% | +18.0 (+1.6 to +34.4) | 51% | 70% | +29.8 (+12.8 to +46.8) |
| DeepSeek V3.2 | 25% | 54% | +18.0 (+3.3 to +32.8) | 38% | 60% | +19.1 (+2.1 to +36.2) |
| GPT-6.1 Sol | 33% | 52% | +16.4 (0.0 to +32.8) | 51% | 64% | +23.4 (+6.4 to +42.6) |
| Amazon Nova 2 Lite | 36% | 52% | +16.4 (+4.9 to +27.9) | 47% | 51% | +10.6 (−2.1 to +23.4) |
| GPT-6 Luna | 43% | 51% | +14.8 (0.0 to +29.5) | 55% | 64% | +23.4 (+6.4 to +40.4) |
| GLM-5 | 36% | 51% | +14.8 (0.0 to +29.5) | 49% | 68% | +27.7 (+8.5 to +46.8) |
| gpt-oss-120b | 25% | 51% | +14.8 (−1.6 to +31.1) | 40% | 57% | +17.0 (0.0 to +34.0) |
| GPT-6 Sol | 28% | 49% | +13.1 (−1.6 to +27.9) | 49% | 64% | +23.4 (+4.3 to +42.6) |
| Jev | 31% | 48% | +11.5 (−4.9 to +27.9) | 40% | 60% | +19.1 (0.0 to +38.3) |
| Jev, one request per passage | 31% | 46% | +9.8 (−6.6 to +24.6) | 40% | 57% | +17.0 (0.0 to +34.0) |
| MiniMax M2.5 | 34% | 44% | +8.2 (−8.2 to +26.2) | 47% | 68% | +27.7 (+12.7 to +44.7) |
| Qwen3 235B A22B 2507 | 33% | 44% | +8.2 (−6.6 to +23.0) | 51% | 60% | +19.1 (+2.1 to +38.3) |
| Mistral Large 3 | 30% | 44% | +8.2 (−4.9 to +21.3) | 38% | 43% | +2.1 (−12.8 to +17.0) |
| Llama 4 Maverick | 36% | 41% | +4.9 (−9.8 to +19.7) | 47% | 49% | +8.5 (−8.5 to +27.7) |
| Claude Haiku 4.5 | 33% | 39% | +3.3 (−13.1 to +18.1) | 40% | 60% | +19.1 (+2.1 to +36.2) |
| Kev-4B | 31% | 30% | −6.6 (−19.7 to +8.2) | 38% | 47% | +6.4 (−8.5 to +23.4) |
| Qwen3-Reranker 0.6B | 21% | 28% | −8.2 (−23.0 to +6.6) | 36% | 45% | +4.3 (−10.6 to +19.1) |
| Qwen3-Reranker 4B | 23% | 23% | −13.1 (−29.5 to +4.9) | 40% | 47% | +6.4 (−8.5 to +21.3) |
| Cohere Rerank 3.5 | 13% | 23% | −13.1 (−27.9 to +1.6) | 36% | 51% | +10.6 (−4.3 to +27.7) |
| Amazon Rerank 1.0 | 16% | 20% | −16.4 (−31.1 to −1.6) | 36% | 45% | +4.3 (−12.8 to +21.3) |
| Laya | 23% | 16% | −19.7 (−31.1 to −8.2) | 28% | 23% | −17.0 (−29.8 to −4.3) |
4.4 How LLM judges fail
Stuga reads the scores wherever an answer states them unambiguously: as the JSON array it asks for,
as the same objects one per line or followed by a note, or as [n] score lines. Four failure modes
remained, all from producing the ranking as text. An unusable answer leaves the first-stage order.
- The answer ignores the format. GPT-6 Luna sometimes returned a bare list of numbers, only the
best passage, or
[index, score]pairs: 11% of its MIRACL answers were unusable. - The answer is corrupted. DeepSeek V3.2 sometimes wrote a stray token where a score belongs, or text unrelated to the task, in 3% of its MIRACL answers.
- The answer does not end. Mistral Large 3 repeated its list until the output limit, and MiniMax M2.5’s reasoning ran into it: 9% and 10% of their MIRACL answers were unusable.
- Scores tie. Two passages shared the top score in 38% of GPT-6 Sol’s MIRACL answers and 24% of Claude Opus 5.5’s; the first-stage order broke the tie. Jev’s probabilities tied in 13%.
Reasoning costs time. Grok 4.6 wrote 1,623 output tokens per search on average for a list of 24 scores, where GPT-6 Sol wrote 198, and took 12 s.
4.5 Cost
Jev counted 8,801 input tokens per MIRACL search and costs input only. Claude Opus 5.5 read 7,093 and wrote 273. Cohere Rerank 3.5 is billed per query, whatever its length.
4.6 Local judges
Qwen3-Reranker 4B reached 81% on MIRACL, level with the hosted leaders, in 5.3 s: a cross-encoder reads the question with each of the 24 passages, and Thai, Japanese and Korean text is long in tokens. The 0.6B model reached 75% in 987 ms.
Two open projects serve the System One API: Kev [11] and Laya [12]. Kev-4B reached 64% in 5.4 s. Laya, whose English checkpoint is ModernBERT-large and whose multilingual one is mmBERT-base, reached 49% in 646 ms; its README places further accuracy in fine-tuning.
5. Choosing a first stage and a judge
In this study’s setup, the results suggest the following defaults. None was tested on a team’s own documents, with other chunking or with real users’ questions, and each holds only as far as two benchmarks of this size allow (Section 7).
Table 5. Choosing a first stage and a judge, from the results above.
| Situation | Suggestion | Evidence |
|---|---|---|
| By default | Add embeddings to BM25 | On BRIGHT, search with embeddings found a gold passage for 61 rather than 47 of 117 questions, losing none that BM25 alone found; a judge cannot rank what search missed. |
| Questions that ask for facts | Rank with Jev | On MIRACL, 77% in 131 ms at $0.37 per 1,000 searches; no ranker was clearly better. |
| A few points more, whatever the wait and cost | An LLM judge such as Claude Opus 5.5 | 4.2 points more than Jev on MIRACL, not shown to be better, at 27 times the latency and 100 times the cost; some LLM judges return unusable answers (Section 4.4). |
| Text that must stay on the machine | Qwen3-Reranker 4B, locally | 81% on MIRACL, but 5.3 s per search on an Apple M6; the 0.6B model 75% in 987 ms. Kev and Laya reached 64% and 49%. |
| Long passages | Show the judge the part that matches the question | On BRIGHT, 21 of 24 judges rose reading the best-matching 1,200 characters instead of the first 1,200, Jev from 31% to 48%; exploratory, measured on the questions it was designed on. |
| Questions that need reasoning | An LLM judge, and better retrieval | On BRIGHT, reading the best-matching 1,200 characters, Claude Opus 5.5 reached 61% against search’s 36% and Jev 48%; search still missed a gold passage for 48% of questions. |
For Stuga’s defaults this means Semantic search on and Reranking with Jev, reading the part of each passage that best matches the question. A team whose questions are mostly troubleshooting may gain more from an LLM judge, at its latency and cost; rewriting the question before retrieving helped on BRIGHT in its authors’ tests [15], and we did not test it in Stuga.
6. In Stuga
Reranking is set up under Settings → This node → AI providers, beside Built-in AI and Semantic search, with a TypeSafe key, directly or through OpenRouter. Every search from Ask and from agents then goes through it. Without it, the chat model judges while Built-in AI is on; with neither, fusion order stands.


A hosted judge sees what it ranks: the question and up to 24 passages, 1,200 characters of each, go to TypeSafe or OpenRouter, as they go to the chat provider when it judges. Reranking accepts any System One endpoint, so a server such as Kev on the same machine keeps everything local, at 64% on MIRACL and 5.4 s per search.
Ranking decides what an agent reads. What it writes back to a document or table waits, by default, for a person to accept or reject: how that review works.
7. Limitations
- MIRACL’s candidates are MMTEB’s lists, not Stuga’s retrieval, and they count every candidate no annotator marked relevant as not relevant, so a relevant passage nobody saw counts against whichever ranker puts it first.
- Both benchmarks are public and older than every model tested, so their questions may be in the models’ training data, which would favour the LLM judges.
- With 61 rankable BRIGHT questions and 40 per MIRACL language, differences of a few points are within noise. A paired interval that includes zero means this test cannot tell two rankers apart, not that they are equivalent.
- We measured the ranking, not the answers Ask writes from it: a better first passage need not mean a better answer or citation.
- BRIGHT’s authors show that reasoning about a query before retrieving changes what is found [15]; Stuga retrieves from the question as asked, and so did we.
- Wikipedia articles and StackOverflow documentation are not a team’s own documents; other collections may rank differently.
- Latency depends on vendor load and the network; it was measured from one client.
- Reading the part of each passage that best matches the question was designed after the first BRIGHT results and measured on the same questions; it needs testing on other data.
- We did not vary the number of candidates, how many characters a judge sees of each, or the prompt, and varied which characters on BRIGHT only; Section 4.3 shows that choice moves the results.
- Every LLM received the same prompt, untuned. Tuning prompts or reasoning per model would move some results.
8. Reproducibility
Code, the two test sets, every candidate list and every call’s result are at stuga-dev/rerank-bench (code MIT; data under each source’s licence). The test sets are built from pinned revisions of the public datasets.
git clone https://github.com/stuga-dev/rerank-bench && cd rerank-bench && pnpm installBENCH_KEYS=keys.json node src/run.ts --set miracl --only jev,gpt-6-sol --limit 5node src/report.ts results/miracl/<date>A ranker maps a query and 24 passages to 24 scores; pull requests with new rankers are welcome.
Stuga has no business relationship with TypeSafe or any other vendor named here; all calls were paid at list price.
References
- D. Kahneman. Thinking, Fast and Slow. Farrar, Straus and Giroux, 2011.
- J. Devlin, M.-W. Chang, K. Lee, K. Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL, 2019.
- R. Nogueira, K. Cho. Passage Re-ranking with BERT. arXiv:1901.04085, 2019.
- W. Sun et al. Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents. EMNLP, 2023.
- S. Robertson, H. Zaragoza. The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval 3(4), 2009.
- G. V. Cormack, C. L. A. Clarke, S. Büttcher. Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods. SIGIR, 2009.
- B. Warner et al. Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference. arXiv:2412.13663, 2024.
- TypeSafe. Introducing System One models and Jev. 15 September 2026.
- B. Efron, R. J. Tibshirani. An Introduction to the Bootstrap. Chapman & Hall, 1993.
- Amazon Web Services. Amazon Bedrock pricing, retrieved 28 September 2026.
- J. Palmer. Kev and kev-4b, 2026.
- Convai Innovations. Laya and Laya multilingual (models); laya 0.3.20 (code), 2026.
- X. Zhang, N. Thakur, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, X. Li, Q. Liu, M. Rezagholizadeh, J. Lin. MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages. TACL 11, 2023.
- K. Enevoldsen et al. MMTEB: Massive Multilingual Text Embedding Benchmark. arXiv:2502.13595, 2025.
- H. Su et al. BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval. arXiv:2407.12883, 2024.