Inside instructkr/reranker-simple-benchmark: The Korean Reranker Benchmark That Starts Before the Reranker
A lightweight evaluation harness that shows how BM25, Korean tokenization, and a two-stage funnel shape what rerankers can actually prove.
- The repo argues that reranker evaluation is really a test of the retrieval funnel that feeds the model.
- In Korean search, tokenizer choice is part of the benchmark, not a preprocessing footnote.
- The main engineering value is a thin wrapper layer that lets many rerankers share one evaluation path.
- Its simplicity is the point, because a local reproducible harness is easier to trust than a heavyweight benchmark stack.
This benchmark looks modest on the surface, but its thesis is sharper than a standard leaderboard. It says a reranker never sees the full corpus, only the candidate set that stage 1 hands it. Once you accept that, the interesting question stops being "which model won" and becomes "what kind of retrieval funnel made that win possible?"
The benchmark is really measuring the funnel
The project's own README makes the design plain: a stage 1 retrieval step narrows each query to a limited corpus, then stage 2 reranks the smaller pool. That matters because candidate quality is not neutral. If the first pass misses the right documents, the reranker cannot recover them, no matter how strong the model looks on paper.
본 프로젝트는 Reranker Benchmark Evaluation을 최소한의 의존성으로 경량화하여, 누구나 쉽게 실행하고 즉각적인 결과를 얻을 수 있도록 설계되었습니다.
Why this repo exists at all
The gap it fills is simple to describe and annoying to solve. MTEB gives broad benchmarking discipline, but this repository turns that discipline into a Korean reranking workflow that is lighter to run and easier to adapt. It also extends the MTEB ecosystem with custom tasks, which makes the project feel less like a wrapper and more like a local benchmark layer built for a specific language problem.
본 프로젝트에서는 BM25 기반의 Stage 1 Retrieval을 통해 각 벤치마크 query 당 retrieval corpus를 1000개로 제한합니다.
How the pipeline works, stage by stage
The implementation is a clean two-stage stack. Stage 1 uses BM25 to retrieve a wide candidate set, and the project's notes show that Mecab was chosen as the preferred Korean tokenizer after comparing recall and F1 against alternatives like Kiwi and Okt. Stage 2 then reranks the smaller set, and the repo patches the evaluation path so it can reuse the precomputed retrieval results instead of reindexing everything from scratch.
class BaseRerankerWrapper:
def predict(self, query: str, documents: list[str]) -> list[float]:
raise NotImplementedError
class JinaRerankerV3Wrapper(BaseRerankerWrapper):
def predict(self, query: str, documents: list[str]) -> list[float]:
return self.model.rerank(query, documents)
class Qwen3RerankerWrapper(BaseRerankerWrapper):
def predict(self, query: str, documents: list[str]) -> list[float]:
prompt = self.build_prompt(query, documents)
return self.model(prompt)
The tokenizer is not a footnote
This is where the project gets quietly opinionated. In English retrieval, tokenization often fades into the background. In Korean, morphological segmentation changes the shape of the input itself, which means it changes the benchmark before the reranker even gets involved. Mecab, Kiwi, Okt, and Kkma are not interchangeable here. They are competing interpretations of where the words begin and end.
One interface, many rerankers
The repo's other real contribution is abstraction. Its wrapper layer normalizes different model families so they can be evaluated through the same pipeline, whether the model is a prompt-driven Qwen-style reranker, a Jina API wrapper, or a BGE-family implementation. That keeps the benchmark focused on ranking behavior instead of on adapter glue.
What this replaces, and what it does not
| Project | Primary job | Strength | Limit |
|---|---|---|---|
| instructkr/reranker-simple-benchmark | Evaluate Korean rerankers | Controls the retrieval funnel and tokenizer choice | Narrow by design |
| MTEB | Broad benchmark suite | Standardizes many tasks across embedding models | Not tailored to this reranking workflow |
| rerankers | Unified reranker API | Makes model swapping easy | Does not benchmark datasets |
That contrast is the project's product idea in one frame. It turns a complicated evaluation problem into a local workflow that a researcher can actually run, inspect, and modify. In a space where benchmarks often feel ceremonial, that simplicity is the innovation.