Inside instructkr/reranker-simple-benchmark: The Korean Reranker Benchmark That Starts Before the Reranker

A lightweight evaluation harness that shows how BM25, Korean tokenization, and a two-stage funnel shape what rerankers can actually prove.

11 min read • View on GitHub • More from instructkr

A query slips through a mechanical sorting system. It passes first through a tokenizer gate, then a BM25 sieve, then into a smaller chamber where only a handful of documents are judged again. The scene explains that reranking depends on the documents chosen earlier in the pipeline.
The benchmark's real subject is the funnel that feeds the reranker.
Key Takeaways

This benchmark looks modest on the surface, but its thesis is sharper than a standard leaderboard. It says a reranker never sees the full corpus, only the candidate set that stage 1 hands it. Once you accept that, the interesting question stops being "which model won" and becomes "what kind of retrieval funnel made that win possible?"

The benchmark is really measuring the funnel

The project's own README makes the design plain: a stage 1 retrieval step narrows each query to a limited corpus, then stage 2 reranks the smaller pool. That matters because candidate quality is not neutral. If the first pass misses the right documents, the reranker cannot recover them, no matter how strong the model looks on paper.

본 프로젝트는 Reranker Benchmark Evaluation을 최소한의 의존성으로 경량화하여, 누구나 쉽게 실행하고 즉각적인 결과를 얻을 수 있도록 설계되었습니다.

README, Project Documentation · instructkr repo README

Why this repo exists at all

The gap it fills is simple to describe and annoying to solve. MTEB gives broad benchmarking discipline, but this repository turns that discipline into a Korean reranking workflow that is lighter to run and easier to adapt. It also extends the MTEB ecosystem with custom tasks, which makes the project feel less like a wrapper and more like a local benchmark layer built for a specific language problem.

본 프로젝트에서는 BM25 기반의 Stage 1 Retrieval을 통해 각 벤치마크 query 당 retrieval corpus를 1000개로 제한합니다.

README, Project Documentation · instructkr repo README

How the pipeline works, stage by stage

The system is not just retrieval plus ranking. It is retrieval, filtering, normalization, reranking, and measurement.

The implementation is a clean two-stage stack. Stage 1 uses BM25 to retrieve a wide candidate set, and the project's notes show that Mecab was chosen as the preferred Korean tokenizer after comparing recall and F1 against alternatives like Kiwi and Okt. Stage 2 then reranks the smaller set, and the repo patches the evaluation path so it can reuse the precomputed retrieval results instead of reindexing everything from scratch.

class BaseRerankerWrapper:
    def predict(self, query: str, documents: list[str]) -> list[float]:
        raise NotImplementedError


class JinaRerankerV3Wrapper(BaseRerankerWrapper):
    def predict(self, query: str, documents: list[str]) -> list[float]:
        return self.model.rerank(query, documents)


class Qwen3RerankerWrapper(BaseRerankerWrapper):
    def predict(self, query: str, documents: list[str]) -> list[float]:
        prompt = self.build_prompt(query, documents)
        return self.model(prompt)

The tokenizer is not a footnote

This is where the project gets quietly opinionated. In English retrieval, tokenization often fades into the background. In Korean, morphological segmentation changes the shape of the input itself, which means it changes the benchmark before the reranker even gets involved. Mecab, Kiwi, Okt, and Kkma are not interchangeable here. They are competing interpretations of where the words begin and end.

A close-up of a single Korean sentence being split by several small hand tools into different token pieces. One tool leaves a neat set of fragments, while others produce messier cuts. The scene explains that token boundaries can change retrieval quality before reranking begins.
Tokenization is part of the measurement, not a preprocessing afterthought.

One interface, many rerankers

The repo's other real contribution is abstraction. Its wrapper layer normalizes different model families so they can be evaluated through the same pipeline, whether the model is a prompt-driven Qwen-style reranker, a Jina API wrapper, or a BGE-family implementation. That keeps the benchmark focused on ranking behavior instead of on adapter glue.

What this replaces, and what it does not

ProjectPrimary jobStrengthLimit
instructkr/reranker-simple-benchmarkEvaluate Korean rerankersControls the retrieval funnel and tokenizer choiceNarrow by design
MTEBBroad benchmark suiteStandardizes many tasks across embedding modelsNot tailored to this reranking workflow
rerankersUnified reranker APIMakes model swapping easyDoes not benchmark datasets
A split scene shows two approaches to evaluation. On the left is a heavy, generalized benchmarking machine with many gears and cables. On the right is a compact local workstation with a clean leaderboard and a few simple controls. The contrast explains the value of a focused, reproducible benchmark.
The point is not more machinery. The point is a clearer evaluation path.

That contrast is the project's product idea in one frame. It turns a complicated evaluation problem into a local workflow that a researcher can actually run, inspect, and modify. In a space where benchmarks often feel ceremonial, that simplicity is the innovation.