One Citation, Two Checks
Presets:
1. Paraphrase (Low Score)
2. Keyword Inversion
3. Decimal Bug (3.5)
Run Step
Reset
01.
Input Sentence
Sentence with Citations
The model exhibits robust performance across standard benchmarks [doc_1:chunk_0], but fails on edge cases like section 3.5 when overloaded [doc_2:chunk_1]. Invalid ref [doc_9:chunk_9].
Active Fixture Chunk
doc_1:chunk_0 (Genuinely Relevant)
doc_2:chunk_1 (Weakly Relevant)
doc_3:chunk_x (Unrelated Chunk)
doc_1:chunk_0
"Standard benchmarks show high robustness and consistency across evaluation suites under normal parameters."
Stage Progress:
Split
Validate
Score
Bucket
02.
Structural Validity
Layer 1
Parsed Citation Markers
Flagged / Invalid Bin
03.
Lexical Groundedness
Layer 2
Boundary Split Engine
⚠ Bug: Splitter fragmented on decimal number "3.5"
TF-IDF Cosine Similarity