Shortest: the LLM that decides how much certainty your test run needs
A tiny shell pipeline turns diffs, specs, and a confidence setting into the smallest test set that still feels safe enough to ship.

Shortest is an open-source AI-powered testing framework that allows developers to write end-to-end tests in plain English, as simple as: shortest(“user can sign up and create a $5 product”) More tests = fewer regressions as AI writes and ships code.
- Shortest does not try to replace testing. It turns test scope into a tunable confidence budget.
- The repo's real trick is a small shell pipeline that gathers diffs, filters candidate specs, and narrows to exact examples with strict JSON.
- Structured outputs matter because they let a probabilistic model feed deterministic tools like `jq` and `rspec` without turning the shell into guesswork.
- The project bets that teams will accept calibrated certainty when the feedback loop is fast and the decision path stays auditable.
The interesting thing about Shortest is not that it uses an LLM. It is that it refuses to treat test selection as binary. Instead of asking whether a change is good or bad, it asks a cleaner question: how much evidence do you need before you trust it?
The confidence slider
That framing changes the job of the tool. The public pitch leans on plain-English tests, but the repo's sharper contribution is more operational: it tries to price certainty. At 80 percent confidence, you can move faster. At 99.9 percent, you pay for more coverage. The point is not to eliminate risk. The point is to make the risk visible enough to choose.
Most teams shrink test runs in blunt ways. They run everything and wait. They hand-pick files and hope they remembered the edge cases. They add path filters or impact analysis, which help until the code shape changes. Shortest sits in a smaller, stranger lane. It uses the model to choose a narrower suite, but keeps the workflow legible enough to inspect.
| Workflow | Scope decision | Speed | Risk | Best fit |
|---|---|---|---|---|
| Run the full suite | No selection step. Everything runs. | Slowest. | Lowest selection risk, highest time cost. | Release gates and high-stakes merges. |
| Manual filtering | A human picks likely tests. | Fast when the codebase is familiar. | Easy to miss coupling or stale assumptions. | Small teams with strong code intuition. |
| Shortest | An LLM narrows scope against a confidence target. | Fast enough for everyday iteration. | Tunable. Lower confidence is quicker, higher confidence runs more. | Changes where speed matters but blind spots still hurt. |
Inside `shortest.sh`
The implementation is deliberately small. In `ruby-on-rails/shortest.sh`, the script gathers local and PR diffs, filters them down to relevant paths, asks the model which spec files and specific specs matter, then executes `bundle exec rspec` on the result. The supporting cast is pure Unix: `gh`, `awk`, `curl`, `jq`, and `set -euo pipefail` to keep the pipeline honest.
# 1. collect context
combined_diff="$(get_combined_diff)"
# 2. ask which spec files are worth considering
spec_files="$(determine_spec_files "$combined_diff")"
# 3. narrow to exact examples or line ranges
specs_to_run="$(determine_specs "$combined_diff" "$spec_files")"
# 4. execute only the selected tests
bundle exec rspec "${specs_to_run[@]}"
Why structured outputs matter more than the LLM itself
The clever part is not that a model can make a guess. It is that the guess comes back as strict JSON. That means a shell script can trust the shape of the answer, parse it with `jq`, and fail fast if the output is malformed. In other words, the model is treated like a decision component, not like a chatty teammate. Probabilistic inference feeds deterministic automation, and that bridge is the whole product.
That matters because the downside of test selection is false confidence. If the model misses a critical example, the run looks clean when it should not. Shortest answers that by making confidence an explicit setting, not a promise of perfection.
Why Sahil's product instinct matters
Sahil Lavingia's projects usually prefer small surfaces to sprawling platforms, and Shortest follows that instinct. The repo reads less like a moonshot and more like a tool you can paste into a terminal and reason about line by line. That is why the shell script matters. It leaves the hard parts visible.
The origin also explains the tone of the project. This is not a company building a giant testing platform first and a workflow second. It is a founder's bet that a tighter loop, even with calibrated uncertainty, is more valuable than a slower guarantee. Relative to Playwright wrappers and broader AI testing platforms, Shortest stays narrower on purpose. It wants to choose what should run, not own the whole QA stack.
How do you know that the test is actually doing what you think it should?
The right objection
That question is the one that matters. Any test reducer can create a comforting illusion of coverage. Shortest does not answer with certainty. It answers with a knob and a workflow, which is more honest than pretending the risk is gone.
That is also the broader market signal. The ecosystem keeps adding more AI test tools, but many of them try to be the whole framework. Shortest is more disciplined. It narrows scope, documents the decision, and gets out of the way.
What Shortest is really betting on
The bet is not that AI replaces disciplined testing. The bet is that software teams will pay for calibrated confidence if the system is fast, local, and legible. That is a useful shift. Instead of chasing total certainty, teams can choose a threshold that fits the moment, then revisit it when the stakes change.
In that sense, Shortest is a very Unix answer to an LLM problem. Small pieces. Clear inputs. One narrow job. If it works, you get a faster loop without surrendering control.