Pattern-miner-mm2: Mining Surprising Logic Inside AtomSpace
A MeTTa pipeline that turns frequent patterns into canonical identities, scores them by information-theoretic surprise, and feeds them back into a symbolic reasoning stack.
- Pattern-miner-mm2 treats pattern identity as a first-class object, so alpha-equivalent logic can be cached, reused, and scored instead of rediscovered from scratch.
- The pipeline is intentionally split into frequency first and surprise second, which keeps candidate generation separate from the question of whether a pattern is actually informative.
- Its main novelty is not graph mining in the generic sense, but symbolic discovery inside AtomSpace where outputs can feed back into reasoning.
- The repo is a port in motion, and that matters because the MM2 and MORK stack is being shaped to support logic-native performance rather than row-oriented analytics.
Most miners ask one question: what occurs often? Pattern-miner-mm2 asks a sharper one: what occurs often enough to matter, and weirdly enough to be worth noticing? That difference sounds small until you realize the whole system is built to answer it inside AtomSpace, where patterns are not just rows of data. They are candidate knowledge.
That is why the repo feels unusual. It is not a conventional graph-mining library wrapped around a scoring function. It is a MeTTa pipeline that makes pattern discovery, canonicalization, caching, and surprisingness part of the same symbolic workflow.
One particular Hyperon module is Pattern Miner, which uses information-theoretic surprisingness criteria. Thus, while Hyperon doesn't insist on implementing AGI systems on top of universal induction, it facilitates the use of elements of algorithmic information theory.
Why surprising matters more than frequent
In a symbolic reasoning stack, pure frequency is a blunt instrument. A pattern can be common and still be boring. Pattern-miner-mm2 instead tries to surface regularities that are statistically informative, which is a better fit for knowledge graphs where the interesting thing is often the relation between parts, not just repetition.
The repo’s framing is straightforward: mine frequent candidates, then judge them with information-theoretic surprisingness. That sequencing matters because it separates discovery from evaluation. The first stage asks what exists. The second asks what is worth keeping.
Pattern discovery through Pattern Miner finds proof motifs (induction schemas, commutativity tricks) ranked by I-surprisingness.
The pipeline: first mine, then judge
The orchestration layer in src/miner/pattern-miner.metta makes the handoff explicit. It runs the frequent miner, then conditionally calls the surprisingness stage. If surprise mode is off, the output stops at a plain frequent pattern. If it is on, the same candidate gets enriched into a frequency-plus-surprise result.
(DEF run-pattern-miner
(if (= SURP-MODE none)
(run-frequent-pattern-miner)
(frequent-with-surprisingness
(run-frequent-pattern-miner)
(do-surp))))
That architecture is the repo’s quiet discipline. It does not mix counting and judging in one opaque step. It makes the pipeline legible, which matters when the downstream consumer is a reasoning system rather than a dashboard.
The real trick: canonical variables become cache keys
This is the conceptual center of the project. In logic, two patterns can be alpha-equivalent, meaning they are structurally identical even if their variable names differ. Pattern-miner-mm2 uses canonical indexing, via vars_to_indices, to collapse those variants into one machine-readable identity.
That identity is more than a cleanup step. It becomes the key for support caching, which means the miner can attach frequency and score data to the pattern itself rather than to one arbitrary naming version of it. That is a major leverage point in symbolic search.
| Approach | What it treats as the unit of work | What it can reuse |
|---|---|---|
| Raw pattern names | Each variable-renamed version as separate | Almost nothing beyond a single occurrence |
| Canonical identity | One alpha-equivalence class of patterns | Support counts, cache entries, and surprise scores |
| Pattern-miner-mm2 | The normalized symbolic object itself | The full discovery pipeline across stages |
Once you see it, the design becomes obvious in hindsight. The miner is not just evaluating patterns. It is building a shared symbolic address book.
How the frequent miner expands the search space without getting lost
The frequent-mining side lives in src/freq/, where the code grows patterns from concrete structure into generalized candidates. The interesting part is not that it searches a large space. It is that it uses shallow abstraction, recursion checks, and triplet expansion to keep that search from collapsing into noise.
One module abstracts constants into variables. Another prevents runaway nesting. Another grows connected triplets into larger conjunctions. Together they form a small discovery engine that behaves more like a symbolic compiler pass than a generic database scan.
; Conceptual flow from the freq tree
(get-vars-for-tree pattern)
(get-value-at-position tree path)
(has-nested-expression-step expr)
(conjunction-expansion-triplet seed support-cache)
That matters because symbolic data is sparse, nested, and relational. A flat frequency counter would miss the structure that makes AtomSpace useful in the first place.
What surprisingness actually measures
The repo offers two scoring paths: ISurp and JSD. The naming sounds abstract, but the practical question is simple. Does the pattern’s observed behavior diverge from what its parts would predict?
ISurp compares empirical probability with estimated probability derived from components. JSD compares distributions of truth values. Both turn the miner from a raw enumerator into a ranking system for symbolic significance.
| Metric | What it compares | Why it helps |
|---|---|---|
| ISurp | Observed pattern probability vs expected probability | Highlights patterns that are unexpectedly informative |
| JSD | Truth-value distributions across patterns and parts | Surfaces divergence in symbolic belief structure |
| Support alone | Only occurrence counts | Finds commonality, but not significance |
The point is not that one metric is universally better. The point is that the repo is building a scoring layer that can distinguish a familiar motif from a useful one.
Why this repo exists in MeTTa instead of Python
This codebase is not using MeTTa by accident. The language choice matches the problem. Pattern-miner-mm2 needs a representation where patterns, rules, and execution steps are all native objects in the same symbolic world.
That is why the repository’s structure matters. MM2 and MORK are not just runtime details. They are the environment that makes a logic-native pipeline plausible in the first place. A row-oriented stack would force this into a different shape, and probably a less interesting one.
Completing the picture a MeTTa sub-language called MM2 has also been launched, specialized for extremely efficient implementation of performance-critical AI algorithms directly against the MORK infrastructure.
| Dimension | Mainstream graph mining | Pattern-miner-mm2 |
|---|---|---|
| Primary unit | Graph motifs or feature vectors | Canonical symbolic patterns |
| Core ranking signal | Frequency or structural score | Frequency plus information-theoretic surprise |
| Execution model | Usually external to the graph | Native to the MeTTa and AtomSpace stack |
| Downstream use | Analytics or retrieval | Further reasoning and rule generation |
The maintenance story: a port in motion
The presence of modular src/surp/ code alongside older references such as isurp-old.metta suggests a repo in transition, not a frozen demo. The architecture is being reshaped to separate frequency, surprise, and orchestration more cleanly while preserving behavior.
That kind of porting work is easy to underestimate. It is not just code cleanup. It is how a research prototype becomes a maintainable subsystem. The value is in making the old logic survive a new runtime without losing its meaning.
Seen that way, Pattern-miner-mm2 is really about infrastructure for symbolic discovery. It tries to make discovery objects durable enough to be reused, compared, cached, and fed back into reasoning.
The broader lesson is bigger than this repository. If you want a reasoning stack to learn from its own graph, you need more than counting. You need canonical identity, cacheable support, and scores that tell you when a pattern is not just present, but informative.
That is what Pattern-miner-mm2 is reaching for. It is a miner, but it is also a memory system for symbolic regularities. In AtomSpace, that may be the more important half.