Engine 02 of eleven

Analytical and data-mining intelligence.

Patterns in your ledger, labelled as patterns.

Data mining is not prediction. It describes what is already in your records — which customers behave alike, which products move together, when a series changed, what does not fit — and it says so as description, never as cause.

Descriptive · Effect size and stability · Never rendered as cause
How it works

Ten kinds of pattern, over a frozen dataset.

Every mining run works from a dataset that is frozen and fingerprinted before the run starts, so the same question asked twice gives the same answer and the population behind a finding can be reconstructed months later.

  • Segmentation and cohorts — groups of customers, suppliers or products that behave alike on the dimensions you actually trade on.
  • Frequent patterns and associations — what is bought together, what is returned together, which codings recur on which supplier.
  • Sequence patterns — the orders in which things tend to happen, before process mining tells you whether that was allowed.
  • Change-point detection — the date a series genuinely changed behaviour, rather than the date somebody noticed.
  • Robust outliers — values that do not fit, computed with methods that are not themselves thrown by the outlier.
  • Duplicate and near-duplicate discovery — the same party or product entered twice under two spellings.
  • Missingness and data-quality patterns — where the gaps cluster, which is usually a process problem rather than a data problem.
  • Model residual and blind-spot analysis — where the predictive engine is consistently wrong, which is how it gets better.
  • Candidate explanatory drivers — what covaries, offered as candidates for investigation and never as findings.
  • Graph-community discovery, where the relationships make that safe.
What it produces

A mining insight, with its own honesty built in.

The envelope is deliberately hard to quote out of context. It carries the population, the method, the effect size, the stability and the comparison period, so a finding cannot be reduced to a number on a slide without the caveats travelling with it.

  • The business question it was run to answer, and the population and filters it ran over.
  • The frozen dataset and its fingerprint — the run is reproducible.
  • The method and parameters, and the pattern type.
  • Support and prevalence, the effect size, and the stability or uncertainty around it.
  • The comparison period, and the multiple-testing treatment where it applies.
  • Representative evidence — the actual records behind the pattern.
  • Data-quality warnings, the limitations, and a descriptive-association label that cannot be removed.
The boundary

Association is never represented as causation.

This is the hardest boundary in the system to hold, because association reads like cause in plain English and a language model will happily make that translation for you. It is blocked in two places: the envelope carries a descriptive-association label that travels with the finding, and the language engine is explicitly forbidden from rewording a mined association into a causal claim. If you want cause, that is a different engine with a different evidential bar.

Where it shows up

Wherever the question starts with “which” or “what changed”.

  • “Why did this quote lose margin?” — the discount pattern behind it, with the comparison period that makes it meaningful.
  • “Which customers behave like this one?” — a cohort, with what defines it and how stable it is.
  • “What is also bought with this?” — the co-purchase suggestions on a sales order, with support and prevalence.
  • “Where are our data gaps?” — missingness clustered by who entered it and when, which is a process finding wearing a data costume.
  • “Have we got this supplier twice?” — near-duplicate discovery into the master-data stewardship queue.
FAQ

Frequently asked questions.

How is this different from machine learning?

Mining describes what is in the data you already have. Machine learning estimates something you do not yet know — a future value, an unobserved class. They use overlapping mathematics and answer completely different questions, which is why they are separate engines with separate boundaries.

Why not just say the pattern causes the outcome?

Because it usually does not, and an ERP that quietly upgrades correlation to cause will eventually cost you money. Causal claims are a separate engine and need a defensible identification strategy, not a strong correlation.

Can a mined pattern change a rule automatically?

No. It can propose one. Ratifying a rule is a human decision, recorded as such.

What stops a finding being cherry-picked?

The envelope carries the multiple-testing treatment, the comparison period and the stability of the effect. A pattern that only holds in one window says so.

See this engine on your own records.

Join the waitlist and ask it something real. Every answer names the engines it used and the records they read.

Join waitlist
No lock-in — export your data anytime.