Spec: smarter product-to-stream matching for ledger imports¶
Status: proposal. Nothing here is implemented. It describes two ways to
improve on the rule-based guess that ships with the ledger upload
(ledgerService.guessStream), when it should be used, and what it would
cost. Written 11 Oct 2026 at Tom's request.
The problem the rules do not solve¶
The shipped matcher decides a line's deal stream from the stream's own name. "HIB Branded" finds lines whose description or product code carries HIB. That works when the stream name contains the brand or product word that also appears on the invoice line. It fails when:
- the stream is defined by a concept the invoice never states: "Own label", "Core range", "Promotional lines", "Heavyside", "Kitchen furniture";
- the brand appears under another spelling or abbreviation: "A.Shanks", "ArmShanks", "AS" for Armitage Shanks;
- two of a supplier's streams share words: "Vado Elements" and "Vado Elements Air";
- the description is a bare product code and the brand is only knowable from the catalogue.
In those cases the rules fall back to a catch-all or the supplier's only stream, which is the point where a learned approach earns its place.
What the group already knows¶
Every group site has signal that the rules ignore:
| Source | What it gives | Where |
|---|---|---|
| Product catalogue | code, description, brand, category, supplier for tens of thousands of products | product index / PIM tables |
| Stream to category links | which product categories a deal stream covers | figuresEntry.categoryID, deal CMT categories |
| Past reclassifications | a human said "lines with code X belong to stream Y" | figuresEntryCustomerMapping, erpMapping |
| Other members' ledgers | the same supplier's product described many ways, already placed | bmnet_* edi docs with streamMatch |
| Intact ERP lines | descriptions with a known stream, for merchants on the integration | edi docs with erpSource = intact |
A learned matcher is mostly a way of applying this signal to a new line.
Two candidate approaches¶
A. Embedding similarity (recommended first)¶
Represent each candidate stream and each incoming line as a vector and pick the nearest stream above a threshold. No generation, no prompt, cheap per line, deterministic for a given model.
Stream representation. For each supplier stream, build a text from the stream name, the deal name, the names of its linked product categories, and a sample of descriptions already known to belong to it (past reclassifications, Intact lines, confirmed ledger lines). Embed once per stream; refresh nightly or when a reclassification lands.
Line representation. Description plus product code plus supplier name. Embed per distinct description (ledgers repeat descriptions heavily; a 1 GB file might hold 50k distinct descriptions out of 5M lines).
Decision. Cosine similarity to each of the supplier's live streams. Accept when the best score clears a threshold and beats the second-best by a margin; otherwise fall back to the existing rules. Thresholds are tuned per group against a held-out set of reclassified lines.
Model. A small embedding model is enough for short product strings. Options, in order of preference: a self-hosted open model (bge-small, e5-small, nomic-embed) on the existing EC2 fleet, which keeps ledger data in the VPC and costs nothing per call; or a hosted API (OpenAI text-embedding-3-small, Voyage, Cohere) at roughly a hundredth of a cent per thousand descriptions.
Storage. Vectors are small (384 to 1024 floats). Streams: a handful per supplier, held in a table. Lines: cache per distinct description per supplier in a table keyed by hash, so repeat uploads cost nothing. Elasticsearch 1.4 has no vector search; the comparison is done in CFML or a tiny sidecar, which is fine at this scale because the candidate set is one supplier's streams, not the whole catalogue.
Fit for the task. Good when a stream has example lines to learn from. Weak for a brand-new stream with only a name; there it degrades to roughly what the rules do.
B. Jev (TypeSafe System One decision API), recommended for the residue¶
Jev is not a chat model. It is a decision endpoint: you send a piece of
text as state and a set of typed questions, and it returns an answer per
question with a probability distribution and a calibrated confidence. No
free text is generated, so it cannot invent a stream that does not exist,
and the answer is always one of the options you declared. Docs:
https://docs.typesafe.ai, model jev-latest (1.13 at time of writing).
The call. POST https://api.typesafe.ai/v1/systemone with a bearer
key. One request per distinct description per supplier:
{
"state": "Supplier: HIB Bathrooms Ltd. Product code: HIB77400. Description: HIB Ambience 60 Mirror H80 x W60 x D4cm",
"model": "jev-latest",
"questions": {
"stream": {
"type": "choice",
"instructions": "Which of this supplier's deal streams does the invoice line belong to",
"criteria": {
"fe_1021": { "what": "HIB Branded: HIB's own-brand mirrors, cabinets and lighting", "examples": ["HIB Ambience 60 Mirror", "HIB Qubic 50 Cabinet"] },
"fe_1022": { "what": "All other products: anything HIB supplies that is not HIB branded", "not_for": "HIB branded items" },
"other": "None of these streams, or not a product line (carriage, credit, deposit)"
}
},
"is_product": {
"type": "noul",
"instructions": "The line is a sale of goods rather than carriage, a credit note, a deposit or an adjustment"
}
}
}
Response (abridged): answers.stream.choice = "fe_1021", confidence 0.91,
probabilities { fe_1021: 0.93, fe_1022: 0.05, other: 0.02 },
answers.is_product.noul = 0.98, plus token usage. The confidence is a
calibrated number, which is what lets us set an acceptance threshold and
mean it.
Why it fits. The candidate set is small (a supplier's live streams,
rarely more than ten; the limit is 255 options), the criteria can carry
exactly the knowledge the rules lack (category names, "not for" notes, a
few known examples), and the other option gives the model a legitimate
way to say none. The is_product question is a free extra: it flags
carriage and credit lines so they stop polluting stream totals.
Cost and speed. Pricing is input-only at $0.084 per million tokens; output is not charged. A request with ten streams and their criteria is about 400 tokens, so 50,000 distinct products cost under $2, once, because the decision is stored on the product mapping and never repeated for that product unless the model version changes. Latency is 70 to 500 ms per call; with modest parallelism a 50,000-description ledger classifies in well under an hour, and small uploads can be classified inline before the sync step.
Data handling, unverified. The public privacy page is a template and
the docs say nothing about retention, training use or region. Before
sending member ledger lines, get written answers on: whether state is
stored and for how long, whether it trains the model, where it is
processed, and whether a data processing agreement is available. Until
then, this is suitable for a pilot on a consenting group's data only.
C. General LLM classification¶
Ask a language model (Claude Haiku 4.5, or Sonnet 5 for hard taxonomies) to classify a batch of descriptions into the supplier's streams, returning JSON with a stated confidence. Same prompt content as B, batched 50 to 100 descriptions per call.
Fit. Handles everything B does and can also explain its reasoning or
propose a new stream, which B cannot. Costs roughly a hundred times more
per description, is slower, and its confidence is self-reported rather than
calibrated, so thresholds are less trustworthy and results can drift across
batches. Keep it for cases B returns other on with high confidence and a
human wants a second opinion, or drop it entirely if B's pilot is good.
Comparison¶
| Rules (shipped) | Embeddings | Jev | LLM | |
|---|---|---|---|---|
| Needs examples per stream | no | yes, improves with them | no, helps | no, helps |
| Handles "Own label", abbreviations | no | partly | yes | yes |
| Cost per 50k distinct descriptions | nil | pence (self-hosted: nil) | about $2 | $5 to $15 |
| Deterministic | yes | yes | yes for a model version | no |
| Confidence | none | cosine score, uncalibrated | calibrated probability | self-reported |
| Data leaves VPC | no | no if self-hosted | yes, terms unverified | no if Bedrock |
| Latency | none | ms | 70 to 500 ms per description | seconds per batch |
| Can say "none of these" | catch-all only | threshold only | explicit other option |
yes |
| Explainable to a user | "name contained HIB" | "similar to known lines" | probability per stream | model's reason |
Recommendation, revised. Rules first, as shipped. Jev second, for every
line the rules leave on a catch-all or unplaced, accepting the answer when
confidence clears a per-group threshold (start at 0.75) and the choice is
not other. Embeddings become optional: they only beat Jev on cost when a
group has a very large catalogue and self-hosting is already in place, and
Jev's criteria can carry the same known-line examples an embedding would
learn from. Keep the LLM as a manual "ask for a second opinion" rather than
a pipeline stage. Everything is cached so each distinct description is
classified once per model version.
Where the decision lives: the product, not the transaction¶
Whatever makes the guess, the result is stored once per product and applied to every line of that product. A ledger with five million lines has perhaps fifty thousand distinct products; the stream decision is the same for all lines of one product; and a product is something a user can find and correct in PIM. Jev therefore runs once per product rather than once per description, and a correction fans out to the lines rather than being repeated line by line.
Store the category, resolve the stream¶
Deal streams are rows of one deal, and deals renew, so a product pinned to
a specific stream row goes stale every year. Streams already reference a
category in the group's category tree (figuresEntry.categoryID, the same
contactGroup tree used across the platform). The durable fact to hold on
a product is therefore its category: "this is an HIB Branded product".
The import resolves category to whichever of the supplier's streams is
live on the invoice date, and caches that resolved stream on the mapping
with the date range it is valid for.
Where a supplier's streams have no category set, which is common today, the mapping falls back to holding the stream directly, and on renewal the resolver looks for a stream of the same name under the supplier's next deal. Setting categories on streams becomes a small housekeeping task for group administrators that pays back in durable product mappings.
The product may not exist yet¶
A merchant's PIM index (pim.<merchantSiteID>) is populated by the Intact
product import or by supplier price files. A CSV uploader's lines
reference product codes with nothing behind them, which is precisely the
population this feature targets. The import therefore creates a stub
product when it meets an unknown code: supplier, supplier product code,
description as first seen, source ledger, nothing else. Lines with no
product code at all get a pseudo-product keyed on supplier plus normalised
description. Stubs are visibly stubs in PIM and are enriched if a price
file or Intact later supplies the real product.
Storage¶
A SQL table is the source of truth, mirrored onto the product document for display and search. SQL because the recategorisation job and the audit trail need to join on it, and because Elasticsearch 1.4 cannot update by query.
| Column | Meaning |
|---|---|
| merchantSiteID | the member's merchant site |
| supplierID | the group supplier the product belongs to |
| productKey | supplier product code, or desc:<hash> for the pseudo-product |
| categoryID | the group category the product is placed in (nullable) |
| figuresEntryID | the stream resolved from the category, or set directly when no category exists |
| validFrom, validTo | the deal period the resolved stream applies to |
| source | nominal, rule, jev, user |
| confidence | 0 to 1 (Jev's calibrated confidence; 1 for user) |
| setBy, setAt | who or what last set it |
| uploadID | the upload that first created it |
Precedence when sources disagree: user always wins and is never
overwritten by a later import; nominal beats jev beats rule. A later
import with a higher-precedence source updates the row; a lower one does
not. The product document carries a copy of category, stream, source and
confidence so PIM can filter "products placed by Jev under 0.8" for review.
Correction fans out¶
A user changes a product's category in PIM. That writes the mapping with
source user, then a job: finds the lines in the merchant index with that
supplier and product key; re-indexes them with the new stream code; rewrites
their erpMapping rows; and recomputes the cached ERP totals for both the
old and the new stream. Lines are already searchable by supplier and
product code, so this is a bounded query per correction, not a re-import.
The same job runs when a group administrator sets a category on a stream
that had none, re-resolving every product mapping that was holding a stream
directly.
PIM as the insight surface¶
The product is the only join between a purchase line, which carries cost and, through its stream, the rebate rate, and a sales line, which carries price. Once both ledgers are indexed against products, PIM can show per product, and rolled up to supplier, category and stream:
- purchases in the period and the rebate earned on them, using the deal's rebate calculation on the stream's turnover;
- sales in the period, at net value;
- margin after rebate: sales net, less purchase net, plus rebate attributable to the purchases of that product.
One dependency to plan for: the sales ledger will use the merchant's own
internal product codes, not the supplier's. Intact products carry both
(id and supplierProductCode), so Intact merchants join cleanly. CSV
uploaders need a one-off internal-to-supplier code mapping, which is the
same remembered-mapping pattern the wizard already uses for columns and
suppliers, applied to products. Until a product is mapped its sales lines
sit against the pseudo-product and contribute to totals but not to
per-product margin.
How it slots into the import¶
The import already records streamMatch on every line. With the product
model above, the per-line step becomes a lookup: resolve the line's
product key, read its mapping, stamp the line. Only products with no
mapping go through the decision pipeline, once each: nominal code, then
rules, then Jev, then catch-all, each writing a mapping row with its source
and confidence (Jev's other answer leaves the product unmapped, and a low
is_product marks the product as non-product so its lines are excluded
from stream totals). Nothing about the index, the ERP totals or the sync
step changes; the sync step's "matched by product name" table becomes a
list of products, which is also what the user corrects.
Two things do change:
- Async classification. For a small residue (a few hundred distinct
descriptions) Jev can run inline between import and sync. Beyond that,
unplaced lines are indexed with
streamMatch = pending, a job classifies distinct descriptions with bounded parallelism, and the affected documents are re-indexed. The wizard's sync step shows "N lines awaiting classification" and the push channel announces completion. - A review surface. Guesses become proposals with a confidence. The sync step (or the ERP lines modal) lets the user confirm or correct per stream and description. Every correction is written back as a reclassification, which feeds the next embedding refresh and the LLM's examples. This loop is what makes the system improve; without it, A and B are only as good as their first run.
Guardrails¶
- Classification only chooses among the supplier's own live streams. It can never move a line to another supplier's deal.
- Any score below threshold is
none, not a guess. A confident wrong answer costs the member more than an unplaced line. - One decision per product, stored with its source and model version. A model upgrade re-runs unmapped and machine-mapped products as a job with a diff report, never user-set ones, and never silently.
- Confirmed human reclassifications always override machine output.
- Log the method and score on the line so the "how" column in the sync step can say "by product name", "similar to known lines" or "classified".
Open questions¶
- How many of each group's streams have a category set today? That decides how much falls back to direct stream mapping and how much housekeeping administrators face.
- Rebate attribution per product for margin after rebate: pro-rata of the stream's rebate by purchase value is simple; stepped and growth deals make the marginal rate differ from the average. Pick one and state it in PIM.
- Jev's data handling: retention of
state, training use, processing region, availability of a DPA. Nothing published answers this yet. - Jev's rate limits and whether a batch endpoint exists; the docs show one
stateper call with several questions, which is the right shape here but means one call per distinct description. - Which groups have enough reclassification history to tune thresholds? NBG almost certainly does; a new group will start on rules only.
- Should classification also run over the existing Intact
edilines withstreamCodeempty? Same machinery, large one-off benefit.
Rough sizing¶
| Piece | Effort |
|---|---|
| Stream text builder + embedding refresh job | 2 days |
| Description embedding cache + inline nearest-stream step | 2 days |
| Threshold tuning harness against reclassification history | 1 day |
| Product mapping table, stub-product creation, category-to-stream resolver | 3 days |
| Jev client, criteria builder per supplier, threshold config | 2 days |
| Correction fan-out job (re-index lines, erpMapping, totals) | 2 days |
| PIM: category field, source/confidence filters, margin-after-rebate panel | 4 days |
| Sales ledger internal-to-supplier product code mapping step | 2 days |
| Pending state, classification job with parallelism, re-index on result | 2 days |
| LLM second-opinion action (optional) | 1 day |
| Review surface on the sync step with write-back | 3 days |
| Self-hosted embedding sidecar (if not using an API) | 1 to 2 days |
About three weeks for the product model, Jev, the fan-out and the PIM surface together; embeddings and the LLM second opinion remain optional extras of a few days each.