Skip to content

Spec: smarter product-to-stream matching for ledger imports

Status: proposal. Nothing here is implemented. It describes two ways to improve on the rule-based guess that ships with the ledger upload (ledgerService.guessStream), when it should be used, and what it would cost. Written 11 Oct 2026 at Tom's request.

The problem the rules do not solve

The shipped matcher decides a line's deal stream from the stream's own name. "HIB Branded" finds lines whose description or product code carries HIB. That works when the stream name contains the brand or product word that also appears on the invoice line. It fails when:

  • the stream is defined by a concept the invoice never states: "Own label", "Core range", "Promotional lines", "Heavyside", "Kitchen furniture";
  • the brand appears under another spelling or abbreviation: "A.Shanks", "ArmShanks", "AS" for Armitage Shanks;
  • two of a supplier's streams share words: "Vado Elements" and "Vado Elements Air";
  • the description is a bare product code and the brand is only knowable from the catalogue.

In those cases the rules fall back to a catch-all or the supplier's only stream, which is the point where a learned approach earns its place.

What the group already knows

Every group site has signal that the rules ignore:

Source What it gives Where
Product catalogue code, description, brand, category, supplier for tens of thousands of products product index / PIM tables
Stream to category links which product categories a deal stream covers figuresEntry.categoryID, deal CMT categories
Past reclassifications a human said "lines with code X belong to stream Y" figuresEntryCustomerMapping, erpMapping
Other members' ledgers the same supplier's product described many ways, already placed bmnet_* edi docs with streamMatch
Intact ERP lines descriptions with a known stream, for merchants on the integration edi docs with erpSource = intact

A learned matcher is mostly a way of applying this signal to a new line.

Two candidate approaches

Represent each candidate stream and each incoming line as a vector and pick the nearest stream above a threshold. No generation, no prompt, cheap per line, deterministic for a given model.

Stream representation. For each supplier stream, build a text from the stream name, the deal name, the names of its linked product categories, and a sample of descriptions already known to belong to it (past reclassifications, Intact lines, confirmed ledger lines). Embed once per stream; refresh nightly or when a reclassification lands.

Line representation. Description plus product code plus supplier name. Embed per distinct description (ledgers repeat descriptions heavily; a 1 GB file might hold 50k distinct descriptions out of 5M lines).

Decision. Cosine similarity to each of the supplier's live streams. Accept when the best score clears a threshold and beats the second-best by a margin; otherwise fall back to the existing rules. Thresholds are tuned per group against a held-out set of reclassified lines.

Model. A small embedding model is enough for short product strings. Options, in order of preference: a self-hosted open model (bge-small, e5-small, nomic-embed) on the existing EC2 fleet, which keeps ledger data in the VPC and costs nothing per call; or a hosted API (OpenAI text-embedding-3-small, Voyage, Cohere) at roughly a hundredth of a cent per thousand descriptions.

Storage. Vectors are small (384 to 1024 floats). Streams: a handful per supplier, held in a table. Lines: cache per distinct description per supplier in a table keyed by hash, so repeat uploads cost nothing. Elasticsearch 1.4 has no vector search; the comparison is done in CFML or a tiny sidecar, which is fine at this scale because the candidate set is one supplier's streams, not the whole catalogue.

Fit for the task. Good when a stream has example lines to learn from. Weak for a brand-new stream with only a name; there it degrades to roughly what the rules do.

Jev is not a chat model. It is a decision endpoint: you send a piece of text as state and a set of typed questions, and it returns an answer per question with a probability distribution and a calibrated confidence. No free text is generated, so it cannot invent a stream that does not exist, and the answer is always one of the options you declared. Docs: https://docs.typesafe.ai, model jev-latest (1.13 at time of writing).

The call. POST https://api.typesafe.ai/v1/systemone with a bearer key. One request per distinct description per supplier:

{
  "state": "Supplier: HIB Bathrooms Ltd. Product code: HIB77400. Description: HIB Ambience 60 Mirror H80 x W60 x D4cm",
  "model": "jev-latest",
  "questions": {
    "stream": {
      "type": "choice",
      "instructions": "Which of this supplier's deal streams does the invoice line belong to",
      "criteria": {
        "fe_1021": { "what": "HIB Branded: HIB's own-brand mirrors, cabinets and lighting", "examples": ["HIB Ambience 60 Mirror", "HIB Qubic 50 Cabinet"] },
        "fe_1022": { "what": "All other products: anything HIB supplies that is not HIB branded", "not_for": "HIB branded items" },
        "other":   "None of these streams, or not a product line (carriage, credit, deposit)"
      }
    },
    "is_product": {
      "type": "noul",
      "instructions": "The line is a sale of goods rather than carriage, a credit note, a deposit or an adjustment"
    }
  }
}

Response (abridged): answers.stream.choice = "fe_1021", confidence 0.91, probabilities { fe_1021: 0.93, fe_1022: 0.05, other: 0.02 }, answers.is_product.noul = 0.98, plus token usage. The confidence is a calibrated number, which is what lets us set an acceptance threshold and mean it.

Why it fits. The candidate set is small (a supplier's live streams, rarely more than ten; the limit is 255 options), the criteria can carry exactly the knowledge the rules lack (category names, "not for" notes, a few known examples), and the other option gives the model a legitimate way to say none. The is_product question is a free extra: it flags carriage and credit lines so they stop polluting stream totals.

Cost and speed. Pricing is input-only at $0.084 per million tokens; output is not charged. A request with ten streams and their criteria is about 400 tokens, so 50,000 distinct products cost under $2, once, because the decision is stored on the product mapping and never repeated for that product unless the model version changes. Latency is 70 to 500 ms per call; with modest parallelism a 50,000-description ledger classifies in well under an hour, and small uploads can be classified inline before the sync step.

Data handling, unverified. The public privacy page is a template and the docs say nothing about retention, training use or region. Before sending member ledger lines, get written answers on: whether state is stored and for how long, whether it trains the model, where it is processed, and whether a data processing agreement is available. Until then, this is suitable for a pilot on a consenting group's data only.

C. General LLM classification

Ask a language model (Claude Haiku 4.5, or Sonnet 5 for hard taxonomies) to classify a batch of descriptions into the supplier's streams, returning JSON with a stated confidence. Same prompt content as B, batched 50 to 100 descriptions per call.

Fit. Handles everything B does and can also explain its reasoning or propose a new stream, which B cannot. Costs roughly a hundred times more per description, is slower, and its confidence is self-reported rather than calibrated, so thresholds are less trustworthy and results can drift across batches. Keep it for cases B returns other on with high confidence and a human wants a second opinion, or drop it entirely if B's pilot is good.

Comparison

Rules (shipped) Embeddings Jev LLM
Needs examples per stream no yes, improves with them no, helps no, helps
Handles "Own label", abbreviations no partly yes yes
Cost per 50k distinct descriptions nil pence (self-hosted: nil) about $2 $5 to $15
Deterministic yes yes yes for a model version no
Confidence none cosine score, uncalibrated calibrated probability self-reported
Data leaves VPC no no if self-hosted yes, terms unverified no if Bedrock
Latency none ms 70 to 500 ms per description seconds per batch
Can say "none of these" catch-all only threshold only explicit other option yes
Explainable to a user "name contained HIB" "similar to known lines" probability per stream model's reason

Recommendation, revised. Rules first, as shipped. Jev second, for every line the rules leave on a catch-all or unplaced, accepting the answer when confidence clears a per-group threshold (start at 0.75) and the choice is not other. Embeddings become optional: they only beat Jev on cost when a group has a very large catalogue and self-hosting is already in place, and Jev's criteria can carry the same known-line examples an embedding would learn from. Keep the LLM as a manual "ask for a second opinion" rather than a pipeline stage. Everything is cached so each distinct description is classified once per model version.

Where the decision lives: the product, not the transaction

Whatever makes the guess, the result is stored once per product and applied to every line of that product. A ledger with five million lines has perhaps fifty thousand distinct products; the stream decision is the same for all lines of one product; and a product is something a user can find and correct in PIM. Jev therefore runs once per product rather than once per description, and a correction fans out to the lines rather than being repeated line by line.

Store the category, resolve the stream

Deal streams are rows of one deal, and deals renew, so a product pinned to a specific stream row goes stale every year. Streams already reference a category in the group's category tree (figuresEntry.categoryID, the same contactGroup tree used across the platform). The durable fact to hold on a product is therefore its category: "this is an HIB Branded product". The import resolves category to whichever of the supplier's streams is live on the invoice date, and caches that resolved stream on the mapping with the date range it is valid for.

Where a supplier's streams have no category set, which is common today, the mapping falls back to holding the stream directly, and on renewal the resolver looks for a stream of the same name under the supplier's next deal. Setting categories on streams becomes a small housekeeping task for group administrators that pays back in durable product mappings.

The product may not exist yet

A merchant's PIM index (pim.<merchantSiteID>) is populated by the Intact product import or by supplier price files. A CSV uploader's lines reference product codes with nothing behind them, which is precisely the population this feature targets. The import therefore creates a stub product when it meets an unknown code: supplier, supplier product code, description as first seen, source ledger, nothing else. Lines with no product code at all get a pseudo-product keyed on supplier plus normalised description. Stubs are visibly stubs in PIM and are enriched if a price file or Intact later supplies the real product.

Storage

A SQL table is the source of truth, mirrored onto the product document for display and search. SQL because the recategorisation job and the audit trail need to join on it, and because Elasticsearch 1.4 cannot update by query.

Column Meaning
merchantSiteID the member's merchant site
supplierID the group supplier the product belongs to
productKey supplier product code, or desc:<hash> for the pseudo-product
categoryID the group category the product is placed in (nullable)
figuresEntryID the stream resolved from the category, or set directly when no category exists
validFrom, validTo the deal period the resolved stream applies to
source nominal, rule, jev, user
confidence 0 to 1 (Jev's calibrated confidence; 1 for user)
setBy, setAt who or what last set it
uploadID the upload that first created it

Precedence when sources disagree: user always wins and is never overwritten by a later import; nominal beats jev beats rule. A later import with a higher-precedence source updates the row; a lower one does not. The product document carries a copy of category, stream, source and confidence so PIM can filter "products placed by Jev under 0.8" for review.

Correction fans out

A user changes a product's category in PIM. That writes the mapping with source user, then a job: finds the lines in the merchant index with that supplier and product key; re-indexes them with the new stream code; rewrites their erpMapping rows; and recomputes the cached ERP totals for both the old and the new stream. Lines are already searchable by supplier and product code, so this is a bounded query per correction, not a re-import. The same job runs when a group administrator sets a category on a stream that had none, re-resolving every product mapping that was holding a stream directly.

PIM as the insight surface

The product is the only join between a purchase line, which carries cost and, through its stream, the rebate rate, and a sales line, which carries price. Once both ledgers are indexed against products, PIM can show per product, and rolled up to supplier, category and stream:

  • purchases in the period and the rebate earned on them, using the deal's rebate calculation on the stream's turnover;
  • sales in the period, at net value;
  • margin after rebate: sales net, less purchase net, plus rebate attributable to the purchases of that product.

One dependency to plan for: the sales ledger will use the merchant's own internal product codes, not the supplier's. Intact products carry both (id and supplierProductCode), so Intact merchants join cleanly. CSV uploaders need a one-off internal-to-supplier code mapping, which is the same remembered-mapping pattern the wizard already uses for columns and suppliers, applied to products. Until a product is mapped its sales lines sit against the pseudo-product and contribute to totals but not to per-product margin.

How it slots into the import

The import already records streamMatch on every line. With the product model above, the per-line step becomes a lookup: resolve the line's product key, read its mapping, stamp the line. Only products with no mapping go through the decision pipeline, once each: nominal code, then rules, then Jev, then catch-all, each writing a mapping row with its source and confidence (Jev's other answer leaves the product unmapped, and a low is_product marks the product as non-product so its lines are excluded from stream totals). Nothing about the index, the ERP totals or the sync step changes; the sync step's "matched by product name" table becomes a list of products, which is also what the user corrects.

Two things do change:

  1. Async classification. For a small residue (a few hundred distinct descriptions) Jev can run inline between import and sync. Beyond that, unplaced lines are indexed with streamMatch = pending, a job classifies distinct descriptions with bounded parallelism, and the affected documents are re-indexed. The wizard's sync step shows "N lines awaiting classification" and the push channel announces completion.
  2. A review surface. Guesses become proposals with a confidence. The sync step (or the ERP lines modal) lets the user confirm or correct per stream and description. Every correction is written back as a reclassification, which feeds the next embedding refresh and the LLM's examples. This loop is what makes the system improve; without it, A and B are only as good as their first run.

Guardrails

  • Classification only chooses among the supplier's own live streams. It can never move a line to another supplier's deal.
  • Any score below threshold is none, not a guess. A confident wrong answer costs the member more than an unplaced line.
  • One decision per product, stored with its source and model version. A model upgrade re-runs unmapped and machine-mapped products as a job with a diff report, never user-set ones, and never silently.
  • Confirmed human reclassifications always override machine output.
  • Log the method and score on the line so the "how" column in the sync step can say "by product name", "similar to known lines" or "classified".

Open questions

  • How many of each group's streams have a category set today? That decides how much falls back to direct stream mapping and how much housekeeping administrators face.
  • Rebate attribution per product for margin after rebate: pro-rata of the stream's rebate by purchase value is simple; stepped and growth deals make the marginal rate differ from the average. Pick one and state it in PIM.
  • Jev's data handling: retention of state, training use, processing region, availability of a DPA. Nothing published answers this yet.
  • Jev's rate limits and whether a batch endpoint exists; the docs show one state per call with several questions, which is the right shape here but means one call per distinct description.
  • Which groups have enough reclassification history to tune thresholds? NBG almost certainly does; a new group will start on rules only.
  • Should classification also run over the existing Intact edi lines with streamCode empty? Same machinery, large one-off benefit.

Rough sizing

Piece Effort
Stream text builder + embedding refresh job 2 days
Description embedding cache + inline nearest-stream step 2 days
Threshold tuning harness against reclassification history 1 day
Product mapping table, stub-product creation, category-to-stream resolver 3 days
Jev client, criteria builder per supplier, threshold config 2 days
Correction fan-out job (re-index lines, erpMapping, totals) 2 days
PIM: category field, source/confidence filters, margin-after-rebate panel 4 days
Sales ledger internal-to-supplier product code mapping step 2 days
Pending state, classification job with parallelism, re-index on result 2 days
LLM second-opinion action (optional) 1 day
Review surface on the sync step with write-back 3 days
Self-hosted embedding sidecar (if not using an API) 1 to 2 days

About three weeks for the product model, Jev, the fan-out and the PIM surface together; embeddings and the LLM second opinion remain optional extras of a few days each.