Skip to main content
Jabal Al Noor Pharmacy Digital Commerce Platform
Case StudyAI / ML

Teaching a machine what a medicine is

Classifying ~19,800 pharmacy SKUs that arrive as seven fields of free text, after embeddings alone failed for a structural reason.

Client
Jabal Al Noor Pharmacy
Industry
Pharmacy retail
Outcome
20 of 22 correct in ~2.2s

The problem

The catalogue lived in a billing and POS system whose sync feed carries exactly seven fields per item:

  • itemId, description, unit, qty, cost, salePrice
  • an optional base64 image

No barcode. No category. No prescription flag. So roughly 19,800 products arrive as free text like PANADOL EXTRA 500MG TAB, and somebody has to decide where each one belongs before it can go on sale. That is weeks of pharmacist time on work that is both tedious and consequential.

Why embeddings alone failed

The obvious answer is semantic similarity: embed the product description, embed each category name, take the nearest. It does not work here, and the reason is structural rather than a tuning problem.

The taxonomy is deliberately shopper-facing. Its buckets are things like "Pain & Fever" and "Allergy & Antihistamines" — the categories a customer browses, not the pharmacological classes a pharmacist thinks in. Drug classes live on products as tags instead.

So the category names contain no brand names and no active ingredients. There is no lexical or semantic bridge from PANADOL EXTRA to Pain & Fever. Cosine similarity has nothing to grip.

Then the economics went wrong

The first implementation embedded the description plus every category name, per row. With around 220 categories, a 200-item batch requested roughly 44,000 embeddings to classify 200 products — enough to exhaust a hosted free tier inside a single sync.

What we built

A three-layer suggester chain where each layer degrades into the next rather than throwing.

  1. LLM (Gemini) — asks what the product is, batched, at temperature: 0, with a 0.6 confidence floor below which it abstains rather than guesses.
  2. Embeddings — cosine similarity, now with a category-vector cache keyed by a content signature, so a rename or an addition invalidates it but a timer never does. The same batch now costs ~220 embeddings once, plus one per description.
  3. Keyword matching — local, no network, always available.

The interface contract is the important part. suggest() must not throw on a remote failure, and returning { categoryId: null } is a valid answer, because suggestion is advisory and an admin confirms every row. A suggester is never load-bearing.

Result

20 of 22 correct in about 2.2 seconds on the measured benchmark. Both misses were the model choosing a parent category where no child genuinely fitted, which is arguably the right answer rather than a miss.

What we deliberately did not do

Google Search grounding was evaluated and rejected on four grounds, recorded in the code so the decision survives staff turnover.

ObjectionDetail
Paid-tier onlyOn a free key every request returns 429, which this chain absorbs into a silent fallback
Cost~$35 per 1,000 grounded queries against ~$0.25 to classify the whole catalogue
Reintroduces driftSearch results change weekly, so the same row re-run twice can land differently
No measured gainThe benchmark showed no gap for it to close

The drift objection is the one that mattered most: temperature: 0 was set precisely so that an admin re-running "uncategorised only" sees the queue settle rather than churn.