Surviving the quota
A degradation path where the failure mode is worse than an outage, because a fallback hides the evidence of itself.
- Client
- Jabal Al Noor Pharmacy
- Industry
- Pharmacy retail
- Outcome
- ~750 rows/min, restart-safe
The problem
Any hosted AI dependency will refuse you eventually, and on a free tier that happens routinely.
The naive failure mode is worse than an outage. A keyword fallback still sets a category, so the row is no longer uncategorised — which is precisely the filter an admin would reach for to fix it. The degradation hides the evidence of itself.
What we built
When a layer fails for a reason a later attempt could fix — a rate limit, an exhausted daily quota, a provider outage — it still returns a usable answer from the layer below, but tags it with deferUntilMs, taken from the provider own retry hint where one is given. The suggester never throws; the processor decides whether to bank the weaker answer or wait.
Four-step escalation
| Step | Delay | What it targets |
|---|---|---|
| 1 | 2 minutes | A per-minute rate limit, so a run barely stutters |
| 2 | 15 minutes | A short provider outage |
| 3 | 2 hours | A sustained outage |
| 4 | 6 hours | A daily quota reset |
Delayed jobs live in Redis, so the schedule survives a worker restart or a redeploy. Deliberately not a cron job, because the API runs PM2 cluster mode where an @Cron() decorator fires once per worker.
After the last step the batch stops waiting and banks the fallback. A degraded re-run scope then finds those rows later by selecting on the recorded strategy rather than on "uncategorised" — because that scope would skip them silently.
Making it visible
The admin panel reports which strategy answered each row, in four buckets that always sum to the classified total. Two bugs in that shape were worth fixing:
- The dashboard punished success. The
llmstrategy used to fall intonone, which the panel folded into "By keyword" and coloured as a warning. So the better the AI performed, the more alarming the dashboard looked. - A silent fallback was silent to operators too. The embedding layer degrades quietly by design, but that meant a misconfigured key or a retired model endpoint looked identical to normal operation: suggestions stayed keyword-based forever with nothing in the logs. It now warns once per failure streak and resets on success, rather than emitting one line per row and burying the signal in a 30,000-row sync.
Throughput
Bounded deliberately: 25 rows per job, 30 jobs per minute, one worker — about 750 rows per minute, tunable by environment variable without a deploy.