Option 2 — Not Optimized sends your question straight to the model together with our entire product catalog, every single time, and waits for one reply:
Option 1 — Token Friendly has the backend talk to the model more than once behind the scenes: first a cheap call figures out what you're actually asking about, then a further call (or calls) answers using only that relevant slice of the catalog: