Should You Buy a Personalization Engine Yet?
In short: Personalization engines learn from behaviour. They need many shoppers, repeat visits, and products that stay in the catalog long enough to accumulate a history. Fashion gives them none of that. About a third of a fashion catalog is new at any moment, apparel has the smallest basket of any retail category, and most shoppers buy once and leave. If your store does under roughly 50,000 sessions a month, an engine that costs $200 to $500 a month is being asked to find patterns in data you have not generated yet. This post is about how to tell where you sit, not about which vendor is best.
Someone asks me most weeks whether they should install Nosto or Rebuy. It is a fair question and the honest answer is usually "not yet, and here is how you would know." That answer is worth more than a comparison table, because the vendors are mostly fine. The mismatch is between what they need and what a young fashion brand has.
What a personalization engine actually requires
The pitch is that the software learns each shopper and shows them the right thing. Underneath, most of these systems run some blend of collaborative filtering (people who bought this also bought that) and behavioural segmentation (this visitor looks like that cohort). Both are pattern-finders. Both need volume before the patterns are real rather than noise.
Three numbers decide whether you have that volume.
Sessions per month. A recommendation model is fitting a curve. Under a few tens of thousands of sessions the curve is drawn through a handful of points, and the engine falls back to showing bestsellers, which your theme already does for free.
Repeat visit rate. Personalization means knowing one shopper across visits. If most of your traffic is first-time, there is no history to personalize against. The engine sees a stranger and guesses.
Purchase history per product. This is the one fashion breaks worst, and it deserves its own section.
Fashion catalogs go stale faster than models learn
Dressipi, which sells fashion personalization and therefore has no reason to undersell the problem, reports that the cold-start problem affects on average 33% of a fashion catalog, ranging between 20% and as high as 50% depending on the retailer (vendor-reported). A third of what you sell is new enough to have no behavioural history at all.
Think about what that means in practice. You drop an autumn collection. Those are the products you most want recommended, because they are full price and you just paid to photograph them. They are also the products the engine knows least about. By the time enough people have bought the linen shirt for the model to be confident about it, the linen shirt is on markdown.
Then there is basket size. Apparel averages about 2.78 items per transaction, the lowest of any retail category, according to Capital One Shopping. Dynamic Yield's ecommerce benchmarks put it at 2.8. Collaborative filtering learns from co-purchase, and a two-item order carries almost no co-purchase signal. Compare that to a grocery basket with forty lines in it, which is the shape of data these techniques were built on.
Small catalog, fast turnover, shallow baskets, mostly one-time shoppers. That is four independent reasons the same class of software struggles here, and none of them is the vendor's fault.
The independent research says the same thing, from a different direction
The most useful study I have found on this question is not about fashion at all. Peukert, Sen and Claussen published "The Editor and the Algorithm" in Management Science in 2024. It is a field experiment on a large German news site, comparing an automated personalized recommender against human editors deciding what to show.
On average the algorithm won, by about 2.5% on clicks. That is a real result and a small one. The interesting part is where it did not win. Human curation performed better when the algorithm lacked enough personal data on the reader, and when reader preferences varied a lot. The editor advantage held up to roughly the 35th percentile of personal data, which the authors put at around three visits per user.
Three visits per user. Ask yourself what share of your traffic has been to your site three times.
The same paper found that the best outcome came from combining the two, with the optimal mix producing up to 13% more clicks than either approach alone. So this is not an argument that algorithms are bad. It is an argument that they need feeding, and that until they are fed, a human choosing what to show is the stronger option.
There is a cost signal worth adding. In 2019 Gartner predicted that by 2025, 80% of marketers who had invested in personalization would abandon their efforts due to lack of ROI, the perils of customer data management, or both. Treat that as a comment on execution difficulty rather than on the technique. Execution difficulty is exactly what a two-person brand cannot absorb.
What the vendors cost, as of September 2026
Prices move and tiers get renamed, so check before you decide. As of September 2026:
| Public pricing | Roughly where it starts | Notes | |
|---|---|---|---|
| Rebuy | Yes, on the App Store listing | Free tier, then tiers rising toward around $500/month | Scales with order volume, so the bill grows as you do |
| LimeSpot | Yes, on the App Store listing | Low tens of dollars per month at entry | Broad feature set, not fashion-specific |
| Nosto | No, quote only | Commonly reported around $500/month | Enterprise motion, expect a demo and a contract |
| Kleep | No, quote only | Not published | Fashion-specific, but its core product is sizing and fit rather than cross-sell |
The number that matters is not the monthly fee. It is the fee divided by the incremental revenue the engine produces over what you were already doing. McKinsey's Next in Personalization report puts typical personalization lift at 10 to 15 percent revenue, with a range of 5 to 25 percent depending on sector and ability to execute. Take the low end, because a small team with sparse data is the low end. Ten percent of a small number is a smaller number, and it has to clear $500 a month before the software has paid for itself.
How to tell where you actually sit
Open your analytics and answer four questions honestly.
Do you get more than 50,000 sessions a month? Below that, most engines are pattern-matching on too little.
Is your returning visitor rate above 30%? Below that, there is not much of a "person" to personalize to.
Do more than half your products have at least 20 orders behind them? If most of your catalog is newer than that, the engine will keep falling back to bestsellers.
Is your catalog turnover under 20% a quarter? Fast fashion and frequent drops make the cold-start problem permanent rather than temporary.
If you answered yes to three or four, buy the engine and give it a proper run. If you answered yes to one or none, spend the same money on the thing that works without data: deciding, yourself, what goes with what. A person who knows the brand can style a coordinated look on the day a product launches, which is precisely when an algorithm knows nothing about it. We wrote about why the two approaches are different problems in why fashion recommendation engines don't create outfits.
Full disclosure: Angadi is our product.
Buy it later, not never
None of this is permanent. Traffic grows, repeat rates climb, and the same engine that would have guessed at 5,000 sessions gets genuinely good at 200,000. Set the trigger now so you are not relitigating it every quarter. Pick the two thresholds that matter most for your store, write them down, and revisit when you cross them.
The failure I see is not brands buying personalization too late. It is brands buying it at 3,000 sessions a month, watching it recommend bestsellers for six months, and concluding that recommendations do not work for fashion. They do. They just need something to learn from first.
If you want the adjacent question of whether your existing widget is doing anything, most complete the look widgets aren't styling anything covers a two-minute check you can run on your own product pages. And for the basket-size question underneath all of this, units per transaction is the metric to watch rather than AOV.
Frequently asked questions
Do I need a personalization app for my Shopify fashion store? Not until you have the traffic to feed it. Personalization engines learn from behavioural data, which means many sessions, repeat visitors, and products with purchase history. Under roughly 50,000 sessions a month most engines fall back to showing bestsellers, which your theme already does. Curated outfit merchandising works from day one because it does not depend on shopper data.
How much traffic do you need for personalization to work? There is no official threshold, but the useful proxy comes from the 2024 Management Science study by Peukert, Sen and Claussen, which found human curation outperformed the algorithm up to roughly the 35th percentile of personal data, about three visits per user. Practically, look for more than 50,000 monthly sessions and a returning visitor rate above 30% before expecting an engine to beat a sensible default.
Why do recommendation engines struggle with fashion? Four reasons compound. Catalog turnover is high, with Dressipi reporting that the cold-start problem affects on average 33% of a fashion catalog. Baskets are shallow, with apparel averaging about 2.78 items per transaction according to Capital One Shopping, which is the lowest of any category. Most fashion shoppers buy once. And the products you most want recommended, new full-price arrivals, are the ones with no history.
What does Nosto or Rebuy cost for a small brand? As of September 2026, Rebuy publishes tiered pricing that starts free and rises toward roughly $500 a month as order volume grows, and LimeSpot publishes entry pricing in the low tens of dollars a month. Nosto does not publish pricing and is commonly reported to start around $500 a month. Check current pricing directly, since tiers change.
Is curated styling better than algorithmic recommendations? They solve different problems, and the research suggests the combination beats either alone. The 2024 Management Science field experiment found the optimal mix of human curation and automated recommendation produced up to 13% more clicks than either approach on its own. For a small brand the practical sequence is curation first, because it works without data, then personalization once there is enough behaviour to learn from.
Should I wait to install any recommendation app at all? No. The question is which kind. Rule-based and curated approaches work at any traffic level because a person is making the decision. Behavioural personalization needs volume. Installing the second kind too early tends to produce six months of bestseller recommendations and a false conclusion that recommendations do not work in fashion.
Angadi builds complete outfits from your catalog and places them on every product page. It installs on Shopify with a 14-day free trial, and nothing goes live without your approval. See it on your store →