When should a return rate trigger a customer review?
A single number can focus your whole returns operation: the return rate at which a customer gets a human review. Set it too high and serial returners drain margin for months before anyone looks. Set it too low and your CX team drowns in false positives while good customers feel surveilled. The right threshold is not a guess. It comes from your own data.
Start with the distribution, not an industry average. Pull every customer's keep rate over the last two quarters: orders placed against units kept. In most apparel catalogs the shape is the same. A large cluster keeps most of what they buy, a long tail returns most of it, and a thin middle sits around half. The review threshold belongs at the point where the tail separates from the middle, not at some round number borrowed from a blog post.
Rate alone is a blunt instrument, so pair it with volume. A customer who placed two orders and returned one sits at 50 percent but means nothing. A customer who placed forty orders and kept six is a pattern. Require a minimum order count before the threshold can fire, typically five to ten orders in the window, so the percentage is measuring behavior and not luck.
Time windows matter as much as levels. A rolling 90-day window catches customers whose behavior is getting worse right now, which is exactly who you want to talk to first. Annual averages hide the customer who was fine for a year and started returning everything in August. Run both: the 90-day window for action, the annual view for context.
Category skew should move the line. Occasionwear and premium denim run hotter return rates than basics for legitimate reasons, so a flat threshold across the catalog punishes the categories your best customers shop. Segment thresholds by product category, and review customers against the threshold for the category they actually buy. Someone returning 70 percent of evening dresses is a different case from someone returning 70 percent of socks.
The trickiest part is bracketing. Multi-size orders returned in one batch inflate a customer's return rate without any bad intent, and penalizing them trains customers to stop buying your full size range. Strip bracketing batches out of the rate before the threshold check: if three sizes of the same SKU go out and two come back together, count it as a fitting event, not three returns. What remains after that adjustment is the number worth reviewing.
When a customer crosses the threshold, the review itself should be quick and structured. Look at three things: the category mix, the timing pattern, and the condition of returned items. Wear-and-return leaves marks that show up in inspection notes. Wardrobing clusters around events. Genuine fit problems spread across sizes and categories. A five-minute structured review separates these far better than an hour of scrolling order history.
Document the decision. Every review should end in one of four outcomes: clear, warn, restrict, or deny, with a note on the evidence. This record is what protects you if a customer complains, and it is what makes the program improve: after a quarter you can see which thresholds produced warnings that stuck and which produced noise, and move the line accordingly.
Thresholds are a starting point, not a policy. The brands that keep returns fraud down review the line itself every quarter, because catalogs change, customer mix changes, and the people gaming the system adapt. A threshold that worked in January can be stale by October. Treat it as a living number and it keeps working.