Building a Return Reason Taxonomy from Customer Language at SKU, Size, and Colorway Level
Drop-down return reason codes produce data you cannot act on because shoppers pick the fastest available option, not the most accurate one, so the record captures behavior at the return portal, not the actual product failure.

Key highlights
- Drop-down return reason codes produce data you cannot act on because shoppers pick the fastest available option, not the most accurate one, so the record captures behavior at the return portal, not the actual product failure.
- A real return reason taxonomy maps customer language to three layers: a macro category, a functional reason, and the specific product attribute the customer names.
- The third layer, the named product attribute, is where a size exchange becomes a pattern correction or a fit note on the product page.
- The honest return reasons live in contact center conversations, where a shopper describes the problem in their own words rather than selecting from a list someone else wrote.
- You turn what a customer says into a merchandising decision at SKU level by mapping natural language to a structured return reason taxonomy and routing the resulting clusters directly to design, sourcing, and planning teams.
- Agentic AI keeps a return reason taxonomy accurate as volume grows by continuously mining conversation data, mapping customer intent to return codes, and flagging category drift before it corrupts downstream reporting.
- Return reason capture rides on the conversation record itself, because Orvera AI generates an automated summary and a transcript for every single interaction.
- A return reason taxonomy is working when the share of contacts landing in "other" or "unclassified" falls, the return rate moves at SKU and size level, and CSAT reflects that recurring product complaints are being corrected.
Why do drop-down return reason codes produce data you cannot act on?
Drop-down return reason codes produce data you cannot act on because shoppers pick the fastest available option, not the most accurate one, so the record captures behavior at the return portal, not the actual product failure.
This is the "changed my mind" trap. A shopper whose jacket sleeves are two inches too short scrolls a five-item list, sees "does not fit," and clicks it. The return is processed. The product reason, that a specific cut runs short in the sleeve, never reaches the person who could correct it. The drop-down closed the loop on logistics while leaving the merchandising problem untouched.
The cost of that ambiguity is a reorder cycle built on bad signal. A short standard list was designed for speed at the portal, not for the kind of granular return pattern analysis (opens in a new tab) that produces a correctable SKU-level finding.
What does a real return reason taxonomy look like at SKU, size, and colorway level?
A real return reason taxonomy maps customer language to three layers: a macro category, a functional reason, and the specific product attribute the customer names.
Standard return reason codes stop at layer one. A drop-down that captures "sizing issue" tells a merchant that something went wrong. It does not tell a merchant that the 32-inch inseam on style W204 in midnight navy runs two centimeters short, or that the issue disappears entirely in the charcoal colorway cut from a different fabric lot. That distinction is the difference between a taxonomy and a list.
Layer one: the macro category. Fit, quality, appearance, wrong item, and changed mind. They exist to route volume and report at the executive level. On their own, they support dashboards. They do not support a buying decision.
Layer two: the functional reason. Within fit, a customer may mean overall size, a specific measurement, or ease of movement. Within quality, they may mean durability after washing or a defect present at unboxing. The functional reason connects the macro category to the team responsible for the fix.
Layer three: the named product attribute. This is where precision matters. "Size too small" is a functional reason. "Sleeves too short" is a product attribute, and it changes the merchandising response entirely. One answer is a size exchange. The other is a pattern correction or a fit note added to the product page for that SKU. The colorway dimension adds another variable. The same cut in two colorways can fit differently when different fabric suppliers are used, and a taxonomy that does not separate them masks a supplier problem inside an aggregate size complaint.
Mapping customer language onto real product attributes requires connecting what a shopper says to what exists in a product specification. That connection cannot happen inside a drop-down. It requires listening to how customers describe the problem in their own words, then anchoring those phrases to the attributes that define each SKU. Agentic AI applied to customer conversations (opens in a new tab) is one structured path to that kind of extraction.
Where do the honest return reasons actually live if not in the returns portal?
The honest return reasons live in contact center conversations, where a shopper describes the problem in their own words rather than selecting from a list someone else wrote.
The missing piece is the source material. A shopper who clicks "other" in your returns portal will often call, chat, or email your contact center and say exactly what went wrong. "The stitching separated after one wash." "The heel sits two inches too high for my foot." Those phrases carry the specificity that a drop-down code discards.
“Contact center conversations hold this language across every channel. Voice calls, live chat transcripts, and email threads each capture the moment a customer decides to explain a defect rather than absorb it. And because these conversations happen before, during, and after a return is filed, they frequently contain detail that the returns portal never sees. A shopper who received a partial refund and moved on still described the product failure to a rep.”
The practical problem has always been volume. Manual tagging of a small conversation sample introduces the same selection bias that drop-down codes do. Automated return reason extraction (opens in a new tab) changes the equation by running intent analysis across the full conversation set a retailer already has. An agentic AI platform that operates across voice, chat, email and the other digital channels can surface the recurring phrases customers use about a specific SKU, flagging the language clusters that point to a single construction or fit issue. That output feeds directly into the taxonomy layers described above, giving your merchandising and sourcing teams something they can act on.

How do you turn what a customer says into a merchandising decision at SKU level?
You turn what a customer says into a merchandising decision at SKU level by mapping natural language to a structured return reason taxonomy and routing the resulting clusters directly to design, sourcing, and planning teams.
Early-season signals matter more than end-of-season audits. Return clusters name the problem SKU, not just the return volume, so the correction can be made on the product rather than at markdown. A buyer who sees a concentration of "seams pulling" or "color not as shown" returns on one style can act on sizing, photography, or production before the full run ships. That is a measurable cost reduction in both return volume and margin erosion.
The feedback loop closes when findings leave the contact center. Understanding how to build a return reason taxonomy is only part of the work. The taxonomy output has to reach design, sourcing, and planning in a format those teams can act on. And just as Agentic AI now routes status calls by intent and context (opens in a new tab) rather than menu selection, the same planning logic can route return intelligence to the right team automatically, without a manual reporting step in between.
How does agentic AI keep a return reason taxonomy accurate as volume grows?
Agentic AI keeps a return reason taxonomy accurate as volume grows by continuously mining conversation data, mapping customer intent to return codes, and flagging category drift before it corrupts downstream reporting.
A single-label classifier cannot manage that task. A shopper who says "the color looked completely different on my screen and the zipper broke after one wash" has given you two distinct, actionable reasons attached to the same SKU. A classifier picks one. An agentic system applies planning logic, assigns both codes, and links them to the correct attribute level. Separating a photography problem from a manufacturing defect is what lets each one get its own remediation path.
Voice of Customer analysis is the mechanism that catches what structured forms never surface. Orvera AI mines every conversation, across every channel the operation runs, to extract the themes, drivers, and sentiment behind returns. When a cluster of conversations about a single colorway shifts in tone or introduces a new descriptor your current taxonomy does not cover, Voice of Customer reporting carries that pattern.
Governance-first design is what keeps that process auditable. Every categorization decision traces back to approved knowledge.
How does return reason capture fit into a live contact center conversation?
Return reason capture rides on the conversation record itself, because Orvera AI generates an automated summary and a transcript for every single interaction.
The conventional approach adds a tagging step after the call closes. The rep selects a code from a dropdown, often under time pressure, and the accuracy of that code reflects the rep's memory and the dropdown's limited options more than the customer's actual words. That gap is where taxonomy data degrades fastest.
Agentic AI for return processing removes that gap. Agent Assist works live during the conversation, surfacing approved knowledge, recommending next-best actions, and generating an automated summary of what the customer described. Resolution stays the priority. The categorization happens in the background.
While the AI agent handles reason categorization, it simultaneously surfaces the exchange policy, the SKU's size-run history, or the next-best offer. The rep gives a faster, more accurate answer. Orvera AI's reporting carries full report logs, conversation summaries and transcripts for every single interaction.
The final piece is where that data lands. Orvera AI integrates across CCaaS, CRM, and helpdesk platforms. Merchandising teams work from those systems already. The taxonomy feeds the tools they already open every morning, which is what determines whether the insight actually changes a stocking decision.
Which KPIs show that a return reason taxonomy is working?
A return reason taxonomy is working when the share of contacts landing in "other" or "unclassified" falls, the return rate moves at SKU and size level, and CSAT reflects that recurring product complaints are being corrected.
Returns analytics without those three measures tells you volume. It does not tell you whether the taxonomy is doing its job.
Unclassified share is the most direct signal. When a meaningful portion of returns still resolve to "other," the taxonomy has gaps, and the conversation data your reps captured is not reaching the merchandising team in usable form. A well-maintained taxonomy, continuously refined by agentic AI as new language patterns surface, should drive that unclassified share toward a small residual. Any increase is a prompt to review recent conversation clusters for emerging reason types that the current codes do not yet cover.
Return rate at SKU and size level is the number the taxonomy exists to move. If a specific colorway in a size-12 athletic shoe is generating repeat returns for the same fit complaint, that pattern should be visible inside your returns analytics. Action taken on that signal, whether a fit note added to the product page or a size recommendation updated, is what makes the taxonomy commercially valuable rather than operationally decorative.
CSAT and broader CX scores provide the customer-side confirmation. Correcting a recurring product defect that the taxonomy surfaced should show up in scores tied to that product category. If CSAT holds flat after several rounds of taxonomy-informed changes, the product corrections themselves may need review. The taxonomy surfaces the signal. Your team decides what to do with it.

What are the key takeaways for building a return reason taxonomy from customer language?
Building a return reason taxonomy from customer language works when every layer of the program treats a customer-selected code as a starting point, uses agentic AI to read the language shoppers actually produce, and ties the result to SKU, size, and colorway so each return drives a merchandising decision.
Understanding why customers return products is not a reporting exercise. It is an operational one. Dropdown codes record the fastest click, not the actual reason. Treat them as a signal that something happened, not as the finding itself.
Agentic AI reads the language shoppers already produce across voice, chat, email, and every other channel the contact center runs, and maps it to a specific, structured reason without adding a survey step or changing the resolution flow. That mapping is where the real answer lives. A caller who says "the toe box crushed my foot on the left side in a size 9" is giving you a fit defect on a single colorway variant. A code labeled "size issue" would bury that signal entirely.
Taxonomy data tied to SKU, size, and colorway is what turns a return contact into a merchandising decision. When enough contacts share the same granular reason on the same product variant, the data supports a size chart correction, a supplier conversation, or a listing update. The return rate moves because the root cause is visible.
A governed, managed platform keeps categorization consistent and auditable as conversation volume grows. Without that consistency, the taxonomy drifts. Categories overlap, analysts reconcile manually, and the signal you built the program to capture gets lost in the noise.
How does Orvera AI build and run a return reason taxonomy for a retailer?
Orvera AI builds, deploys, and runs a return reason taxonomy directly on a retailer's existing contact center stack, pulling from live voice, chat, email, and every other channel the contact center runs to produce a SKU-level classification system that reflects what customers actually say.
The program is a managed service. Orvera AI, headquartered in San Francisco and carrying 18+ years of contact center operating experience, does the build, trains the contextualization models on de-identified conversation data, and runs the ongoing classification and reporting. Orvera AI runs onboarding, knowledge-base setup, agent training, and change management. Full deployment lands in three to six weeks on the stack you already have. Your QA analysts work from the results.
What comes back is a return reason structure built from your callers' and chat customers' own language, mapped to SKU, size, and colorway, and updated as return drivers shift. The "other" bucket shrinks. The categories that remain carry enough contact volume to move a merchandising decision or a product description. And when a new return pattern surfaces, it appears in the reporting as its own category rather than inside an aggregate.
If you want to see that taxonomy built from your own conversations, talk to the team at Orvera AI (opens in a new tab).
Frequently asked questions
Standard return reason codes were built for routing, not for merchandising decisions, and the gap costs product teams the insight they need most. A code like "too small" combines two distinct problems into one label. It does not separate a sizing standard that runs narrow across an entire line from a single production run where the cut shifted. A merchandiser reading that code cannot tell which problem to fix or which vendor conversation to start. The short list makes the problem worse. Representatives select the fastest option, not the most accurate one. Accuracy often loses to speed. The detail merchandising needs lives in the shopper's own words, on voice calls and chat transcripts. Unstructured data analysis for retail returns is what turns those conversations into a signal the buying team can act on, pulling the specific product attributes a shopper names rather than the bucket a representative clicked.



