Somewhere between "the design team presents three concepts" and "the product ships with one of them printed on it," something goes quietly wrong more often than anyone likes to admit. Not always - sometimes the process works and the right concept wins clearly, everyone leaves the review aligned, and six months later the numbers back up the call. But often enough that almost anyone who's sat through a packaging review has a story about the time the wrong design won, and the strange part of that story is never that nobody could tell the concepts apart. It's that the way the decision actually got made had very little to do with which design would perform best on a shelf, and a great deal to do with dynamics nobody in the room would ever describe out loud as "how we make decisions here" - even though, in practice, it is exactly how the decision got made, every single time.
This isn't really a story about bad taste, or under-skilled designers, or a lack of research budget, and it's worth being honest about that up front, because it's tempting to reach for one of those explanations when a packaging launch underperforms. The design team did competent work. The concepts on the table were, by any reasonable measure, good options. What actually failed was something upstream of the concepts themselves: a decision process that was never actually designed. It just happened, the same way it happened last time, because nobody stopped to ask whether the way five people look at three JPEGs on a screen and talk for twenty minutes is a good way to figure out which one will move product on a shelf next to a dozen competitors. It usually isn't, and the reasons why are worth understanding in detail, because the failure repeats with almost mechanical predictability once you know what to look for - and because, once you understand the mechanism, the fix turns out to be a lot less mysterious than "just have better taste."
Table of Contents¶
- The Real Cost of a Bad Packaging Call
- Why It Keeps Happening
- How This Actually Gets Fixed
- A Composite Example
- FAQ
The Real Cost of a Bad Packaging Call¶
It's tempting to treat a packaging decision as low-stakes compared to, say, a product recall or a pricing mistake - it's "just" the artwork, after all, and artwork feels like the kind of thing that can be quietly fixed later if it doesn't land. That framing understates what's actually on the line, and it understates it in a way that only becomes obvious after the fact. A packaging decision, once printed, is expensive and slow to reverse: print runs are committed in bulk because that's how the unit economics work, retail listings get built around final artwork with lead times measured in weeks, and a mid-cycle redesign means either eating the cost of existing inventory outright or running two versions in market simultaneously - which confuses exactly the customers you were trying to build recognition with in the first place, at the exact moment you need them to recognize you fastest. Get it wrong on a seasonal or limited-run product, and there's often no second chance at all. The window closes before anyone even has clean data on whether the design worked, and the next opportunity to fix it is a full year away.
There's a quieter cost too, one that never shows up in a P&L line item but is felt by everyone who sat through the process. A packaging decision that clearly got made on politics rather than merit - the loudest stakeholder's favorite, the concept nobody in the room had the standing to push back on - teaches the team something uncomfortable, whether or not anyone says it out loud: that the actual criteria for winning a creative argument in this company are seniority and persistence, not evidence. That lesson doesn't stay contained to one project. It shapes what the next round of concepts even looks like, because designers and brand managers learn, quickly and rationally, to design for the room they'll be presenting to rather than the customer standing in front of a shelf, and once that adaptation sets in, it's very hard to undo without deliberately changing how the room actually makes its decisions.
Why It Keeps Happening¶
None of the patterns below are exotic. They show up in company after company, category after category, almost regardless of how talented the design team is or how much the organization genuinely cares about getting it right - which is itself the clue that something structural, not personal, is driving the outcome.
The HiPPO Effect¶
HiPPO - the highest-paid person's opinion - is a well-worn term for a reason: it describes something that happens in almost every organization that hasn't deliberately built a process to prevent it. It's not usually malicious, and it rarely feels, from the inside, like anything other than a normal meeting. A senior stakeholder genuinely believes their instinct is good (and it might be - HiPPOs are sometimes right), and everyone else in the room has entirely rational incentives not to contradict them in front of their peers, especially when the disagreement is about something as inherently hard to argue objectively as visual taste. The result is a "discussion" that's really a ratification, dressed up as a debate, complete with genuine-sounding back-and-forth that never actually changes the outcome. Nobody voted, exactly, but everyone in the room could tell you which concept was going to win before the meeting started, and the meeting itself mostly exists to make that foregone conclusion feel collaborative.
Death by Committee¶
The opposite failure mode is just as common and does just as much damage, even though it looks nothing like the HiPPO effect on the surface. Instead of one dominant voice, you get six moderate ones, none willing to fully commit to a direction, all offering incremental notes - "can we try the logo a bit bigger," "I'm not sure about that color," "what if we tested a version with the claim moved up." Individually, every single one of these notes is reasonable. Collectively, they sand every distinctive edge off a concept until what's left is inoffensive to everyone in the room and genuinely compelling to no one standing in an aisle. Committees are very good at eliminating obviously bad options and very bad at picking a genuinely bold one, because boldness, almost by definition, always has at least one detractor - and consensus-seeking processes are built to route around detractors rather than empower someone to overrule them and actually commit.
Familiarity Bias¶
The people making the call have, by the time they're making it, looked at these concepts dozens of times - in early sketches, in revised comps, in the version with the tweaked color, in the final files sitting in the shared drive. The brand team has lived with them for weeks, sometimes months. That exposure itself changes how the concepts read, independent of anything about the designs themselves - mere repeated exposure reliably makes almost anything feel more acceptable, a well-documented effect in psychology that applies just as much to a packaging comp as it does to a song that felt strange the first time you heard it and completely normal by the tenth. The concept that struck someone as slightly odd in week one often feels entirely unremarkable by week four, not because it improved, but because the internal team quietly stopped being a fair proxy for a first-time shopper the moment they became familiar with it - and nobody on the team can feel that shift happening from the inside, which is exactly what makes it so persistent.
"I'll Know It When I See It"¶
Ask most teams what specifically they're testing for in a packaging review, and you'll get something vague - "does it feel premium," "does it pop on shelf," "does it feel like us." None of these are wrong exactly; they're all reasonable things to want. But none of them are testable either, which means every review becomes a fresh, unstructured argument about undefined criteria rather than a check against something agreed on in advance. Without a specific brief - what should this packaging communicate, to whom, and how would we know if it succeeded - "I'll know it when I see it" is doing all the actual work of the decision, and it's doing that work inconsistently from one reviewer to the next, since everyone's private, unstated definition of "premium" is at least slightly different from everyone else's.
Sunk Cost on a Favorite Concept¶
Somewhere in most packaging projects, someone - often someone senior, often someone who commissioned the work in the first place or championed a particular direction to the agency early on - falls for a specific concept and stays attached to it regardless of what later feedback suggests. Walking that attachment back gets harder the more time, money, and internal capital has already gone into defending it in prior meetings, which means the attachment tends to get stronger, not weaker, as the process drags on and more has been invested in it. By the final review, the question quietly being decided often isn't "which concept is strongest" but "does anyone in this room have the standing to tell this person their early favorite isn't working" - and usually, by that point, nobody does, because raising it now would mean unwinding weeks of momentum that everyone has quietly agreed not to disturb.
Surface-Level Feedback¶
Even when a team does gather outside opinions - a hallway poll, a quick customer email blast, a handful of Slack reactions from people outside the project - the question asked is often just "which one do you like" or "rate this 1 to 5." That produces a number or a preference, but not a reason, and a packaging decision made on an unexplained preference is barely more defensible than one made on a senior stakeholder's gut feel; it just comes with a thin coat of "we asked people" painted over the top of it. The teams that get burned worst by this are often the ones who did technically gather feedback and took real, genuine comfort from having done so, and still shipped the wrong concept, because the feedback never told them why people preferred what they preferred - so when the launch underperforms, there's no diagnostic trail left to learn from, only a vague sense that the testing "should have caught it."
Timeline Pressure¶
Packaging timelines are almost always tighter at the decision stage than anywhere else in the process, and this compounds every failure mode above rather than sitting apart from them. By the time concepts are ready for review, there's usually a print deadline breathing down everyone's neck, which means the review itself gets compressed into whatever slot is left on an already-full calendar. Rushed decisions default to whichever option requires the least debate to agree on, which is not remotely the same thing as the strongest option on the table - it's simply the path of least resistance under time pressure, and it gets dressed up afterward, in the retelling, as a decision everyone was genuinely aligned on all along.
How This Actually Gets Fixed¶
None of what follows requires a particular tool, a large research budget, or an outside agency's help. These are general principles - most of them decades old in the research methodology and decision-science literature - that address each failure mode above directly, and they tend to work best applied together rather than picked one at a time.
Get Outside Reaction Before Internal Familiarity Sets In¶
Since familiarity bias is purely a function of exposure time, the fix is structural rather than a matter of individual willpower: gather reaction from people seeing the concepts for the first time, and do it early, before the internal team's judgment has had weeks to drift away from a genuinely fresh first impression. This doesn't mean ignoring internal expertise, and it isn't an argument for outsourcing the decision to strangers - it means sequencing internal judgment correctly, so it gets applied to real outside signal instead of operating in a closed loop that never gets checked against anyone unfamiliar with the project. In practice, this usually means building the outside check into the calendar from the start of the project, not bolting it on defensively once someone in leadership starts to have doubts about the emerging favorite.
Separate Diagnosis From Decision¶
Bad packaging reviews conflate two genuinely different questions into one messy, time-pressured conversation: "what's working and not working about each concept" and "which concept do we actually ship." Splitting them into two distinct steps keeps each one honest. First comes structured, element-specific feedback on each concept individually - is the logo reading as premium, is the claim legible, does the color fit the category - gathered without anyone in the room needing to defend a final position yet. Only after that diagnostic pass is complete does the group move to a separate, focused decision step, informed by what the diagnosis actually surfaced rather than by whoever spoke first in a single blended discussion.
Use Forced Comparison, Not Independent Ratings¶
This is one of the oldest and most consistently replicated findings in psychometrics, and it's worth understanding why it holds up so reliably: asking people to rate options independently on a scale produces compressed, noisy data, because everyone brings a slightly different internal reference point for what a "4" means, and that reference point drifts across the several judgments a single reviewer makes in one sitting. Asking people to choose between pairs of options, by contrast, produces a much cleaner and more decisive signal, because a relative judgment - "which of these two" - doesn't require the respondent to hold an abstract scale steady in their head at all; it only requires them to notice which one they'd actually pick. The method, variously called paired comparison or forced-choice judgment, was formalized by the psychologist Louis Thurstone in the 1920s and still underpins market-research techniques like MaxDiff today, precisely because the underlying mechanism hasn't gotten any less true in the intervening century: a sequence of relative judgments composes into a reliable ranking far more consistently than a batch of independent ratings ever does.
Write the Brief Down Before the Review¶
A single written sentence - what this packaging needs to communicate, to whom, and what would count as success - turns "I'll know it when I see it" into something a disagreement can actually be resolved against, rather than something that gets invented retroactively to justify whichever concept ends up winning anyway. It doesn't need to be elaborate, and it doesn't need sign-off from a dozen stakeholders to exist. It needs to exist before the review starts, committed to on paper, so that when two people disagree about a concept in the room, there's a third thing in the conversation besides their two competing opinions.
Give One Person Clear Ownership of the Final Call¶
Broad input and a single accountable decision-maker aren't in conflict with each other - in practice, they work best together, and the mistake most teams make is assuming they have to choose one or the other. The failure mode isn't having many voices in the room; it's having many voices and no clearly named owner of the final synthesis, which is exactly what produces both HiPPO capture (one voice quietly becomes the de facto owner without anyone ever agreeing that should be the case) and death by committee (no voice is empowered to actually decide, so the group drifts toward whatever nobody objects to). Naming the decision-maker explicitly, in advance, while still deliberately gathering wide input from the rest of the team, heads off both failure modes at the same time rather than trading one for the other.
Time-Box the Decision, Don't Let the Deadline Make It For You¶
Set the internal decision deadline several days - ideally a full week or more - before the actual print deadline, not at it. This sounds like a scheduling detail, but it changes the character of the decision itself: a decision made with real time still on the clock can afford a genuine comparison process, a diagnostic round, a moment to sit with the result before committing. A decision made in the shadow of a hard deadline defaults, almost automatically, to whichever option is fastest to agree on - which is a description of convenience, not of quality, no matter how confidently it gets framed afterward as the team's considered choice.
How These Six Fit Together¶
None of the six principles above are alternatives - a menu a team picks one item from and skips the rest. They sit at different points in the process, which means the real question isn't "which one should we use" but "when does each one kick in," and the honest answer is that they're meant to run together, in sequence, as one connected process rather than six independent tips.
Two of them are settled before anyone even looks at a concept: writing the brief, and naming who owns the final call. Get these in place first, and several of the failure modes above never get the room to start - there's already a shared standard and a clear owner before the first opinion gets voiced.
Two of them shape how outside reaction actually gets collected, once concepts exist to react to: sourcing that reaction from people seeing them for the first time, and asking for it through forced comparison rather than a rating scale, so what comes back is a decisive signal instead of a handful of similar-looking scores.
One governs how the review itself unfolds once opinions are in hand: diagnosing each concept element by element first, and only moving to the comparative decision once that diagnostic pass is complete, rather than trying to do both in the same breath.
And one runs underneath the whole thing rather than sitting at any single point in it: keeping the internal deadline well ahead of the print deadline, so the first three groups of principles actually have the room they need to happen properly, instead of getting compressed into whatever's left once the calendar has already made most of the decision.
| When | What happens |
|---|---|
| Before anyone sees a concept | Write the brief. Name the decision-maker. |
| While gathering reaction | Source it from fresh eyes. Collect it through forced comparison. |
| Once opinions are in | Diagnose each concept element by element, then make the comparative decision. |
| Running underneath all of it | Keep the internal deadline well ahead of the print deadline. |
A Composite Example¶
How It Actually Played Out¶
Picture a mid-sized food brand relaunching a product line with three packaging directions in front of the team two weeks before a print deadline. The VP of Marketing has quietly favored one direction since the first internal review, though nobody's said so out loud - not even the VP, who would probably describe themselves as genuinely open-minded going into the final meeting. The creative review runs forty-five minutes: the VP asks a few clarifying questions, two brand managers offer supportive comments that build on whatever the VP just said, and a third hesitates for a moment on one point before agreeing rather than pushing back with the deadline looming and the room's energy already settled. The favored concept is approved, and everyone leaves feeling like the discussion was thorough.
Six weeks after launch, sell-through on the new packaging is flat versus the previous design, and nobody can say with any real confidence why. There was never a specific hypothesis about what the new design was supposed to improve - more premium perception, clearer flavor communication, better shelf standout against a specific competitor - so there's nothing concrete to check the actual result against now, only a general sense that something didn't work as hoped. The retrospective conversation, inevitably, starts with "maybe we should get outside opinions earlier next time," which was exactly as true before the first review as it is now. It was just much harder to act on once forty-five minutes and a looming deadline had already made the call for everyone in the room.
The Same Decision, Run Differently¶
Run the same relaunch through the six principles instead, starting from the same point: three packaging directions, the same VP with the same early lean toward one of them, the same two-week window before print. The difference starts with the calendar - the internal decision deadline gets set for one week out, not two days out, and a single written sentence exists before the first review: the new packaging needs to read as more premium than the outgoing design without losing the loyal buyers who already recognize the current look. The head of brand, not the VP alone, is named as the person who owns the final call.
Reaction gets gathered from people outside the company, seeing all three concepts for the first time, and it's collected in two clean stages rather than one blended conversation. First, structured, element-specific feedback on each of the three concepts individually - the logo, the flavor callout, the color block - with nobody defending a final position yet. That diagnostic pass surfaces something the internal team, deep in familiarity with all three options, had stopped noticing: the VP's favored concept has a flavor callout that a meaningful share of first-time viewers read as ambiguous, unclear which variant of the product it's even for. Only after that diagnostic round is complete does a forced comparison run across all three concepts to settle which one wins outright - and it isn't the VP's original favorite, though it's close, and the reason it isn't is now sitting in the diagnostic data rather than being anyone's guess.
The team fixes the callout on the winning concept before it goes to print, and this time the launch has an actual hypothesis behind it: clearer flavor communication should show up in fewer product-mix-up returns and steadier repeat purchase, something concrete to check the results against in six weeks instead of a general hope that "it feels right." Whether or not sell-through beats the old design, the team will know specifically why - which is the one thing the original version of this story never had.
FAQ¶
Is a bad packaging decision usually the designer's fault?
Rarely. The failure modes described here - HiPPO effect, committee dilution, familiarity bias, sunk cost, vague criteria - live in the decision process, not in the quality of the concepts being decided between. Strong concepts lose to weak ones inside a broken process all the time, and a team can easily walk away blaming the design work for what was actually a process failure.
How many stakeholders should be involved in a packaging decision?
There's no universal number, but the more people with a vote and no clear tiebreaker, the more the process tends toward consensus-driven dilution rather than a confident, defensible choice. A smaller group with a clearly named decision-maker, informed by structured outside input, generally outperforms a large committee trying to reach agreement entirely on its own.
Should internal opinion be ignored entirely?
No - internal teams often carry real strategic context an outside respondent simply doesn't have, and discarding that would be its own mistake. The problem isn't that internal opinion exists; it's when internal opinion is the only input, gathered without structure, from people too close to the project to react the way a first-time shopper actually would.
What is paired comparison or forced-choice testing, in plain terms?
Showing people two options at a time and asking for a direct pick, rather than asking them to rate several options independently. It's a long-established method in psychometrics precisely because relative judgments are more reliable and decisive than absolute ones - it's much harder to sit on the fence when the only real options in front of you are "this one" or "that one."
What's the biggest single fix a team can make without changing tools or budget?
Writing down, before the review, specifically what the packaging needs to communicate and to whom - even a single sentence turns "I'll know it when I see it" into something an actual disagreement can be resolved against, rather than something invented after the fact to justify whatever won.