Direct answer: Use Good-Better-Best pricing only when evidence supports three distinct value bands. Start with customer segmentation and a tier hypothesis, test tier count and feature trade-offs with conjoint or choice research, pressure-test price points with Van Westendorp or Gabor Granger, and verify the final structure through a controlled rollout.

Decision map showing when a product team should research, build, test, or launch a three-tier pricing menu.
Use the map to decide whether the next step is tier-count research, menu design, price testing, or controlled launch.

The decision is conditional: do not build a three-tier menu until research demonstrates distinct customer groups or measurable value differences. A three-option menu can be a useful structure, but the existence of an entry, middle, and premium edition is not evidence that customers experience three coherent offers. [Harvard Business Review's framework] describes the central distinction: tier design should reflect differences in customer value rather than force a standard menu shape.

Four-stage process map from customer segmentation through tier mapping, fencing, and research validation.
Follow the sequence to keep tier construction separate from price-point selection.

When Three Tiers Are Justified

Four comparison panels matching four pricing research methods to the questions each method can answer.
Choose the panel that matches the decision still uncertain, and do not treat modeled or stated results as observed behavior.

Good-Better-Best pricing presents three offers arranged around increasing value, capability, or fit. The construction decision is not simply whether three choices look orderly on a pricing page. It is whether the product team can explain who each tier is for, what outcome or value driver separates it from the next tier, and why the difference should matter to purchase decisions.

Limitation sequence showing research quality checks, interpretation safeguards, launch monitoring, ethical review, and the option to return to two tiers.
Use these checks as release gates before committing engineering and go-to-market resources to three tiers.

A defensible three-tier structure therefore needs three distinct value bands or customer groups. Those groups may differ in usage, needs, budget, or other value drivers. The relevant question is not whether customers can be described as “basic,” “standard,” and “advanced.” It is whether the differences are measurable enough to guide package design and research.

This is where [Goldilocks pricing] can become a misleading shortcut. A middle option may be useful, but a weak middle tier can create an arbitrary fence between otherwise similar offers. A premium tier can have the same problem if it adds features that do not represent a meaningful outcome for a distinct audience. If the evidence supports fewer groups, forcing three tiers creates a construction problem rather than solving one.

The same rule applies to the price spread. A higher price should correspond to a difference in features, usage, support, or outcome that customers can evaluate. A ratio between prices is a hypothesis to test, not a universal rule. [Stripe's guidance on tiered pricing] supports mapping the offer to willingness to pay instead of applying a fixed multiplier.

Method: Build the Tier Hypothesis Before Pricing It

Method: Separate tier construction from price-point selection. First, segment customers using the differences relevant to the product decision: usage, needs, budget, or value drivers. Then map each proposed tier to a specific audience or outcome. The team should be able to complete a sentence such as: “This tier is for customers who need this outcome and receive this measurable value.”

Next, choose the fences that make the differences concrete. Possible fences include seats, support, API access, or usage limits. A fence is not automatically persuasive because it is easy to implement. It should communicate a difference that can be connected to an audience or outcome. The package map should also make clear what is shared across tiers and what changes between them.

At this stage, prices remain hypotheses. Draft two-tier and three-tier alternatives before selecting a final menu. Compare how each alternative maps features and value to willingness to pay. Labels should communicate audience or outcome rather than generic status alone. This keeps the research focused on the decision the team actually faces: which package differences are coherent, and whether the proposed tier count reflects those differences.

The process is: segment by relevant customer differences, map tiers to audiences or outcomes, select defensible fences, and test tier count, feature trade-offs, and candidate prices. [BCG framework for pricing strategy] is useful strategic context because pricing basis, offer structure, and pricing mechanism should be coordinated rather than treated as isolated choices.

Choose the Research Method for the Question

Different methods answer different uncertainties. A perception study should not be asked to prove observed conversion, and a choice experiment should not be treated as a live-market result.

  • Van Westendorp: Uses four price-perception questions to estimate an acceptable price range. It is appropriate when the team needs to understand perceived price boundaries before selecting candidate points.
  • Gabor Granger: Tests purchase intent at specified price points. It is appropriate when the menu is defined enough to compare stated intent across selected prices.
  • Conjoint or choice research: Evaluates trade-offs among bundled features and price. This makes it relevant to testing proposed tier structures, package differences, and alternative menu designs. [Controlled selection experiments] can help the team examine whether the proposed fences create coherent choices.
  • Field experiments: Test the structure in market conditions and provide observed behavior such as conversion, offer mix, upgrades, and churn.

These outputs must remain separate. Stated purchase intent is what respondents say they may do. Modeled or simulated share and revenue are outputs of a research model. Observed conversion, upgrades, churn, and lifetime value are field outcomes. One category cannot be relabeled as another.

A team can use [pricing research and benchmarks] to organize the method choice, but the unresolved question should determine the method. If the uncertainty is tier count or package design, conjoint or another choice experiment is the strongest perception-based option in this framework. If the uncertainty is an acceptable range, use Van Westendorp. If the uncertainty is response to specified price points, use Gabor Granger. If the uncertainty is actual behavior, use a controlled rollout.

Interpretation: Turn Evidence Into a Tier Menu

Interpretation: Start by asking whether the research shows clusters or value differences that support three offers. Do not infer a customer group merely because respondents select a particular package in a survey. Examine whether the proposed features and outcomes explain the selection and whether the differences remain coherent across relevant segments.

For choice research, use utilities or modeled share to compare package designs. These are modeled outputs, not observed demand. For price research, use the acceptable range or purchase-intent responses to pressure-test candidate prices. These are perception or stated-intent results, not a guarantee of conversion.

Treat every tier-price ratio as a testable hypothesis. The important interpretation is the relationship between the price, the fence, and the value driver. A higher tier may need more support, greater usage capacity, API access, or another measurable difference. The evidence should show whether the proposed spread helps communicate value or instead makes one tier look difficult to justify.

The live check is a controlled rollout. Monitor conversion, average order value, upgrades, churn, and lifetime value as separate metrics. The rollout does not validate only the price. It tests the combined effect of tier count, package structure, fences, labels, and price points.

Validate the Menu Before a Full Launch

Use a staged sequence:

  1. Run a perception or choice study to evaluate the tier hypothesis and package differences.
  2. Test candidate price points with Van Westendorp or Gabor Granger when the price uncertainty is narrower.
  3. Use a narrow A/B test or limited rollout to observe the menu in market conditions.
  4. Monitor conversion, mix, average order value, upgrades, churn, and lifetime value before a full rollout.

Set release gates before the test begins. Revisit the menu if the fences appear weak, the cheapest tier attracts excessive selection without supporting the intended value structure, or customers show confusion and downgrade requests. Evidence that does not support three coherent offers is a reason to return to a two-tier alternative, not a reason to add more explanation to a weak menu.

This sequence also prevents the team from engineering the full pricing architecture before it knows whether the architecture is defensible. [Stripe's guidance on tiered pricing] can frame the offer structure, while the research and rollout determine whether the proposed configuration is suitable for this product and audience.

Limits: What Tier Research Cannot Prove

Limits: Tier research cannot establish a universal optimal tier count, a universal price multiplier, a conversion lift, a revenue result, or a customer segment distribution. The conclusions depend on the prices, attributes, sample, and segments included in the study.

Stated intent is not observed behavior. Modeled share or revenue is not a field result. A field test may also lack enough traffic for a clean comparison. Choice experiments can compare proposed selections, but selection alone cannot determine how many editions the team should build. That remains a product and pricing decision informed by the evidence.

Before launch, review disclosure of all-in prices, cancellation and renewal terms, fee transparency, and personalization practices. Also review safeguards related to protected characteristics. These checks belong alongside research quality, segment quality, and interpretation quality.

Next Step: Run the Smallest Defensible Study

Next step: Write the one-sentence evidence test first: identify the distinct customer groups and the measurable value difference each group receives. If the statement is not defensible, research tier count or price perception before drafting a three-tier menu. If it is defensible, test two-tier and three-tier alternatives, compare package and price trade-offs, and define rollout guardrails.

Choose the smallest study that matches the uncertainty. Use Van Westendorp for an acceptable price range, Gabor Granger for specified price points, and conjoint or choice research for package and tier structure. Then plan a narrow rollout with predefined checks for conversion, average order value, upgrades, churn, and lifetime value.

Kinetic Pricing provides a Price range finder priced at $149.00 one time, a Price point tester priced at $199.00 one time, a Feature value ranker priced at $279.00 one time, and a Package and price builder priced at $499.00 one time. [Kinetic's plans] are available for teams evaluating which research step to run next.

Additional context on these methods is available from controlled selection experiments, Good better best, Pricing strategy and pricing decisions | BCG and Tiered pricing 101 | Stripe.

Run this method with your users

Start 30-day free trial to run every method with your users, or Buy one study when one decision needs evidence now.

Sources

Harvard Business Review, “The Good-Better-Best Approach to Pricing

Stripe, “Tiered Pricing 101: A Guide for a Strategic Approach”

Boston Consulting Group, “Pricing Strategy”

Software Pricing, “Good-Better-Best Pricing”

Kinetic Pricing, “Research”

Kinetic Pricing, product and pricing information