Match the method to the decision and the evidence available: use a controlled A/B experiment for observed purchase behavior when live traffic is available, Van Westendorp for an early acceptable price range, Gabor-Granger for a single-product demand curve, and conjoint for feature, package, and price trade-offs. Survey methods help form and narrow hypotheses before live validation. They should not be presented as equivalent to observed transactions.

Decision map routing pricing teams from their available evidence and business question to A/B testing, Van Westendorp, Gabor-Granger, or conjoint analysis.
Start with the evidence available and the decision required, then follow the branch to the method whose output matches that decision.

Start With the Evidence You Can Collect

Comparison panels showing the evidence type, primary output, and best use case for six price testing methods.
Compare methods by the evidence they collect and the decision their output can support, not by speed alone.

Price testing is not one universal procedure. It is a way to collect evidence for a defined pricing decision. The first distinction is between what people say they may do and what they do in a live purchase context.

Annotated progression from an acceptable price range and stated demand curve to observed revenue evidence from a live experiment.
Read the progression as increasing behavioral validation, while keeping survey intent and observed transactions as separate evidence types.

Survey responses represent stated purchase intent. They can support modeled or simulated revenue estimates, but those estimates are not observed revenue. A controlled live experiment provides observed behavioral outcomes. This distinction matters when a pricing team is deciding whether to change a price, redesign tiers, or compare offers. Survey responses and live experiments produce different evidence types.

Sequential preflight checks for hypothesis, metrics, sample size, price persistence, stopping rules, and delayed business effects.
Use the sequence as a release gate so an apparently positive result is not treated as a decision before its design and downstream effects are checked.

For a live-traffic decision, evaluate revenue per user alongside conversion rather than conversion alone. A lower price that lifts conversion can still shrink total revenue per user. Conversion is therefore a supporting measure when the decision concerns revenue impact. Retention, contribution margin, and lifetime value can also be supporting measures when the pricing change may affect outcomes beyond the initial purchase. Pricing experiments should connect conversion to revenue per user.

Method

Use this process:

  1. Define the pricing decision and the outcome that would determine success.
  2. Classify the available evidence as live behavioral data or survey responses.
  3. Select the method whose output matches the decision: observed revenue impact, an acceptable range, a demand curve, or feature-level trade-offs.
  4. Predefine revenue per user as the primary metric where applicable, with conversion, contribution margin, retention, and lifetime value as supporting measures.
  5. Establish the population, randomization, sample-size or power requirement, price-assignment integrity, and stopping rule before collection.
  6. Analyze point estimates with uncertainty and separate stated intent, modeled or simulated revenue, and observed behavior.
  7. Validate a survey-led price hypothesis with a controlled experiment when sufficient traffic becomes available.

Compare the Core Price Testing Methods

A/B experiments are the fit when the decision requires observed purchase behavior and live traffic can be assigned to price conditions. The result can be evaluated through observed revenue per user and conversion, with downstream measures considered where relevant. The design depends on controlled assignment and reliable delivery of the assigned price.

Van Westendorp is suited to an early-stage question: what price range may be acceptable? It uses four price-perception questions to estimate an acceptable price range and an optimal price point from intersecting curves. This makes it useful for forming a candidate range before a live experiment. Van Westendorp uses four price-perception questions.

Gabor-Granger is suited to comparing price points for one product or subscription. It asks purchase-intent questions across price points and produces a demand curve that can support price-point comparison. Its output is stated intent across prices, not observed transaction behavior. Gabor-Granger produces a purchase-intent demand curve.

Conjoint analysis is appropriate when the decision involves several attributes at once. Respondents choose among bundles that combine features, packages, and price, allowing the analysis to estimate feature and package trade-offs. Use it when the question is not simply “Which price?” but “Which combination of offer elements and price should we test?” Conjoint estimates trade-offs among bundles.

Monadic testing shows each respondent one price rather than asking for a direct comparison. This reduces direct comparison and anchoring effects, but it requires substantial sample per price point. It is a survey design choice, not a substitute for observed purchase behavior. Monadic testing is most useful when isolated stated responses are needed across several price conditions.

Dynamic pilots and fake-door tests can be considered when a team needs a limited behavioral signal or is evaluating a live offer path. Their usefulness still depends on the decision, the assignment design, and what outcome is actually observed. A method name alone does not establish that a result measures revenue, retention, or margin.

Design a Test That Can Support a Pricing Decision

Begin with one falsifiable hypothesis. For example, define the price change, the population to which it applies, and the outcome that would support or reject the change. Then declare the primary and secondary metrics before collection. If revenue is the decision target, revenue per user should be primary where applicable, while conversion and other business measures remain secondary or diagnostic.

Plan the population and assignment before exposing respondents or visitors to prices. A live A/B test needs reliable randomization and price assignment that persists from the landing page through the invoice. A survey comparison needs a sample-size plan that accounts for the number of price points or conditions. An underpowered A/B pricing test cannot support a confident decision merely because one condition appears higher.

Set a stopping rule before the test starts. The rule should specify the duration or evidence required for review rather than allowing the team to stop when an early result looks attractive. The broader requirement is straightforward: a price test needs a falsifiable hypothesis, predeclared metrics, an adequate sample-size plan, reliable price assignment, and a defined stopping rule. The design should be specified before collection.

These checks are especially important when a price change is delivered across multiple customer paths. If the price shown does not persist through checkout and invoicing, the assigned condition is not the condition that was purchased. That weakens the connection between the test and the pricing decision.

Interpret Results Without Confusing Intent With Behavior

Interpretation begins with the evidence label. A survey result is stated purchase intent. A revenue number derived from survey responses is modeled or simulated revenue. A live transaction outcome is observed behavior. Keep those categories separate in the analysis and in the decision memo.

For an A/B experiment, review the point estimate with uncertainty rather than relying on a single observed difference. Revenue per user should accompany conversion because conversion alone does not describe the revenue outcome. Review retention, contribution margin, or lifetime value when the decision could have delayed effects. Early revenue results may not reveal later churn or lifetime value effects.

For Van Westendorp, interpret the intersecting curves as an acceptable range and an optimal price point estimate. For Gabor-Granger, interpret the curve as purchase intent across price points. For conjoint, interpret the estimated trade-offs among features, packages, and price. None of these survey outputs should be described as observed transactions.

Interpretation

The central interpretation is that method quality depends on question fit. A/B testing is the strongest option for observed purchase behavior when live traffic and controlled assignment are available. Van Westendorp provides an acceptable price range for early-stage decisions. Gabor-Granger supports a demand-curve question for one product or subscription. Conjoint addresses feature, package, and price trade-offs. The appropriate method is the one whose output answers the decision without overstating the evidence.

Use a Decision Rule for Your Next Test

For a pre-launch decision with no live purchase evidence, start with Van Westendorp when the immediate need is an acceptable price range. Use Gabor-Granger when the immediate need is to compare candidate prices for one product or subscription. Use conjoint when the unresolved issue is the structure of the package, including which features and prices belong together.

For subscription pricing, survey methods can narrow candidate prices, while a controlled experiment can validate the revenue outcome once sufficient traffic is available. This combined workflow uses research to prioritize hypotheses and behavioral testing to validate the live result. A survey-first workflow can precede controlled validation.

For a tier redesign, use conjoint when feature and package trade-offs are central. For live-traffic optimization, use an A/B experiment when controlled assignment and reliable price delivery are available. For limited traffic, begin with the smallest defensible survey study that matches the decision, then reserve a live experiment for the hypothesis that warrants validation.

Limits

Survey methods measure stated intent and can be affected by hypothetical bias, anchoring, wording, and detachment from the product experience. A/B experiments require sufficient live traffic, valid randomization, infrastructure that preserves price assignment through invoicing, and enough duration to observe relevant outcomes. Conjoint and monadic designs require greater sample and analytical complexity. Early revenue results may not reveal delayed churn or lifetime value effects. Fairness, customer communication, and jurisdiction-specific legal considerations can constrain unequal pricing pilots.

Next step

Write one falsifiable hypothesis, select one primary outcome, and choose the smallest defensible study that can answer it before committing to a price change. After that method and decision are clear, you can review Kinetic Pricing plans for the next stage of the work.

Additional context on these methods is available from tier-repricing decision, particularly suited to subscription pricing, 10%, research and benchmarks page, Kinetic Pricing's plans, Pricing experiments: A guide for businesses | Stripe and How to A/B test your pricing (Unbounce).

Run this method with your users

Start 30-day free trial to run every method with your users, or Buy one study when one decision needs evidence now.

Sources

Stripe, “Pricing experiments: A guide for businesses”

Unbounce, “How to A/B test your pricing”

Kinetic Pricing, “Van Westendorp Pricing Research for SaaS”

Kinetic Pricing, “Test SaaS Price Points with Gabor-Granger”

Kinetic Pricing, “Adaptive Conjoint Analysis”

Kinetic Pricing, “Price Sensitivity Analysis”

Greenbook, “Monadic Price Testing”