# Value-based pricing with AI: research first, model second

> A practical pricing workflow that uses AI for evidence synthesis and scenario modeling without inventing willingness to pay or replacing customer research.
> Author: Roman Belov · Published: 2026-04-19 · Source: https://futurecraft.pro/blog/value-based-pricing-ai/

The dangerous pricing prompt is: “Here is my product. What should I charge?”

A model will answer. It may even return a precise range, an “optimal” price, and a tidy Van Westendorp chart. None of those numbers came from a buyer.

AI is useful in pricing when it reduces research work: extracting evidence from interviews, normalizing public plans, checking arithmetic, and comparing scenarios. It becomes harmful when plausible prose is mistaken for willingness to pay.

This workflow keeps those two roles separate.

## Four numbers that pricing discussions often mix up

Value-based pricing starts with customer outcomes, but “value” is not one number.

1. **Delivered value** is the economic change the product can create: labor avoided, cost removed, revenue enabled, or loss reduced.
2. **Perceived value** is the part the buyer understands and believes before purchase.
3. **Willingness to pay** is the most the buyer would exchange under a specific offer and context.
4. **Feasible price** must also cover cost-to-serve, sales cost, support, risk, and the margin needed by the business.

A product may deliver $20,000 of annual value while the buyer perceives $5,000 and will approve only $1,200 from the available budget. Applying a universal “capture 10–25% of value” rule hides the actual constraint.

Current [Stripe guidance on value-based pricing](https://stripe.com/resources/more/cost-based-and-value-based-pricing) makes the useful distinction: cost helps define a floor, while perceived customer value informs the ceiling. The price still has to survive market pressure and economics between them.

## Step 1: choose one segment and one job

Do not price “for SMBs.” Start with a [specific customer profile](/blog/icp-definition-ai/) and a recognizable workflow:

```text
Segment: performance marketing agencies with 10–30 employees
Buyer: head of client services
Job: prepare and review monthly client reports
Alternative: analyst work plus two existing tools
Purchase context: annual team subscription
```

The same reporting product can have different value for a solo consultant, an agency, and an enterprise marketing team. Mixing them produces averages that describe nobody.

Record the segment before collecting data. If an interview or competitor serves another segment, label it instead of letting it silently influence the model.

## Step 2: build an evidence ledger

Use a table with one row per claimed outcome. For example:

| Outcome | Baseline | Product change | Annual impact | Evidence | Confidence |
|---|---:|---:|---:|---|---|
| Report preparation | 12 h/month | 5 h/month | 84 hours | 4 interviews + time logs | Medium |
| Existing tool removed | $80/month | $0 | $960 | invoices | High |
| Error rework | 3 h/month | unknown | unknown | no measured pilot | Low |

The last row stays unknown. Do not ask an LLM to fill it with an industry average unless you have a named, relevant source and can explain why it transfers.

For time savings, use the buyer's loaded labor cost and the share of saved time that can actually be redeployed. Saving ten minutes across scattered tasks is not automatically ten minutes of economic value. For revenue claims, separate correlation from causal evidence. For risk reduction, show event probability and loss separately so reviewers can challenge both.

The ledger is a range, not a single total:

```text
conservative value = verified savings only
base value         = verified + supported pilot effects
upside value       = base + explicitly uncertain outcomes
```

That structure lets AI help without becoming the source.

## Step 3: give AI an evidence-bound task

A useful prompt requires provenance and permits “unknown”:

```text
You are organizing pricing research, not estimating missing facts.

Inputs:
- customer interview notes
- current workflow costs
- pilot measurements
- competitor pricing pages with retrieval dates

Tasks:
1. Extract claimed customer outcomes.
2. Link every number to a source.
3. Separate observed facts, customer estimates, and assumptions.
4. Build conservative, base, and upside value scenarios.
5. Flag contradictions and missing inputs.
6. Do not invent benchmarks. Return "unknown" when evidence is absent.

Output:
- evidence ledger
- assumptions register
- questions for the next five interviews
```

Review every source link and calculation. The model is a fast analyst, not an authority. If it browses competitor pricing, store the page URL and date because packaging changes.

This is also where AI can cluster interview language. “Faster reporting,” “fewer revisions,” and “less Friday-night work” may describe one outcome, but preserve the original quotes and segment labels before merging them.

## Step 4: map alternatives and the pricing metric

Competitor price is context, not the answer. Normalize each alternative by:

- target segment;
- billing interval;
- included seats or usage;
- mandatory add-ons;
- implementation and service fees;
- overage rules;
- annual discount;
- the workflow it replaces.

Include the status quo. A spreadsheet plus analyst time may be the strongest competitor.

Then choose the unit that grows with customer value. Common candidates are account, seat, location, transaction, successful workflow, data volume, or a hybrid base fee plus usage.

Score each candidate against evidence:

| Criterion | Question |
|---|---|
| Value alignment | Does the bill rise when the buyer receives more value? |
| Predictability | Can the buyer forecast the invoice? |
| Verifiability | Can both sides measure the unit consistently? |
| Cost alignment | Does heavy usage create material variable cost? |
| Gaming | Does the metric encourage workarounds or shared accounts? |
| Sales friction | Can a buyer explain it to procurement? |

A technically elegant metric can still fail if customers cannot budget for it. A predictable flat tier can still fail if one heavy account consumes all support and infrastructure margin.

## Step 5: use Van Westendorp only with respondents

The Van Westendorp Price Sensitivity Meter asks relevant respondents when a product becomes:

- so cheap that quality is doubtful;
- a bargain;
- expensive but still worth considering;
- too expensive to consider.

The method plots cumulative response distributions and calculates their intersections. It is a **survey instrument**, not a reasoning template. An LLM that generates the answers is simulating a population from its training data, not measuring your segment.

AI can still help:

1. rewrite the offer so every respondent evaluates the same package;
2. check survey logic and currency formatting;
3. remove invalid or inconsistent responses using declared rules;
4. plot the curves and bootstrap uncertainty;
5. compare segments without pooling them prematurely.

Stated willingness to pay has known limitations. A meta-analysis of 77 studies found material hypothetical bias on average, with accuracy varying by method and product context ([Schmidt & Bijmolt, 2020](https://doi.org/10.1007/s11747-019-00666-6)). Do not “correct” every survey with one universal percentage. Use the finding as a reason to pair stated answers with behavior.

## Step 6: test an offer, not a naked number

A price is inseparable from package, term, guarantee, onboarding, and buyer.

For founder-led B2B sales, create two or three coherent offer cards. Keep the core outcome comparable, vary the price-package hypothesis, and record:

- segment and company size;
- quoted price and term;
- package shown;
- whether the buyer advanced;
- discount requested and reason;
- procurement or budget constraint;
- realized contract value;
- activation and retention after the sale.

Do not change the price mid-call based on how wealthy the prospect looks. Define the assignment rule before the conversation.

For self-serve products, randomize eligible visitors or accounts. The primary metric should reflect money, such as revenue per eligible visitor or contribution margin per account. CTR alone can favor a cheap offer that produces less revenue and worse retention.

Calculate sample size from baseline conversion, minimum detectable effect, desired power, and the number of variants. “Three hundred visits per page” is not a significance rule. Low-volume B2B teams often learn faster from repeated sales proposals than from an underpowered landing-page experiment.

If the product is not available, say that before collecting a lead. A transparent waitlist can measure interest; a fake purchase flow followed by a surprise “not built yet” screen damages trust and still does not measure retention or realized revenue.

## Step 7: package only real differences

Three tiers are common, but they are not mandatory. The classic compromise-effect research found that an option can gain share when it becomes the middle choice in a set ([Simonson, 1989](https://doi.org/10.1086/209205)). That result does not imply a universal `1x : 2.5x : 5x` ratio, nor that the middle plan maximizes revenue.

Create a tier only when it maps to a real difference:

- a segment with a different workflow;
- a value metric threshold;
- higher cost-to-serve;
- governance, security, or service requirements;
- a capability that genuinely changes the outcome.

Artificial feature removal may raise short-term upgrades and increase long-term frustration. Start with the smallest package set that sales can explain in one minute.

## Step 8: keep a pricing decision record

The first price is a hypothesis. Write down why it exists. An illustrative record:

```yaml
segment: agencies_10_30
offer_version: 2026-08-a
pricing_metric: client_workspace
price: 249_monthly
annual_term_price: 2490
evidence:
  interviews: 9
  paid_pilots: 3
  competitor_pages_checked: 6
assumptions:
  - 7 hours of report work saved per month
  - no implementation service included
review_when:
  - 20 qualified proposals
  - 10 activated accounts
  - material cost-to-serve change
```

Review by cohort, not by one blended conversion rate. Watch realized revenue, gross margin, discounting, activation, [retention behavior](/blog/retention-curve-pmf/), expansion, support load, and loss reasons. A high conversion rate does not automatically mean price is too low; it may mean acquisition is well targeted. A low rate may reflect weak positioning, poor traffic, or the wrong package rather than price.

Connect the model to the [unit economics worksheet](/blog/unit-economics-calculator-ai/), but keep assumptions visible. LTV derived from a few months of churn data is not a fact merely because a spreadsheet returns a precise number.

## A compact pricing workflow

1. Define one segment, buyer, job, and purchase context.
2. Build an evidence ledger with sources and confidence.
3. Use AI to extract and reconcile evidence, never to invent missing WTP.
4. Normalize alternatives and choose a value-aligned metric.
5. If useful, run Van Westendorp with actual respondents.
6. Test coherent offers through real proposals or randomized purchase behavior.
7. Evaluate revenue, margin, activation, and retention together.
8. Record the decision and the conditions that trigger review.

AI can compress the analysis from messy notes to an auditable model. It cannot remove the part where a customer sees the offer and chooses what to do with real time, budget, and risk.

## Sources

- [Stripe: cost-based and value-based pricing](https://stripe.com/resources/more/cost-based-and-value-based-pricing)
- [Stripe: pricing strategies for new products](https://stripe.com/resources/more/pricing-strategies-for-new-products)
- [Schmidt & Bijmolt: meta-analysis of hypothetical bias in willingness-to-pay research](https://doi.org/10.1007/s11747-019-00666-6)
- [Simonson: attraction and compromise effects](https://doi.org/10.1086/209205)
- [MIT OpenCourseWare: measuring customer reactions to prices](https://ocw.mit.edu/courses/15-818-pricing-spring-2010/resources/mit15_818s10_lec04/)
