Founders rarely set out to break their own survey. They just want the question to sound more natural, more on-brand, a little less like a template lifted from a blog post. So "how would you feel" becomes "how disappointed would you be." "No longer use it" becomes "no longer existed." A neutral option gets added because forcing a choice feels harsh. None of it looks like a big deal.

It is a big deal, for one specific reason: the number you're comparing yourself against, 40%, wasn't calculated from a general vibe about disappointment. It came from Sean Ellis running one specific question, worded one specific way, across roughly a hundred startups. That's the entire basis for the benchmark. Reword the instrument and the benchmark doesn't automatically travel with it.

The exact wording the benchmark is built on

Here's the original question, unmodified:

"How would you feel if you could no longer use [product]?"

With exactly three forced options: Very disappointed, Somewhat disappointed, Not disappointed, plus a separate "N/A, I no longer use it" bucket that gets excluded from the denominator. Rahul Vohra used this same question, word for word, when he ran Superhuman's first PMF survey and got the 22% score that eventually became 58%, a run-through documented in his widely cited First Round Review interview. He didn't invent a new question. He used Ellis's, because using a different one would have made the number incomparable to the benchmark that gave it meaning.

Source: First Round Review, "How Superhuman Built an Engine to Find Product-Market Fit".

A benchmark is only a benchmark for the instrument it was measured with. Change the instrument, and 40% is just a number you made up.

Four "small" edits that aren't small

1. "How disappointed would you be" instead of "how would you feel"

This looks like a synonym swap. It isn't. "How would you feel" is direction-neutral, the respondent supplies the emotion. "How disappointed would you be" presupposes that disappointment is the correct frame before the respondent has answered anything, which nudges toward picking a disappointment level rather than considering "I wouldn't care" as the honest, neutral starting point. In survey design this is a form of a leading question: the wording implies an expected answer rather than leaving the direction open. General survey-methodology guidance on leading questions and how they skew results is well documented; see for example SurveyMonkey's guide to eliminating order and framing bias.

2. "If [product] no longer existed" instead of "if you could no longer use [product]"

These sound like the same hypothetical. They aren't. "No longer use it" is about losing access to one tool: you'd switch, or go back to your old workaround. "No longer existed" invokes something bigger, the company disappearing, the team scrambling, contracts and integrations unwinding. That's a scarier scenario for reasons that have nothing to do with how much anyone loves the product day to day, and it will inflate "very disappointed" answers on switching-cost grounds rather than on genuine product love. The score goes up; the thing it's supposed to measure hasn't changed.

3. Adding a neutral middle option

The original scale is deliberately a forced choice: three options, no escape hatch. Adding "neutral" or "no opinion" feels more respectful of uncertain respondents, but it changes the structure the 40% threshold was calibrated against. Some share of people who would have picked "somewhat disappointed" under duress will instead take the neutral exit, and your "very disappointed" percentage is now drawn from a different pool of forced answers than the one the benchmark used. Unbalanced or restructured response scales are a standard survey-bias concern for exactly this reason: changing the option set changes the distribution, independent of any change in actual sentiment.

4. Substituting NPS ("how likely are you to recommend?")

This isn't a wording tweak, it's a different metric wearing the PMF survey's clothes. NPS measures willingness to promote a product to others. The Sean Ellis question measures how essential a product is to the person using it. The two often move together, but they're not interchangeable, and an NPS score has no defined relationship to the 40% line. We cover the full difference here, but the short version for this post: if you ran NPS, you don't have a PMF score. You have an NPS score.

Run the exact question, not your best guess at it

PMFtracker ships the unmodified Sean Ellis question and scale, pre-built, so your score is comparable to the 40% benchmark from day one.

Start free → 14-day free trial · No credit card

Order matters too, not just wording

Even with the exact question, the order you present the three options in can nudge results, a well-established finding in survey research known as order or primacy bias, where whichever option appears first tends to attract a disproportionate share of picks. Keep the standard order (Very disappointed → Somewhat disappointed → Not disappointed) rather than reshuffling it for visual reasons in your survey tool.

What to do instead

None of this is precious for its own sake. A benchmark is only useful if the thing you're measuring is the same thing it was calibrated on. Reword the question and you can still learn something interesting about your users, just not whether you've crossed 40%.

Get a score you can actually compare to 40%

PMFtracker runs the unmodified Sean Ellis question, scores you against the benchmark, and segments the results, so the number means what it's supposed to mean.

Measure your PMF score free → Set up in 5 minutes · No credit card required