Market Research Survey Questions That Work for a New Product
By Omar Zidan, Founder
Good market research survey questions for a new product ask people to rank options they already know or to report what they did recently. They never ask what people would do. A survey does two jobs well: ranking known alternatives and segmenting a list you already have. Drop hypothetical purchase intent questions entirely, because they overstate demand for new products.
Why is a survey a weak test of demand?
A survey records what people say. Demand is what they do with their money and time, and the gap between the two is widest in exactly the case you care about.
Marketing managers routinely use purchase intentions to predict sales, so Morwitz, Steckel and Gupta studied when that works and when it fails. Their paper in the International Journal of Forecasting found that intentions track purchases more closely for existing products than for new products, and for durable products than for non-durable products. The full list of findings is worth reading as a checklist. According to the paper, intentions predict better for short than for long time horizons, when people are asked about specific brands or models instead of a whole product category, and when purchase intentions are collected in a comparative mode than when they are collected monadically.
Now look at the typical founder survey. It asks about a new product, usually software or a subscription, with no set purchase date (a long horizon). It often asks about interest in a kind of product instead of a specific brand or model (category level). And it shows one concept alone (monadic). That puts it at the worst end of four of the factors the study found.
So the standard "Would you buy this?" survey is weak. Asking it more cleverly won't fix it, because the problem is the premise. But the same research shows where surveys do work: comparisons, specific options and the near term. Those conditions define the two jobs below.
What are the two jobs a survey actually does well?
Job 1: ranking options people already know. When respondents compare things they have used or can picture in detail, you get the comparative, specific kind of answer the research favors. "Which of these five problems cost you the most time last month?" is a fair question. People have lived through all five, and you only need them to put them in order.
Job 2: segmenting a list you already have. If you have a waitlist, a newsletter, a customer base or a community you run, a survey can sort those people into groups by what they do now: the tools they use, what they pay and how often the problem comes up. You aren't asking the survey to prove demand. You're asking it to tell you which slice of your list to test first.
Neither job answers "will this sell?" That question gets answered by a behavioral test such as a pre-sale or a fake door. The survey's job is to make that test sharper and cheaper.
A survey sent to strangers from a panel does neither job well. Strangers can't segment a list you own. And you have no way of knowing whether the people ranking your options are people who would ever buy. Survey your own audience or don't survey.
Which survey questions guarantee useless answers?
Some question types produce data that looks clean in a chart and predicts nothing. Here are the common ones, starting with the worst.
| Question type | Example | Why it fails | Ask this instead |
|---|---|---|---|
| Hypothetical purchase intent | "How likely are you to buy a tool that does X?" | New product, no set date, one concept shown alone: the worst conditions in the research above | "What did you last pay to solve X, and to whom?" |
| Open willingness to pay | "What would you pay for this?" | Costs nothing to answer, so people anchor on whatever feels reasonable | "What do you pay now for your current workaround?" |
| Feature desirability | "Would you use a feature that auto-syncs your calendar?" | Nobody turns down a free imaginary feature | Rank five features against each other, with no "all of them" option |
| Agree/disagree statements | "Scheduling is a major pain for my business. Agree?" | The wording hands people the answer | "How many hours did scheduling take last week?" |
| Future frequency | "How often would you use this?" | People are poor at predicting how often they will do something new, and a medium-to-large change in intention leads to only a small-to-medium change in behavior | "How many times did you do X in the last 30 days?" |
| Category-level interest | "Are you interested in productivity apps?" | Category questions predict worse than specific ones | Name the exact tools and ask which one they opened this week |
The pattern holds across every row. A useless question is about the future, about a hypothetical, or about a single option. A useful one is about the recent past, about something specific, or about a trade-off between options.
Hypothetical purchase intent leads the table because it does the most damage. A top-box score (the share answering "definitely would buy") feels like a forecast. It gets multiplied by a market size and pasted into a pitch deck. Research firms turn top-box scores into forecasts with rules of thumb, such as treating about half of top-box respondents and 10 to 20 percent of the next box down as buyers, and even those rules depend on who you ask and what information you give them. For a product that doesn't exist, in a category you define, sold at a price nobody has paid, you have no such calibration. If you want that number for your product, you have to produce it yourself: put a real price in front of real people and count the payments.
How do you write a ranking survey that tells you something?
Work through this in order.
- List only options respondents already know from experience. These can be problems they have had, tools they have used or tasks they do now. If an option needs a paragraph of explanation, cut it.
- Keep it to five to seven items. Beyond seven, people stop comparing and start skimming, so switch to best-worst for longer lists.
- Force a trade-off. Ask for the single most costly item, or the top two, or a full order. Never let people rate each item on its own 1 to 5 scale. Everything ends up a 4 and you learn nothing. Rating each option alone is the monadic mode the research found weaker.
- Anchor the question in a recent window. "In the last month" is better than "in general." Short horizons predict better.
- Randomize the order of options. People favor the first items they see, so shuffle the order for each respondent.
- Add one behavioral follow-up to the top choice. "What do you currently do about it?" with choices like "nothing," "a spreadsheet," "a paid tool" and "I hired someone." This one question often matters more than the ranking itself.
If you have more than seven items, use a best-worst format (often called MaxDiff). Each screen shows four or five items, and the respondent picks the most and least important. Most survey tools either support it or can approximate it with repeated forced-choice screens.
How do you segment a list you already have?
Segmentation questions are about the present and the past. You want groups defined by behavior, because behavior is what your later test will check against.
Four questions do most of the work:
- What do you use today to handle X? (a fixed list plus "other")
- In the last 30 days, how many times did X happen? (number ranges)
- Do you currently pay for anything to deal with X? (yes or no, then the amount in ranges)
- One firmographic or context question that matches how you would sell. That could be team size, business type, or whether they're buying for themselves or an employer.
Then cross the answers. The segment you want is people with high frequency who already pay for a workaround. They have shown demand with their own money. Someone who says the problem is "very painful" but has spent nothing on it and does it twice a year is telling you a story. Someone paying for a clumsy tool every month is showing you a budget.
How many responses do you need before you trust a ranking?
More than most founders expect, and the gap between your top options is what decides it.
Say you have a waitlist of 1,200 people and 300 answer your survey. You asked which of five problems cost them the most time last month. Problem A gets 40% of first-place votes and problem B gets 34%. That looks like a clear win for A.
Check the margin of error. For a single share at 95% confidence, the formula is 1.96 × √(p(1 − p) / n). For A, that's 1.96 × √(0.40 × 0.60 / 300) = 1.96 × 0.0283 ≈ ±5.5 points. The difference between A and B needs a slightly different formula, because both shares come from the same respondents: 1.96 × √((pA + pB − (pA − pB)²) / n).
Plug in the numbers: 0.40 + 0.34 − 0.06² = 0.7364. Divide by 300 to get 0.00245. The square root is 0.0495, and times 1.96 that's about ±9.7 points.
Your 6-point lead sits inside a ±9.7-point margin. You cannot tell A and B apart at this sample size. To tell a 6-point gap apart from noise, you would need about 0.7364 × (1.96 / 0.06)², which is roughly 786 complete responses. Even then, 786 responses only gives you roughly even odds of detecting a true 6-point gap. That's more than half your list, and far more than a 25% response rate will give you.
Segments shrink even faster. If 90 of your 300 respondents already pay for a workaround, a 50/50 split inside that group carries a margin of about ±10.3 points. Anything finer than "clearly most" or "clearly few" is noise at that size.
What this means in practice:
- Treat a ranking as a shortlist, not a winner. Take the top two or three options into a behavioral test.
- Only act on gaps larger than your computed margin. Run the formula on your own numbers before you present results to anyone, including yourself.
- Remember who answered. The 300 who replied may be your most engaged 300, and they aren't a random draw from the 1,200.
What should you do with the results?
Turn every survey finding into a test that costs the respondent something. If you can't separate problem A from problem B, write two landing pages with prices on them. Send each to the segment that already pays for a workaround, and count clicks on the buy button or deposits taken. The survey has done its work once it tells you which two pages to write and who to send them to.
If your survey has only purchase intent questions, don't try to salvage the data. Rewrite it around the recent past, forced trade-offs and current spending, and send it again. It will look less impressive in a chart, but it will actually predict something.
Questions
Should I survey strangers from a paid panel to validate a new product?
Usually not. Panel respondents can't help you segment a list you own, and you can't tell whether they would ever buy. Surveying your own waitlist, newsletter or customers gives you answers you can act on.
How long should a new product survey be?
Keep it short enough to finish in a few minutes. A ranking question of five to seven items, one behavioral follow-up, and three or four segmentation questions cover most needs. Every extra question lowers completion rates and adds noise.
Is a 1 to 5 rating scale ever useful in a product survey?
Rating each option on its own scale tends to produce similar scores for everything, so it rarely separates priorities. Forced trade-offs, such as picking the single most costly problem or a best-worst format, reveal real preferences. Use rating scales only for tracking one known item over time.
What should replace a purchase intent question?
Ask what people did recently and what they pay now. Questions like what they last paid to solve the problem, and to whom, reflect actual spending. To measure intent to buy your product, run a behavioral test such as a priced landing page or pre-sale.
What response rate should I expect from surveying my own list?
Rates vary widely by audience and engagement, but many lists return somewhere around 10 to 30 percent. Remember that those who reply are often your most engaged members, so results may not represent the whole list.
