Ask a language model for the size of the UK pet grooming market and it will give you a number. It will be plausible, specific, and delivered with complete confidence. It may also be entirely fabricated.
This is not a flaw that better prompting solves. It is a consequence of how the models work: they produce likely-sounding text, and a specific figure is more likely-sounding than an admission of ignorance. The failure is silent by design.
Which is a problem, because AI is genuinely good at the parts of market research that are slow — reading widely, spotting patterns, and structuring findings. The task is separating those from the parts where it will confidently invent.
The Dividing Line
There is a clean test, and it holds up better than any prompt technique:
Can the model verify this, or is it recalling it?
A model asked to summarise ten competitor pricing pages it has just been shown is doing verifiable work — the source is in front of it.
A model asked what those competitors charge from memory is producing a guess shaped like an answer.
| Reliable | Unreliable |
|---|---|
| Summarising documents you supply | Market size figures from memory |
| Structuring and categorising your own data | Growth rates and CAGR |
| Extracting themes from reviews you paste in | Competitor revenue or headcount |
| Drafting research questions | Survey statistics and percentages |
| Reading and comparing pages it can fetch | Anything with a citation it produced itself |
That last row deserves emphasis.
Models can fabricate citations readily — a real-sounding journal, a real-sounding author, or a report that does not exist. If a model hands you a source, the source is a hypothesis until you open it.
Five Research Jobs and How to Run Each One
1. Competitor Pricing and Positioning
Don't ask what competitors charge. Give the model the pages.
Collect the pricing and homepage copy of eight to twelve competitors. Paste them into the model and ask for a comparison table covering:
- Pricing tiers
- Prices
- What's included at each level
- Language each competitor uses to describe its target customer
- Key positioning differences
The output is genuinely useful and grounded because every cell traces back to text you supplied.
What usually emerges are positioning gaps:
- The customer segment nobody addresses
- The feature everyone charges extra for
- The price band that is conspicuously empty
- Messaging that competitors repeat across the market
Time required: Roughly one hour for collection and a few minutes for analysis.
Manually, the same task can take most of a day.
2. Customer Language
One of the most under-used research sources is text your customers have already written.
Useful sources include:
- Customer reviews
- Support tickets
- Sales call notes
- Reddit threads
- Comments on competitor posts
- Customer feedback
- Online discussions
Gather several hundred examples and ask the model to identify:
- Recurring complaints
- Words customers use to describe the problem
- Common frustrations
- What customers are trying to achieve
- The moment they decided to look for a solution
This is pattern-finding across a corpus — exactly what these systems do well.
It can make the difference between writing generic copy such as "streamline your workflow" and using the actual language your customers use to describe their problems.
3. Market Sizing
This is the most dangerous request, and it is also one of the first requests many people make.
Sizing a market from memory is unreliable. But market sizing is not impossible — it simply needs to be built rather than recalled.
Use published figures that you can source yourself, such as:
- National statistics offices
- Industry associations
- Company filings
- Government databases
- Published industry reports
Then have the model perform the arithmetic and help you test the logic.
A defensible bottom-up structure looks like this:
Businesses in the category × share matching your customer profile × realistic annual spend = estimated market opportunity
For example:
Businesses in the category (from a national business register) × the share matching your customer profile (your estimate, clearly stated as an estimate) × realistic annual spend (from your own pricing or published benchmarks)
Every input should be traceable.
Every assumption should be visible and arguable.
That is what makes the calculation useful in a business plan or pitch — not simply the size of the final number, but the fact that a reader can challenge a specific input or assumption.
A number you cannot source is not research. It is a placeholder that will eventually be quoted back to you.
4. Finding Competitors You Don't Know About
Direct competitors are usually easy to find because you already know who they are.
The more interesting alternatives are the solutions customers choose instead of your entire category.
These could include:
- Spreadsheets
- Agencies
- Freelancers
- In-house employees
- Manual processes
- Doing nothing
There are two grounded approaches.
Search Customer Problems
Search the phrases your customers would actually use — the problem, not the product category.
Record everything that ranks.
Search rankings can provide evidence that people are searching for solutions to that problem.
Check AI Recommendations
Ask ChatGPT, Claude, and Perplexity what they would recommend for that specific problem.
Record the companies, tools, or services they mention.
Increasingly, these recommendations can influence the shortlist a potential buyer sees.
The key is to record what the models recommend based on live sources rather than treating their answers as unquestionable facts.
5. Trend and Demand Validation
Do not simply ask a model whether a market is growing.
Instead, give it actual data and ask it to interpret the data.
Useful inputs include:
- Google Trends exports
- Search-volume data from keyword tools
- Website analytics
- Customer acquisition data
- Sales data
- Conversion data
The model can analyse the shape of the data and explain potential patterns.
But the important distinction is that the underlying data came from somewhere real.
Verification, Briefly
Three rules cover most of the risk.
Rule 1: Every Number Gets a Source
Every statistic or figure should have a source before it leaves the document.
Not a source the model named.
A source you personally opened and checked.
Rule 2: Check the Reasoning, Not Just the Answer
Ask the model to explain how it reached its conclusion.
Then check the reasoning.
A bad chain of reasoning is often easier to identify than a bad number.
A useful follow-up question is:
"Where does this figure come from, and what would make it wrong?"
Rule 3: Give the Model Permission to Say "I Don't Know"
Tell the model explicitly:
"If you don't have a reliable source for this, say so."
This encourages the model to distinguish between information it can verify and information it cannot confidently support.
The default behaviour of language models is often to produce an answer rather than stop when information is uncertain.
A Workable Research Sequence
For a small business or solo founder, a full research pass can be completed in about a day.
| Step | Grounded In | Time |
|---|---|---|
| Collect competitor pages | The pages themselves | 1 hr |
| Comparison and gap analysis | Those pages | 30 min |
| Gather customer text | Reviews, tickets, forums | 1–2 hrs |
| Extract themes and language | That corpus | 30 min |
| Assemble sizing inputs | Public statistics | 1 hr |
| Build the sizing model | Your inputs | 30 min |
| Check AI engine recommendations | Live queries | 30 min |
| Write it up | Everything above | 1 hr |
Every step has a named source.
Nothing in the finished document needs to rest on the model's memory.
What Good Output Looks Like
A market research document you can defend has three properties, and they are easy to check.
1. Every Figure Has a Source
Every statistic or numerical claim should have a source that a reader can open and verify.
2. Every Assumption Is Clearly Labelled
If you estimated something, say that it is an estimate.
Do not present an assumption as a fact.
3. Conclusions Follow the Evidence
The conclusions should come from the evidence rather than appearing alongside unsupported claims.
A document that reads smoothly but fails all three tests is the standard output of asking a language model to "do market research for my business."
It looks like the real thing, which is precisely the risk.
Plausible research is more dangerous than obviously bad research because it survives review.
Frequently Asked Questions
Which model is best for market research?
The choice of model matters less than whether you supply reliable sources.
A weaker model reading real documents can be more useful than a stronger model working from memory.
Can AI replace a market research firm?
AI can replace or reduce some of the analysis and synthesis work involved in market research.
It does not replace primary research such as:
- Original surveys
- Customer interviews
- Proprietary panels
- First-party research
If the answer does not exist in accessible information, a language model cannot simply retrieve it from nowhere.
How do I check a statistic a model gave me?
Search for the statistic directly.
If it is real and well-established, you should be able to locate the underlying source or a credible publication reporting it.
If several minutes of searching produces nothing credible, treat the statistic as unverified rather than assuming that it is simply difficult to find.
Is it safe to paste customer data into a model?
It depends on the model, account type, data, and applicable privacy commitments.
Before uploading customer information, check the provider's current data-use and privacy policies.
When there is any doubt, anonymise customer information before using it for analysis and make sure the process is consistent with your own privacy obligations.
How often should market research be redone?
Different parts of market research change at different speeds.
Competitor pricing: Quarterly can be reasonable because pricing can change frequently.
Customer language: Usually changes more slowly and can be reviewed periodically.
Market sizing: Often only needs to be revisited annually unless there is a major change in the market, business model, or available data.
Final Takeaway
AI can be extremely useful for market research, but its value depends on how the research is conducted.
Use AI to read, compare, categorise, analyse, and synthesise real evidence.
Do not treat confident answers, statistics, citations, or market estimates produced from memory as automatically reliable.
The strongest workflow is simple:
Source the evidence → give it to the model → analyse it → verify the important claims → clearly label assumptions → draw conclusions from the evidence.
That approach turns AI from a generator of plausible answers into a practical research assistant.
Questions
Which model is best for market research?
The choice of model matters far less than whether you supply reliable sources. A weaker model reading real documents can be more useful than a stronger model working from memory.
Can AI replace a market research firm?
AI can replace much of the analysis and synthesis work involved in market research. However, it does not replace primary research such as original surveys, interviews, or proprietary research panels. If the answer does not exist in accessible information, no model can simply retrieve it.
How do I check a statistic a model gave me?
Search for the statistic directly. If it is genuine, you should be able to find the primary source or a credible source reporting it. If several minutes of searching produces nothing credible, treat the statistic as unverified rather than assuming it is simply obscure.
Is it safe to paste customer data into a model?
It depends on the model, the type of data, and the provider's data-use policies. Some consumer AI products may retain conversations under their applicable settings, while business offerings may provide different data protections. If there is any doubt, anonymise customer information before sharing it and check your own privacy commitments.
How often should market research be redone?
Competitor pricing can change frequently, so reviewing it quarterly can be reasonable. Customer language generally changes more slowly and can be reviewed periodically. Market sizing typically does not need to be redone more than annually unless there is a significant change in the market or available data.
