Research-backed startup idea validation: how the ranking works
By Eli Fayerman, attorney and software engineer · Last updated
Each discovery runs a three-phase pipeline: market research that surfaces real products, pricing, user complaints, and competitor gaps; opportunity generation where every idea must cite specific evidence from that research; and competitor validation that confirms real alternatives exist with real pricing. Results are then ranked by a business score for standalone viability, so the strongest directions surface first and weak ideas are automatically downranked; an alternate sort by insight quality is available in the workspace.
Why this is different from generic AI idea lists
Generic AI idea lists usually start from common patterns in training data. Research-backed startup idea validation starts from market evidence: named competitors, visible pricing, user complaints, workflow friction, and signals that buyers already spend time or money on the problem. The ranking is meant to answer a practical question: which direction is most worth validating next?
The discovery pipeline
When you enter a market or niche, the system runs these stages before returning results:
A market evidence base built from live web search, at most 24 hours old. Later stages are audited against it.
Ten candidate opportunities. A nurse gets healthcare plays; an attorney gets legal-ops.
The majority answer survives, never an average. When the runs fully disagree, the result says so instead of picking a winner.
The linter demotes competitor and price claims it cannot trace back to the evidence. The model proposes; code verifies.
Scores display only at the precision our reliability tests support: one decimal, with the measured re-run variance (about ±1.5) published beside the method. The three judgments that move the score most are sampled three times and averaged, so one contested judgment cannot swing a score by a point.
10 ranked opportunities, each with the market gap it exploits and a first step for this week. The top 5 are validated further: named competitors, checked prices, and a triple-sampled re-judgment of standalone viability.
Table 1 · Test–retest reliability of scan outputs
Prompt bundle v10 · measured 2026-09-02
| Measure | Estimate | 95% CI | Note |
|---|---|---|---|
| Rank stability, Spearman ρ | 0.75 | 0.47–0.94 | a |
| Reliability, ICC(1,1) | 0.61 | 0.23–0.94 | a |
| Same-item deviation, full scan | ± 1.5 pts | c | |
| Same-item deviation, validation phase | ± 0.45 pts | a | |
| Judgment consensus, k = 3 | 83% | b | |
| Idea recurrence, same topic | ~50% | c |
- 25 items × 3 independent passes; bootstrap CI, 2,000 resamples; validation phase only. ρ tracks whether items keep their order across re-runs, ICC whether they keep their scores.
- n = 240 production judgments; all three judgments are averaged into the score, so the 17% reaching no majority are not forced to a winner.
- 58 repeated-topic production pairs, aligned where two runs describe the same concept under different names; measured before the v10 averaging fix and kept as the conservative estimate. ~50% recur within 24 h, under 20% after research refresh.
Prior: winnability rating, single-sample ρ = −0.05; retired. Raw data and full audit: the scoring audit.
Each scan performs fresh web research. Re-running a topic may surface substantially different opportunities, especially as sources change. Even same-day re-runs vary: measured across repeated production scans, the same idea's score reproduces to about ±1.5 points, which is why scores are shown at one decimal with that error published, and why scores within about 1.5 of each other should be read as tied.
In more detail:
- Research: Searches for real products, real pricing, specific user complaints, quantified market trends, and named competitor gaps. Every finding must include at least one concrete data point: a product name with a price, a complaint with a source, or a trend with a number.
- Generation: Produces 10 opportunities grounded in those findings. Each opportunity must cite a specific research finding in its evidence field, and the set is diversified across different problems and audience segments rather than collapsing into one angle.
- Validation: Searches for real competitors for each generated opportunity and retrieves their actual pricing. If no real competitor with confirmed pricing exists, the system says so instead of fabricating comparisons.
How ranking works
Every opportunity is ranked by its business score. The base weights measure signal quality: speed to revenue, cost efficiency, evidence strength, pricing basis, and scalability. Viability adjustments are applied on top: standalone monetization strength, packaging fit, moat, and pricing confidence. Ideas flagged as free tools, lead magnets, or indirect-monetization plays are downranked automatically, so they do not outrank genuinely monetizable standalone products unless the entire market is weak. An alternate sort by insight quality (the base weights alone, before viability adjustments) is available in the workspace.
What operator fit means
If you specify a role or background, the system adjusts what types of opportunities it generates. Technical users see more SaaS, APIs, and automation tools. Non-technical users see more services, content products, and community plays. This is not cosmetic filtering: the generation prompt, packaging rules, and diversity constraints all change based on who is asking, so the ranked ideas are more likely to fit what the user can actually build and sell.
Limits and caveats
- Research quality depends on what is publicly available about a market. Thin or emerging markets produce weaker findings, and the system should say so rather than invent specifics.
- Scores are directional, not precise. A business score of 7.2 is meaningfully different from 3.1, but the difference between 6.8 and 7.0 is noise.
- Competitor validation confirms the market is active, not that a competitor is successful. Real pricing is evidence of willingness to pay, not proof of revenue scale.
- The system does not replace talking to customers. It gives you a stronger ranked starting point; the next step is always real validation with buyers.
Related questions
See ranked opportunities
See the methodology work on a real market. This opens the workspace with a research-heavy starting point so you can test the ranking flow on your own niche.
Run a research-backed scan →