Research-backed startup idea validation: how the ranking works

By , attorney and software engineer · Last updated

Each discovery runs a three-phase pipeline: market research that surfaces real products, pricing, user complaints, and competitor gaps; opportunity generation where every idea must cite specific evidence from that research; and competitor validation that confirms real alternatives exist with real pricing. Results are then ranked by a business score for standalone viability, so the strongest directions surface first and weak ideas are automatically downranked; an alternate sort by insight quality is available in the workspace.

Why this is different from generic AI idea lists

Generic AI idea lists usually start from common patterns in training data. Research-backed startup idea validation starts from market evidence: named competitors, visible pricing, user complaints, workflow friction, and signals that buyers already spend time or money on the problem. The ranking is meant to answer a practical question: which direction is most worth validating next?

The discovery pipeline

When you enter a market or niche, the system runs these stages before returning results:

1 Researchlive web search

A market evidence base built from live web search, at most 24 hours old. Later stages are audited against it.

2 Generationshaped by your background

Ten candidate opportunities. A nurse gets healthcare plays; an attorney gets legal-ops.

3 Validation3 independent runs

The majority answer survives, never an average. When the runs fully disagree, the result says so instead of picking a winner.

4 The gatedeterministic code

The linter demotes competitor and price claims it cannot trace back to the evidence. The model proposes; code verifies.

5 No fake precisionwhere ideas get killed

Scores display only at the precision our reliability tests support: one decimal, with the measured re-run variance (about ±1.5) published beside the method. The three judgments that move the score most are sampled three times and averaged, so one contested judgment cannot swing a score by a point.

10 ranked opportunities, each with the market gap it exploits and a first step for this week. The top 5 are validated further: named competitors, checked prices, and a triple-sampled re-judgment of standalone viability.

Table 1 · Test–retest reliability of scan outputs
Prompt bundle v10 · measured 2026-09-02

MeasureEstimate95% CINote
Rank stability, Spearman ρ0.750.47–0.94a
Reliability, ICC(1,1)0.610.23–0.94a
Same-item deviation, full scan± 1.5 ptsc
Same-item deviation, validation phase± 0.45 ptsa
Judgment consensus, k = 383%b
Idea recurrence, same topic~50%c
  1. 25 items × 3 independent passes; bootstrap CI, 2,000 resamples; validation phase only. ρ tracks whether items keep their order across re-runs, ICC whether they keep their scores.
  2. n = 240 production judgments; all three judgments are averaged into the score, so the 17% reaching no majority are not forced to a winner.
  3. 58 repeated-topic production pairs, aligned where two runs describe the same concept under different names; measured before the v10 averaging fix and kept as the conservative estimate. ~50% recur within 24 h, under 20% after research refresh.

Prior: winnability rating, single-sample ρ = −0.05; retired. Raw data and full audit: the scoring audit.

Each scan performs fresh web research. Re-running a topic may surface substantially different opportunities, especially as sources change. Even same-day re-runs vary: measured across repeated production scans, the same idea's score reproduces to about ±1.5 points, which is why scores are shown at one decimal with that error published, and why scores within about 1.5 of each other should be read as tied.

In more detail:

How ranking works

Every opportunity is ranked by its business score. The base weights measure signal quality: speed to revenue, cost efficiency, evidence strength, pricing basis, and scalability. Viability adjustments are applied on top: standalone monetization strength, packaging fit, moat, and pricing confidence. Ideas flagged as free tools, lead magnets, or indirect-monetization plays are downranked automatically, so they do not outrank genuinely monetizable standalone products unless the entire market is weak. An alternate sort by insight quality (the base weights alone, before viability adjustments) is available in the workspace.

What operator fit means

If you specify a role or background, the system adjusts what types of opportunities it generates. Technical users see more SaaS, APIs, and automation tools. Non-technical users see more services, content products, and community plays. This is not cosmetic filtering: the generation prompt, packaging rules, and diversity constraints all change based on who is asking, so the ranked ideas are more likely to fit what the user can actually build and sell.

Limits and caveats

Related questions

See ranked opportunities

See the methodology work on a real market. This opens the workspace with a research-heavy starting point so you can test the ranking flow on your own niche.

Run a research-backed scan →