Roadmap: built in public
1mil.app is one founder building an AI tool that finds and pressure-tests business ideas. Here is how it got here, what we learned, and what we are building next. If you want to shape it, the form at the bottom goes straight to us.
Why we publish this
Most tools hide how they are made. We do the opposite, for one reason: the product is about telling you the truth about an idea, including the hard parts. A roadmap that only lists wins would contradict that. So this page includes the things we tried and dropped, and the moment we found a real gap in our own scoring and fixed it.
The story so far
-
March 2026
Formed the company
Registered Million Opportunities LLC in Wyoming and set up the foundation, so this could be a real product rather than a weekend script.
-
April 2026
Built the scanner and shipped the workspace
The core pipeline: research a topic with live web grounding, generate candidate opportunities, then score them. Paired with a redesigned landing page and workspace.
-
May 2026
Made it findable, and made it tougher
Built out the content surfaces (ideas by role, validation guides, comparisons), fixed dozens of technical SEO issues, and unified the brand. When a model we relied on was retired without warning, we hardened the pipeline against the next outage.
-
June 2026
Accounts, billing, and lifecycle email
Added Stripe billing, Clerk accounts with an optional sign-in step (kept behind an instant kill-switch so it can be reversed in seconds), funnel analytics, and a consent-first email system on Resend.
-
June 2026
Tried a billing rail, then dropped it
We explored Whop as a distribution and billing channel and shelved it. It was a payment rail, not a discovery feature, and it pulled focus from the core job. Not everything we start ships.
-
June 2026
Added a winnability rating
A second rating beside the business score: whether a solo builder can actually capture an idea, not just whether people pay for it. More detail below.
-
June 2026
Fixed competitor research that was too optimistic
Our scans were calling markets "open" when they were not. We found ideas reported as having little competition that actually had a dozen or more funded tools already serving them. We rewrote how the scanner checks competition. It now searches for the real alternatives, including platform-native features and do-it-yourself workarounds. Then it names them, and states the gap against those named players instead of claiming an empty market.
-
July 2026
Stopped the scanner from defaulting to "build an AI tool"
The scanner began as an AI-opportunity finder, so its research step was built to look for "AI tools, software, and SaaS" for every topic. That fit when the whole product was about AI. But broader topics, like a snack recipe or a local product, kept returning AI apps that did not belong. So we broadened the research to look for whatever actually serves a topic, including physical products, local services, and offline businesses. On genuinely AI topics it still surfaces AI tools, because that is what fits. The topic decides now, not an old default.
-
August 2026
Stopped taking the model's word for it
Until now, when the scanner reported that a competitor charges $49 a month, that claim came from the model and went straight onto your screen. It was usually right. Usually is not good enough for a number you might price against. So we built a checking layer that reads every scan before you do. It gives the scanner one supervised attempt to fix what gets flagged, and it shows its working instead of asking you to trust a single number. More detail below.
-
August 2026
Measured our own rating, and published the miss
We ran the same 25 ideas through the winnability rating twice and compared. The runs agreed at a rank correlation of −0.05, which is no agreement at all. We stopped displaying decimals, rebuilt the display around repeat checks, and published the full audit instead of fixing it in the dark.
-
August 2026
Briefs now match what an idea actually is
The old build brief assumed every idea was a software product: it handed a nurse's productized-service idea a file structure and a Stripe integration. Brief generation now runs server-side and reads what the idea is first. Software gets a build plan; a service gets a landing page, a payment link, and a first-customer channel; content and templates get their own shapes.
-
August 2026
Showed the machinery, and learned your language
The homepage now shows the five-stage discovery pipeline itself, in place of marketing framing. We rebuilt the FAQ from questions real builders asked on our Indie Hackers launch post. And scans now answer in the language you write them in, down to the competitor research. A Russian scan surfaced Russian competitors with ruble pricing, not translations of American tools.
-
September 2026
Started publishing a monthly edition
Every month we mine public threads where practitioners describe problems their tools do not solve, and publish the resulting product ideas by market: 371 ideas across 26 markets in the first edition, each one quoting the practitioner and linking the discussion so you can judge the evidence yourself. They are hypotheses, not validated businesses, and the pages say so. We also capped free scans per day after someone farmed the anonymous trial by resetting cookies.
-
September 2026 (latest)
One score instead of two
We measured which parts of the business score wobble between runs: over half the noise came from three judgments the model made once, unchecked, during generation. The sampling machinery we built for the winnability rating now defends the score you actually see. Those three judgments are sampled three times and averaged, and the winnability label itself is retired. Two scoring vocabularies on one card confused every reviewer who audited us; the facts behind the label, such as a free alternative existing or a funded incumbent shipping, stay visible in the competitor list. Scores now display at one decimal with the measured re-run variance (about ±1.5) published on the audit page, which also tells the retirement story in full.
A closer look: the Winnability rating (2026, retired)
We ran our own top-rated ideas through a hard willingness-to-pay check, the kind a skeptical investor would. The result was uncomfortable: most of them had real, provable demand and were still bad bets for a solo builder, because a free tier, a funded incumbent, or the platform vendor had already closed the gap.
That exposed a gap in our own scoring. Our score measured demand (do people pay for this?) but was blind to winnability (can you, without an audience, actually capture it?). Those are different questions, and we were only answering one.
So we added a second rating beside the business score. For each top idea it checked four things:
- whether anyone already paid for this job;
- whether a free or native option set the price floor;
- whether a funded player or the platform had already shipped it;
- whether the moat was something code could build.
It showed the single biggest reason an idea was hard, in plain language. The idea generation did not change. We only added the second number.
A year of measurement later, we retired the rating itself: the sampling machinery that made it stable now steadies the main score, and the facts it surfaced live on in each card's competitor list. The audit page tells that whole story, including the miss.
A closer look: checking our own output
This product exists to tell you the truth about an idea. That obligation points inward too, because the scanner is built on a language model, and language models produce confident sentences whether or not the facts behind them hold.
The early version handled this the way most AI tools do: write careful instructions, ask the model to police itself, trust the result. That is a reasonable place to start and it got us a long way. It also has a ceiling. You cannot instruct your way out of a wrong answer, because the instruction and the answer come from the same place.
So we changed the arrangement. The model proposes, and code verifies. A checking layer now reads every scan before you see it, and it does not ask the model whether the scan is right. It compares the claims against the research the scan actually gathered:
- A competitor price survives only if the search genuinely found it. Otherwise it is marked as not found rather than shown as a fact.
- A price found for one company cannot silently attach itself to another.
- A market gap has to name a specific capability. If the sentence would read as true with any other product swapped in, it is flagged as empty.
When something is flagged, the scanner gets one attempt to fix it, and the same code that caught the problem has to accept the repair before it reaches you. Failed repairs leave the flag in place. We keep the original wording either way, so the record of what was changed does not disappear.
None of this makes the scanner smarter. It makes it accountable, which we think matters more when you are deciding where to spend a year.
What is next
- Re-measure the business score's run-to-run reliability now that the sampled re-judgment is live, and let the result set the display precision.
- Check the score against reality, not just against itself: do the ideas people export and pursue actually score higher?
- Keep publishing the monthly edition, and start showing which of those ideas survive scrutiny and which are tempting traps.
- Keep tuning the scoring against real outcomes, not intuition.
What we build next depends partly on you.
Tell us what to build
Have an idea, a complaint, or a feature you want? It goes straight to the founder. Email is optional, only add it if you want a reply.
Want to try it? Run a scan and get ranked, evidence-backed directions.
Open the workspace →