Contents
  1. The three ways ideas happen
  2. The walkthrough, one real run
  3. What the machine actually hands you
  4. Reflections
  5. The loop stays open

The pitch was working. Early traction with the account, a champion forming, the kind of signal that makes a team lean in. Then the buying committee started naming established players, and a proper look around confirmed it. We had walked into a seriously crowded market.

A crowded market, mapped.

The investor’s read was blunt. “I’m not confident we can enter this.”

Fair. Neither was I. I’ve worked difficult buying committees before. I’ve even built simulators of them. But this was something else. So I started running FlowScout, an open tool I built for walking research papers, hunting for patterns that could form the core of a product bet. Promising threads emerged, and with them an intuition for what the business could become in its fully evolved form. But everything I had came from the tech side. Nothing from the other end. No market pull. Ammo for a debate, not conviction for a bet.

So I resurfaced an old trick I picked up working with Mike Taylor, comparing two things at a time and letting the labels emerge from the contrast, and rebuilt it as an agentic system. Multi-agent workflows. Dozens of evals. Scripts that turn rented LLM intelligence into something you can trust at scale. It read the whole market the way a research team would, then handed me the uncomfortable part. A few players were already positioned exactly where the wave would land, and the honest read was that backing them beat building against them. Six months of learning arrived in days.

The investor pulled out. If you’re going to fail, fail fast.

Which left me where many builders, and plenty of investors, eventually land. Staring at a crowded market, holding opinions where conviction should be. Most people read a crowded market as bad news. I’ve come to read it as a dataset. Money is provably moving, buyers are provably buying, a mature pool of dissatisfied customers is forming, so the biggest uncertainty in entrepreneurship, does anyone pay for this, is already answered. What’s missing is an instrument for conviction. Something that shows a builder where their unique take survives, and shows an investor which players are worth backing.

No one hands you that. So I threw attention and token budget at it. Here’s what came out.

The three ways ideas happen

There are only three moves in the idea game.

Induction spots a repeating pattern across specific observations and generalizes it into an opportunity. Scrape a thousand G2 reviews, notice complaints about simulation builders growing ten percent month over month, conclude there’s a growing market for simulation tools.

Deduction takes a principle proven somewhere and applies it to ground where it hasn’t been tried. On-demand apps let people summon services instantly. Uber proved it for rides. Apply it to food and you get DoorDash.

Abduction notices something that should exist and doesn’t, then asks what best explains the gap. Physics allows a fast, long-range electric car, yet none exists. So the barrier must be industry assumption, not technology. That explanation, held against all consensus, is Tesla.

The trade-offs between them aren’t arbitrary. Two old frameworks predict them. The efficient market hypothesis says visible demand is priced-in demand. If you can scrape an opportunity from review sites, so can everyone else, which is why inductive bets reach revenue fastest and compete away their margins fastest too. Markets only misprice what they can’t yet see, and abduction is by definition a bet on an explanation the market hasn’t noticed. That’s why it carries both the highest failure rate and the only shot at monopoly returns. The adoption curve tells the same story in time. Inductive bets sell to a mainstream that already knows what it wants, so money arrives this quarter. Abductive bets start at the far left of the curve, selling to innovators who barely exist yet, so money arrives in years, if at all. Deduction sits between. The principle has crossed the chasm somewhere, but the new domain has to walk the curve from the start.

Founders migrate between modes as ideas make contact with reality. Instagram began as deduction (check-ins were proven, Burbn applied them with a twist) and was saved by induction when usage data showed everyone ignoring everything except photo filters. Slack began as deduction (games were proven, Glitch died anyway) and was saved by abduction. Why does the internal chat tool we can’t stop using exist nowhere as a product?

Now the part that matters for anyone renting intelligence by the token. AI is great at induction. It’s getting better at deduction. It is structurally bad at abduction, the one mode that makes monopolies. And for all three, the conviction to act still has to come from a human.

You can’t rent abduction. But you can farm its raw material. Run induction at a scale no analyst team can, until the anomalies fall out.

The walkthrough, one real run

Here’s what that looks like in practice. I aimed the system, by now a proper swarm of agents with quality gates between them, at a consumer market I had a personal thesis about, apps that help parents capture family moments and guide their child’s development. What follows is the actual run, from its own report.

The scrape. Multi-mode agent search swept the category and came back with 119 candidates. Adversarial vetting killed the vaporware and left 111 real companies, every claim tied to a source URL.

The market universe. 111 companies as nodes, colored by lane. The legend shows a color that never appears.

Breadth is cheap now. Belief is not. Which is why the next pass matters more.

The receipts. The hard part of running a swarm isn’t getting output, it’s getting output you can bet on. So the system is built to distrust itself. Independent agents check each other’s work. Quality gates block anything below threshold instead of waving it through. Claims that arrive without a source get recorded, never believed. Adversarial audit passes attack the finished result the way a skeptical partner would. That machinery is invisible in the final report, which is the point. What survives it is what you’re looking at. The most useful thing my agents did all week was distrust each other.

The feature map. Forty-five capabilities coded across all 111 companies, every cell coded twice by independent agents with the agreement rate measured and reported. It came out at 0.87, which is the difference between a listicle and an instrument. An instrument tells you how much to trust it.

The capability matrix. 45 capabilities coded across 111 companies, twice.

What winners carry. Statistics over the matrix, false-discovery controlled so cherry-picks die. The signals that survived are boring-looking, which is how you know they’re real. Winners disproportionately carry clinical backing, records that whole families can share, and free-writing space. Products led by a chat interface correlate with losing.

What verified-traction winners carry, false-discovery controlled.

The anomaly. Then the interesting thing fell out. Two capabilities, each common across the market, should appear together in about seven companies if features combined at random. The actual count across all 111 is zero.

The empty quadrant. Two common capabilities that never co-occur.

That’s a structural zero, and it’s precisely the shape of abductive raw material, something that should exist and doesn’t, demanding an explanation. A structural zero is how a market hides either a graveyard or a fortune, and no amount of model intelligence can tell you which. The induction engine found it. Explaining it was my job.

The stress test. So I formed my explanation, and instead of falling in love with it, I made the machine attack it. My take became six competing offers, put in front of 45 synthetic parents cloned from a calibrated research panel. Forced trade-offs first, then priced choices against two honest anchors, “keep doing what you already do” and “just use free AI and your camera roll.” The results were humbling in the right ways. The fused offer beat every incumbent shape. Removing its core component collapsed it. And inertia took 56.5 percent of all choice, a reminder that the real competitor is nobody. Then a red-team of mental models (inversion, incentives, base rates, second-order effects, trade-offs) attacked the recommendation itself, and partially won. It found the deepest risk somewhere I hadn’t been looking.

The simulated shelf. Share of choice across six offers and two honest anchors.

The machine ranks and refutes. It still doesn’t believe. Synthetic panels rank options. They do not price reality. That label is printed on every chart, verbatim, because an instrument that hides its own limits is just a slower way to lie to yourself.

What the machine actually hands you

Not an idea. What came out the other end is a testable explanation with kill criteria attached. Pre-registered falsification tests, probability weights on the outcomes, and a rule agreed in advance for what each result means.

The decision, with its falsification battery pre-registered.

That last part matters more than any finding, because of what a business idea actually is. It isn’t one idea, it’s a stack of them. Buyer, budget line, form factor, price, channel, timing, brand. Every dial has a thousand settings, almost all of them dead, and ideas give no partial credit. Get one dial wrong and the whole thing reads as “no demand” even when every other dial was right. Instagram already existed inside Burbn. It looked like a failing check-in app until one variable moved. That makes idea search a combination lock, not a hill climb, and it’s why commitment is epistemic rather than motivational. Switch bets every few weeks and you reset the experiment before it can report. A bet abandoned early is indistinguishable from a bad bet.

The swarm doesn’t remove the uncomfortable part. What it changes is the starting position. For the first time, I get partial credit before betting. The matrix shows which dial combinations exist and win across a whole market. The simulation reads individual dials before a euro moves. I still commit. I just commit smaller, sooner, and with the meaning of the result agreed up front. Turn a dial, or drop the lock.

Reflections

You can’t rent abduction, and neither can anyone else. AI reasoning is a commodity available to every competitor at the same price, so it can’t be a moat. The moat is orchestration. What you point the commodity at, how the passes check each other, and how fast the loop closes. The asymmetric upside lives exactly where the commodity can’t go alone.

Verification is the craft. Every failure in this run was a trust failure. Wrong numbers, degraded outputs, confident nonsense. Every fix was a verification structure. Evidence ladders, blind dual coders, quality gates that block instead of warn, adversarial review of the final answer. Slop is a design failure, not a model failure.

Partial credit exists now. The market grades pass-fail. The swarm grades dial by dial. That changes the economics of idea search more than any individual insight ever could.

The run doesn’t care which chair you sit in. The same machinery that killed our build idea identified who was worth backing instead. Due diligence turns out to be idea search viewed from the other side of the table, which I learned by watching a swarm talk me out of building and into wishing I could write a check.

The loop stays open

Full disclosure. I’m still interpreting what this run produced. The recommendation survived its own red-team, but conviction is the one thing the machine can’t hand me, so I’m fine-tuning the system and running more simulations until I land on a bet I’m willing to hold. I’d rather say that plainly than dress a work-in-progress up as a verdict.

That’s what I’m building toward, a suite of market-intelligence tools whose whole job is to find my next bet. In the meantime, the first of those tools works for paying customers. If you’re staring at a market you can’t see deeply enough to bet on, whether you’re building into it or writing checks into it, I can point the same swarm at yours. It’s called Scout. Book a demo and I’ll walk you through the live report from this run.

If this was useful