Product

Santosh Bishnoi
Founder & CEO
10 mins

The comparison everyone reaches for is the dot-com bubble, and almost everyone reaches for the wrong half of it.
What people picture when they say "dot-com" is the wreckage: the shutdowns, the down rounds, the layoffs, investors who suddenly cared about revenue. But that was the aftermath, the 2000–2002 shakeout. The bubble's actual final months, late 1999 into March 2000, looked nothing like wreckage. They looked like the best party in the world. The Nasdaq rose about 86% in 1999 and kept climbing. The IPO window was thrown wide open: 484 U.S. IPOs in 1999 at an average first-day gain near 70%, and forty-eight companies doubled on their first day of trading in the first quarter of 2000 alone. Fourteen dot-coms bought Super Bowl ads at roughly $2.2 million each. AOL used its inflated stock to acquire Time Warner for about $163 billion. "Eyeballs" were quoted as if they were revenue. Everyone half-knew it couldn't hold, and nobody wanted to be the one who left early.
Mid-2026 is not that. It is the other half, the part that came after the peak, when the financing cycle reversed and capital remembered how to be skeptical. CB Insights and Gartner put the AI startup failure rate at roughly 80% by the end of 2026, above the 70% that has always been normal for tech. The pricing tells the same story more precisely than any think-piece can: a vertical AI company at $500K of ARR that commanded a 15x revenue multiple in 2024 now commands about 8x if it is a thin wrapper on someone else's model, and about 25x if it owns a real workflow. The market has relearned the one skill it forgets at every peak: telling a business from a story.
That distinction changes the conclusion. If this were the melt-up, the right move would be to wait for the air to come out. It is not the melt-up. It is the clearing, and the clearing is when durable companies get built cheaply against weak competition while demand keeps rising underneath. Eighty-eight percent of organizations already use AI in at least one function; the agent market sits around $8.5 to $11 billion in 2026 and is heading past $50 billion by 2030. What is dying is not demand. It is undifferentiated supply. The companies that survive a shakeout share one property, which is that they own something a competitor cannot acquire by signing up for an API. The whole question of what to build reduces to finding that property and standing on it.
Spend any time with what gets sold as "agentic" and you notice almost all of it collapses into two shapes, neither of which feels like the thing the word promises.
The first is the co-pilot. It drafts, suggests, and hands control back. You are still doing the work; it is sitting beside you doing a faster first pass. It is billed by the seat, which is honest about what it is, because it is selling you assistance rather than completion.
The second is the scheduled automation, the cron job with a language model bolted on. An event fires, the system performs one fixed task, and it stops. It feels autonomous for exactly as long as the trigger and the task stay identical — which is to say, until anything changes. It has no standing goal, no memory of last week, no view on what should happen next.
What neither shape does is hold an objective over time and act on it without being poked. That gap is real, and the dissatisfaction it produces is a correct reading of the landscape rather than impatience with it. The mistake is in the response. The natural instinct, once you have felt that gap, is to reach for more autonomy, more generality, more of the thing that would finally make the system feel alive. That instinct points up. The evidence points the other way, and the cleanest way to see it is to look at where autonomy actually lives in working companies, level by level.
It helps to grade agentic systems by how much they decide on their own. Five levels, from a tool that only suggests to a system that sets its own agenda:
Level | Behaviour | Representative companies |
|---|---|---|
L1 Assisted | Suggests; the human acts | Copilot, ChatGPT, Grammarly |
L2 Augmented | Acts on a trigger; the human reviews | Zapier, Make, Fin |
L3 Supervised | Runs a workflow; checks in at decisions | Cursor, Harvey, Agentforce |
L4 Autonomous | Sets sub-goals, adapts, runs largely unattended | Devin, Sierra, HappyRobot |
L5 Self-directed | Pursues standing goals proactively | none, commercially |
The interesting thing is not the taxonomy. It is what happens to the companies as you climb it.
L1 is where the largest revenues sit and the ground is least stable. ChatGPT runs around $25 billion annualized with a billion monthly users; Microsoft 365 Copilot has fifteen million paid seats and implied ARR north of $5 billion; Grammarly clears $700 million ARR. These are among the largest AI products on earth, and every one of them is a suggestion engine that defers to a human on the final action. The instability shows up one tier down in market size. Jasper built AI copywriting to roughly $120 million ARR in 2023 and watched it fall to about $55 million in 2024 once the foundation-model providers made the same output a free byproduct of their chat products. Tome raised $81.6 million, reached twenty million users, converted under 2% of them, and discontinued its core product. Wordtune was shut down; Copy.ai was absorbed. The common cause of death is specific and worth stating plainly: a single-feature layer over a model has no defense for the day the model vendor ships that feature natively, and the model vendor always ships it eventually, because your feature is their roadmap.
L2 is unglamorous and quietly profitable, which is the combination that survives downturns. Zapier runs about $310 million ARR on a near-bootstrapped balance sheet, moving over three billion tasks a month. n8n is growing 5.5x year over year with SAP investing at a $5.2 billion valuation. The clearest signal is the exit: Intercom's Fin, an AI support agent that resolves or escalates within a tightly bounded loop, was acquired by Salesforce for $3.6 billion in June 2026. That is a large outcome for an architecture that never pretends to be more than disciplined plumbing. The cautionary cases at this level are not about the concept but about the economics of running it. Transposit raised $50.4 million and closed at roughly $3.3 million ARR; IFTTT raised $63 million over fifteen years to reach $3.4 million. The pattern at L2 is that the work is real but the margins are unforgiving, and a good concept with bad unit economics still dies.
L3 is where the durable, fast-scaling, defensible companies of this cycle actually are. At this level the system runs a multi-step job end to end but is built to stop at the decisions that matter and let a human approve before proceeding. Cursor reached roughly $4 billion annualized in about three years, the fastest enterprise-software ramp on record, on a loop where the agent writes the code and the developer merges the diff; SpaceX announced a $60 billion acquisition of its parent in June 2026. Harvey drafts and redlines legal work and routes it to a partner for sign-off, at around $300 million ARR and an $11 billion valuation across more than a hundred thousand lawyers. Salesforce's Agentforce ($800 million ARR), Replit, Glean, and Ironclad, the last of which is growing 40% a year with no new capital since 2022, are all the same shape: genuine autonomy, fenced by a review gate, embedded in one industry's system of record. The instructive failure here is Builder.ai, which sold the perfect version of this pitch, that AI would build 80% of an application and humans would finish it, raised about $445 million, and turned out to be roughly 700 engineers writing by hand the code it marketed as generated. It went bankrupt in May 2025. The lesson is not that the L3 shape is weak. It is that the shape is attractive enough to raise nine figures on a fake, which tells you how badly the market wants the real thing and how exactly you have to deliver it.
L4 is real, and it is real specifically when it is narrow. At this level the system sets sub-goals, adapts mid-task, and runs largely unattended. Cognition's Devin decomposes an engineering ticket into sub-tasks and works through them, and its own CEO has been candid that it performs like a junior-to-mid engineer and earns its keep on long-tail maintenance, migrations and refactors and dependency upgrades, rather than self-started projects. Sierra resolves about 72% of inbound customer interactions end to end and bills per successful resolution rather than per seat. HappyRobot runs freight-broker phone calls for DHL, Ryder, and Flexport and grew revenue tenfold in nine months. Each one is autonomous inside a tightly drawn box, and each is careful to describe the box. Set against them is the part of L4 that did not survive, and it is mostly the part that went broad or went into hardware. Humane's AI Pin, pitched as an ambient assistant that would handle anything you asked, burned $230 million, saw returns outpace sales within months, and sold to HP for $116 million as every shipped device stopped functioning. Rabbit's r1 could not reproduce its own demo and had staff on strike over unpaid wages by late 2025. Adept raised $415 million to build a general computer-use agent and was reverse-acqui-hired by Amazon, returning investors their principal and nothing more. Read across the level and the split is unambiguous: the survivors were narrow, software-only, and precise about their limits, and the casualties were broad, often physical, and sold on general capability.
L5 has no commercial tenants at all. This is the level the word "agentic" implicitly promises: a system that wakes up, decides what matters, and pursues it without being asked. As of June 2026 the entire list of companies claiming it consists of Reflection AI, which has raised about $4.5 billion toward a reported $25 billion valuation while publishing no research, shipping no generally available product, and disclosing essentially no revenue; ai.com, which paid roughly $70 million for its domain and launched on a Super Bowl ad with a self-improving-agent claim no independent party has verified; and AutoGPT, the original viral demonstration of the idea, which proved brittle in production, never made meaningful money, and survives as an open-source curiosity. The most aggressive autonomy claim in the entire market belongs to the company that has raised the most money and shipped the least.
Lay the five levels next to each other and one relationship runs straight through them. The companies with real, compounding revenue are the ones that bounded their autonomy and said so. The company with the purest autonomy pitch has no product. Ambition and commercial reality run in opposite directions, and the higher you reach toward the thing that feels most like intelligence, the thinner the revenue gets and the louder the deck has to be to cover the gap. The instinct to climb, the one that the gap in Section 2 provokes, leads directly into the emptiest part of the building. The occupied rooms are two floors down.
It would be convenient if the ceiling were a model problem, because then it would be temporary and someone else would lift it for you with the next release. It is not a model problem. Three constraints hold it down, and none is about raw intelligence.
The first is that the binding constraint shifted from the model to the domain. The reason the failure statistics are so grim, 95% of generative-AI pilots showing no measurable P&L effect, only 39% of organizations reporting any EBIT impact, is almost never that the model was not capable enough. It is that real businesses run on definitions that are scattered and quietly contradictory. "Revenue" means one thing in the CRM, another in finance, a third in the billing system. A competent human notices the ambiguity and asks which one you mean. A model does not ask. It commits to whichever definition it encountered first and produces an answer that is fluent, confident, and wrong in a way nobody catches until it has propagated. Pilots rarely die of visible failure. They die of invisible error, and the only antidote is domain knowledge codified into the system, which is exactly the asset you cannot obtain through an API call.
The second is what researchers have started calling the coherence cliff. Current agents are, in one apt description, brilliant but amnesiac. Over a long task they lose the load-bearing state: the decision made twenty steps ago, the constraint mentioned fifty steps ago, the dead end they committed not to revisit eighty steps ago. The work on fixing this, proactive context folding that compresses a hundred turns into a few thousand tokens, memory systems that summarize and discard before the context window saturates, runtimes that persist agent state across days, is the genuine frontier, and much of it is being published and open-sourced. But until it is solved, open-ended long-horizon autonomy decays into loops and self-contradiction, which is the specific failure that put AutoGPT in the ground.
The third constraint is not technical, and for anyone deciding what to build it is the most useful of the three. Incumbents are structurally unable to build the autonomous version of their own products, because their revenue is denominated in seats and usage, and a system that genuinely removes the human removes the seat that the system was billing. Every SaaS incumbent therefore has a financial immune response to real autonomy. They can add copilots, because copilots sell more seats. They cannot ship the thing that deletes the seat, because it cannibalizes the number they report to the market. The first two constraints explain why L5 does not yet exist. The third explains why, in the zone that does work, the largest competitors will decline to follow you even when they can see exactly what you are doing.
The reframe that follows from all of this is the part most people get backwards. Narrowness reads like a compromise, the smaller ambition you accept because the larger one is out of reach. It is not the compromise. It is the mechanism.
Consider why no one can safely take the human out of a general agent. A general domain is unbounded, poorly instrumented, and saturated with exactly the definitional ambiguity that makes a model confidently wrong. Now draw the domain tightly: one buyer, one workflow, one system of record. Three things change at once. The context stops fragmenting, because a domain that small has a finite set of definitions you can actually pin down and encode. The long-horizon problem softens, because the task is now short and structured enough that the agent can hold its state to the end. And the edge cases become rare and recognizable, which means the agent can be built to escalate the specific situations a human should see and to handle the rest without supervision. Inside that box, removing the human stops being reckless and starts being the obvious design. The agent runs the loop; a person governs the perimeter. The autonomy that feels impossible in general becomes routine in the narrow case, and the narrowness is the reason, not the price.
This is what the survivors at L3 and L4 understood. Sierra did not build a general service brain and then constrain it; it built a bounded resolver and priced it on resolutions. HappyRobot did not build an AI employee and aim it at logistics; it built a freight-call worker wired into the dispatch stack. The narrow scope is the product, not a limitation of it.
It is also where the memory frontier stops being an obstacle and becomes leverage. The hard, unsolved version of long-horizon coherence is the open-ended one, an agent maintaining arbitrary goals across arbitrary tasks forever. You do not have to solve that. In a narrow vertical the state an agent must carry is bounded and largely known in advance, which means you can do the un-glamorous engineering that actually works: an explicit schema for the facts the agent must never lose, deterministic checks that catch when it has drifted, and a defined taxonomy of which situations get escalated to a human and which do not. That combination, a bounded domain, a hardened state model, and outcome-based escalation, produces something that holds a goal across days and acts on it unprompted. It is the thing the gap in Section 2 was asking for, and it is reachable as ordinary engineering rather than as a bet on a capability that does not exist yet. Almost no one is building it, because the field has split into chasing general autonomy, which is too hard, and shipping triggered automations, which is too shallow, and the productive middle is comparatively empty.
The system that results has a consistent silhouette. It is headless: the product is the outcome, a booked appointment or a filed claim or a reconciled ledger, not a chat window. It holds a standing objective rather than waiting for a trigger. It initiates and reports rather than waiting to be asked. It checks its own work, retries, and escalates only genuine ambiguity. And it is unbounded inside its domain precisely because the domain is bounded. That is the line between a smarter cron job and an autopilot for one job.
The thing worth building is a vertical agent that owns one painful, recurring, judgment-heavy workflow end to end for one specific buyer, and that becomes harder to replace the longer it operates.
Not a better model; the model is a commodity, and the gap between the best and the second-best closes a little every quarter. Not a horizontal agent that serves everyone, because more than 70% of horizontal agents never make it from demo to production, where real customer data is messy and a general system has no domain knowledge to survive the mess. The objective is to embed so deeply into one operation that you become its system of record — part of the work rather than a tool sitting beside it. That is the source of the defensibility, and it is worth being concrete about what defensibility consists of, because "moat" is the most abused word in the category.
There are four kinds, and a real company has at least two. The first is a proprietary data loop, where the system gets measurably better at the job from data nobody else can obtain. A prior-authorization agent for physical-therapy clinics is a clean example: after fourteen months of live submissions it knows which insurers approve which procedure codes at which rates for which diagnoses, knowledge that exists nowhere as a dataset and can only be accumulated by having done the work thousands of times. A competitor starting today cannot buy that; they can only begin their own fourteen months. The second is integration depth, a position wired into enough of the surrounding stack that ripping you out means a migration project rather than a cancellation. The third is workflow ownership, where you run the process itself, with the steps and the rules and the exceptions encoded, rather than supplying one feature inside someone else's process. The fourth is distribution into a niche, owning the channel to a specific buyer so completely that reaching those customers is itself the barrier. The number that reveals whether any of this is real is net dollar retention. Wrappers churn; companies that own a workflow run 120% and up, because the workflow expands and the data compounds.
Before any of that, an idea has to pass a buyer test, and the first criterion does most of the work. Is someone already paying a human to do this exact task? If they are, you are replacing a line on a payroll rather than manufacturing demand, which is the strongest signal available that the problem is worth money. The second criterion is timing: is there a 2025–2026 catalyst that makes now the moment rather than last year or next, such as voice-agent APIs cheap enough to make phone work automatable, the Model Context Protocol making integrations tractable, or the collapse in model prices changing the unit economics. The third is competition at the right resolution: not whether "AI for healthcare" is crowded, which it is, but whether your exact segment, insurance billing for mental-health practices under thirty clinicians, has fewer than five funded competitors, which it does.
The segments that pass tend to be the ones nobody puts on a conference stage. Contract review for solo law firms; government-bid hunting for mid-size contractors; no-show recovery for independent medical practices; front-office and billing for solo dental practices; claims triage for independent insurance adjusters; parts procurement for small manufacturers; freight exception-handling for 3PLs; AP reconciliation for mid-market finance teams. It is worth walking through one in full, because the shape only becomes convincing when it is specific.
Take the home-services contractor, the HVAC company or the roofer or the solar installer. The average HVAC business misses 30 to 40% of its inbound calls during peak season, and a missed call in that business is usually a missed job worth hundreds to thousands of dollars; the lost revenue from missed calls alone runs $50,000 to $200,000 a year for a single shop. The owner knows this and cannot fix it, because the fix is a full-time person answering the phone at 7 p.m. in August, and the economics of that hire are marginal. A narrow agent that owns this one workflow answers every inbound call, qualifies the job, quotes from the company's actual pricing, books the slot against the real dispatch calendar, and follows up on the leads that did not close, is not a copilot for the owner and not a triggered automation either. It holds a standing goal, keep the calendar full and the pipeline worked, and it pursues that goal across days without being asked. It accumulates a data loop nobody else has, namely which jobs in this trade and this region close at which quotes and which follow-up timing recovers which leads, and it wires into the dispatch and pricing systems deeply enough that removing it would leave a hole in daily operations. And because what it produces is booked, paid jobs, it can be sold and priced on that outcome rather than on a monthly seat.
That last point is the wedge, and it is structural rather than stylistic. AI is the delivery mechanism, not the product; what you sell is the outcome the AI makes possible, the booked appointment or the recovered revenue or the filed claim. Pricing on outcomes does two things that pricing on seats cannot. It proves you actually removed the human rather than merely assisting one, because nobody pays per resolved outcome for a tool that still needs a person to do the resolving. And it walks straight through the third constraint from Section 4, the seat-cannibalization conflict that prevents every incumbent from following you, because you have no seat to cannibalize. The thing that paralyzes them is the thing you are built on.
The graveyards in Section 3 are not decoration; they are the photographic negative of the prescription, and each one names something to avoid.
Do not build a thin layer the platform will absorb. This is the Jasper and Tome death, and it has a one-line test: if a foundation-model provider shipped your feature natively tomorrow, would you still have a business? If the honest answer is no, you have a wrapper, and the platform's roadmap is your obituary.
Do not chase general autonomy or ambient AGI. This is the Reflection AI and AutoGPT pattern, maximum narrative and minimum revenue, and it is the most seductive form of the climbing instinct precisely because it dresses up as the dream rather than the trap. Technically thrilling and commercially worthless are entirely compatible states. Worth building means a buyer pays repeatedly and cannot leave, which is a different test than whether the demo is impressive.
Do not build hardware unless you are exceptionally well-capitalized, because the L4 casualties are disproportionately physical: Humane at $230 million, Adept at $415 million, Rabbit on unpaid wages. The economics of hardware plus a frontier model are brutal and unforgiving of the iteration that software takes for granted.
Do not sell tools to the companies that are dying. No-code agent builders, multi-agent orchestration frameworks, agent-observability and evaluation platforms, data-labeling shops: in a shakeout the infrastructure vendors fail early, because their customers are the wrapper startups going bankrupt, and a supplier to bankrupt companies is on the same clock.
Do not enter the commodity zones at all without an unfair edge: general SDR and cold-email agents, general coding assistants, low-touch support chatbots, RAG-only document Q&A, horizontal RPA, and any "AI agent for X" with no real vertical depth underneath the slogan.
And do not build anything the buyer can comfortably live without. Mission-convenient dies in the first downturn; only mission-critical renews through one. If losing your product would not hurt the customer inside the current quarter, you do not have a company, you have a feature waiting to be cut from a budget.
A single checklist filters most of this. Run an idea through it honestly. A wrapper clears two or three of these; a real company clears eight or more.
Is a buyer already paying a human to do this exact task?
Is it mission-critical, not merely convenient — would losing it hurt this quarter?
Is the exact segment underserved — under five funded competitors — even if the category is crowded?
Is there a 2025–2026 catalyst that makes now the moment?
Can you accumulate proprietary execution data that compounds over time?
Can you own the workflow and become the system of record — not just add a feature?
Can you price on outcomes rather than seats?
Is the domain narrow and instrumented enough that full autonomy is safe?
If a foundation-model provider shipped your feature natively tomorrow, would you still have a business?
Does it avoid hardware and pure infrastructure tooling?
The starting move, once an idea clears the checklist, is to find the workflow before building the agent, and the highest-signal place to look is wherever a business is currently paying people to do something repetitive and judgment-heavy. That is why a particular path works unusually well: if you run, or can partner with, a services business in some industry, the recurring tasks you already perform by hand for clients are pre-validated product candidates, because the buyer test is already satisfied, someone is paying for the work today.
Pick the narrowest viable vertical rather than the largest available market, because narrowness is what makes the autonomy safe and what gives you a data loop nobody can match. Build that loop from the first deployment, since every action the agent takes is training data the next competitor cannot retroactively acquire, and twelve to eighteen months of it becomes two of your four moats. Apply the long-horizon engineering from Section 5 where it earns its keep, so the agent holds the goal across days and initiates rather than waits. And land the first one or two customers as design partners who will tolerate the rough edges in exchange for the outcome, deliver that outcome, turn it into a reference, and only then productize and repeat. The services-to-product path is well-worn for a reason: it funds itself, it keeps you pressed against real pain, and it gives you distribution before you have a finished product.
The model is a commodity, and intelligence is no longer the scarce input. What is scarce, and therefore what is worth owning, is codified domain knowledge, deep integration into a real operational stack, and proprietary execution data that compounds over time. The autonomy ladder demonstrates this the expensive way: ambition and revenue run in opposite directions, the top level is empty while the narrow rooms two levels down hold the durable companies, and the most over-funded pitch in the market is attached to the least product. The move that fits this moment is not to climb. It is to take one narrow, painful, judgment-heavy workflow that a specific buyer already pays a human to do, and build a headless, goal-holding, proactive agent that runs it end to end and becomes harder to remove every month. The autonomy that feels out of reach in general becomes safe and obvious once the domain is narrow enough, and the long-horizon-memory work that is hard in the open-ended case is tractable in the bounded one. Sell the outcome rather than the seat. Avoid wrappers, frameworks, hardware, and the revolutionary-but-unsellable trap. Done that way, the shakeout is not a hazard to wait out. It is the clearing that gets made, every cycle, for whatever gets built next.
Federal Reserve Bank of St. Louis (FRED), "Nasdaq Composite Index." fred.stlouisfed.org/series/NASDAQCOM
Jay Ritter, University of Florida, "IPO Statistics." site.warrington.ufl.edu/ritter
NBER, "DotCom Mania: The Rise and Fall of Internet Stock Prices." nber.org/papers/w8630
Colrows, "From Copilots to Autonomous Companies: AI-Native Operations." colrows.com
dev.to, "Why So Many AI Startups Fail" (citing MIT and McKinsey figures). dev.to/praveenax
Sharad Jain, "Brilliant but Amnesiac: The Coherence Cliff in Long-Horizon AI Agents." sharadja.in
arXiv, "Memory for Autonomous LLM Agents." arxiv.org/html/2603.07670v1
VC Cafe, "Vertical AI in 2026: The Good, the Bad and the Ugly." vccafe.com
Preuve, "AI Agent Startup Ideas 2026." preuve.ai
Ciela AI, "The 12 Most Profitable AI Automation Agency Niches in 2026." ciela.ai
TechCrunch, "Salesforce acquires AI customer service platform Fin for $3.6B." techcrunch.com
Sacra, "Cursor." sacra.com/c/cursor
Sacra, "Sierra." sacra.com/c/sierra
The Register, "Builder.ai insolvency." theregister.com
TechCrunch, "Humane's AI Pin is dead as HP buys startup's assets for $116M." techcrunch.com
Turing Post, "Inside Reflection AI." turingpost.com
Product
What to Build in AI Agents in 2026 Copy

Santosh Bishnoi
Founder & CEO
