How to Scope an AI MVP That Actually Ships
Most AI MVPs fail on scope, not technology. The specific cuts that get a product in front of real users in weeks instead of quarters.
The failure mode for AI products is rarely that the model isn’t good enough. Usually six months went into building something broad and shallow, and the first real user contact happens only after the money is mostly spent.
Scope is the whole game. Scoping an AI product has a few specific traps that ordinary product scoping does not.
Cut to one user, one job, one path
The version that ships solves one problem, for one type of user, one way.
Not one problem with three input methods. Not one problem for two user types who mostly overlap. One path, end to end, deployed, with real data.
The instinct to resist is breadth as insurance: covering more cases because you are not sure which one matters. It feels like risk reduction. It is the opposite. Every additional case multiplies the surface you have to make reliable, and you learn less from a broad prototype than from a narrow thing that people actually use.
Find the load-bearing screen
Every product has one screen where the real complexity lives. The review queue. The comparison view. The thing that shows the AI’s output next to the evidence for it.
Build that one first, before auth, before settings, before onboarding. It is where all the hard design questions are, and it is the screen that decides whether the product is any good. Everything else is a form.
Scope the AI part narrowly
This is where AI projects diverge from ordinary ones.
Name the exact task. Not “AI-powered analysis,” but something closer to “given a document of this type, extract these eleven fields into this schema.” Vague AI scope produces unbounded AI work, because nobody can say when it’s done.
Decide what wrong looks like before you build. If you cannot describe a wrong output, you cannot evaluate the system, and you will ship on vibes.
Assume it will not be perfect, and design the interface around that. The single biggest factor in whether an AI feature works in practice is what the interface does with uncertainty. Show the evidence. Make correction cheap. Let the user override without friction. A system that’s right seventy percent of the time with a two-second fix beats one that’s right ninety percent of the time with no way to intervene.
Pick the boring model first. Start with the cheapest one that plausibly works, and measure it. Teams routinely start with the most capable model, build everything around its latency and cost, and then discover the cheap one was fine all along.
What to leave out of version one
Onboarding flows: set the first accounts up manually. Watching someone use the product with you teaches more than any wizard.
Settings and configuration: every option is a decision you haven’t made yet. Make the decision, hard-code it, revisit later.
Admin panels: query the database. You are the admin.
Notifications: almost never load-bearing in a first version, and always more work than it looks.
Mobile, unless the job genuinely happens on a phone. Integrations, beyond the one that’s actually needed (manual import is fine at ten users). And fine-tuning: prompt and context engineering gets you further than most teams expect, without locking you to a model.
What has to be in version one
Some things don’t survive being cut, and skipping them is the other way an MVP fails.
Real authentication, not a shared password. It leaks into the data model too deeply to retrofit comfortably.
The real data model. Interfaces are cheap to change later; schemas are not, so this deserves genuine thought up front. If the product will ever serve more than one customer, that thought has to include tenancy from row one, not after the first outside customer signs up: see multi-tenant SaaS architecture.
Evaluation, for any AI component, from day one. You cannot bolt this on later, because by then you have no baseline to compare against.
Logging of every AI input and output. It is your debugging record, your eval set, and your evidence when someone disputes an output.
A correction path wherever the AI produces something a user relies on. They need a way to fix it, and you need to capture that fix.
A realistic timeline
For a focused AI product with one path and one user type, six weeks is a reasonable target. The domain model, the hard screen as a clickable prototype, and an evaluation set come together in week one. Weeks two and three get the core path working end to end against real data, deployed. Weeks four and five are for measuring the AI component against the eval set and getting guardrails in. By week six there are real users doing real work, with instrumentation switched on.
Anything materially longer than six weeks usually means the scope was not cut hard enough. The honest fix is to cut again, not to extend the timeline.
What comes after
The point of shipping narrow is that the roadmap stops being a guess. Six weeks of real usage tells you which of your assumed priorities were wrong, and in our experience roughly half of them are.
Build the next thing from what people actually did, not from the backlog you wrote before anyone had used it.
We scope and build AI products on this shape of timeline. See how we work, or read about evaluating the AI parts properly.