Cancel AI PoC or continue: The quickest way to make a decision is a clean baseline
You can spend weeks on an AI PoC and still not know at the end whether it was “good”. Not because your team works poorly, but because you work without a starting point. Without a baseline, every number seems arbitrary, every optimization seems like activism, every discussion like gut feeling against gut feeling. This is exactly where the lever lies: those who create a baseline early make decisions more quickly, invest more specifically and end PoCs before they become a bottomless pit.
Here's the gist: A PoC rarely fails because of the model. It fails because no one has clearly defined how success is measured and the status quo is not quantified.
Why a baseline will save (or quickly end) your AI PoC
Many companies start a PoC with a vision of “perfect”. The result: disappointment is inevitable because no one knows what “better than today” actually looks like. Alex Grimm puts it in a nutshell: If you don't have the comparison, you're "swimming" a bit.
Cancel or save AI-PoC: Why a baseline determines success or failure
You spend weeks in an AI PoC, everyone is busy, the results are “kind of okay,” and yet the feeling remains: We don’t know whether things are going well here or whether we’re just reassuring ourselves. It is precisely at this point that the most costly mistakes happen. Not because the model is “too bad,” but because no one has clearly defined how “good enough” is actually measured. The lever that creates the most clarity is surprisingly unspectacular: a resilient baseline that you can reach quickly and consistently use as a guide.
Making AI-PoC successful: First baseline, then optimization
Many teams start a proof of concept with an ideal goal in mind. The problem: Without a benchmark, every result seems arbitrary. A baseline is exactly this standard of comparison. It answers three questions that decide whether to continue or stop:
- Where are we really today (time, costs, quality, risk)?
- What is “better than before” and how much does “better” have to be?
- Which change had which effect?
In practice, this baseline is often missing because the current process works informally: gut feeling, empirical knowledge, Excel, short coordination between departments. This can be “good enough” operationally for years, but it is poison for an AI PoC. Because without a baseline, you can neither prove progress nor clearly argue why you need more time, more data or a different approach.
Here is the crucial change in perspective: A PoC is not just a model test. It is often the first moment in which a company structures its process in such a way that it becomes measurable. This measurability is already a result, even if the model is not yet ready for production.
Practical Rule: Define the baseline in business terms, not ML terms. Instead of starting immediately with Accuracy, F1 or MAE, first clarify: How long does the process take today? How much does it cost? Where do errors occur? What subsequent problems are associated with this? Then you translate that into technical metrics.
When you should abort an AI PoC: There must be a signal after a few iterations
A PoC should not feel like an endless loop. At the same time, “abort” is often emotionally charged because teams interpret the word as failure. Operationally, however, you need a sober mechanism that says: We have seen enough to decide.
A helpful approach from project practice: work in short iterations and force yourself to reach a first number early on. Not because this number is “correct”, but because it gives you orientation. If you have defined a target (for example an error measure below a threshold), then the distance to the target is your early warning system:
- If your first model is close to the target, you have the wind at your back. Then optimization is worthwhile.
- If it's far off the mark, you most likely need different requirements: more data, different features, a different measurement concept, sometimes even a different problem formulation.
Timing is important: Many teams waste too much time analyzing data before they even build a first model. This feels thorough, but delays the only insight you really need: Is there a realistic signal that this use case will become viable with the existing data in the foreseeable future?
Practical Rule: Plan for a maximum of three serious iteration loops before making a hard interim decision. The following applies: Either you are in an area where optimization takes effect with reasonable effort, or you need structural changes (data, process, scope).
The most common reason for “failed” PoCs: There is a lack of process and data capture, not AI competence
Many PoCs don't run into a technical wall, but rather against the reality of data collection. A typical pattern: In the PoC it becomes clear that data is recorded at too high a frequency, important influencing factors are missing or the data quality fluctuates because sensors, systems or processes change.
This feels like a step backwards in the project, but it is actually the most valuable insight that a PoC can provide: what homework is necessary so that the problem can be solved at all.
If you measure three times a day in a laboratory, but you need every 15 minutes to make reliable forecasts, then that is not a model problem. Then it's a measurement and process problem. And yes, that can mean that the PoC cannot be achieved in its original objective. Still, you haven’t achieved “nothing.” You have created clarity about what needs to change for a later solution to work.
Practical rule: Always formulate PoC results in two tracks:
- Result for the use case (how close can we get to the goal with today's setup?)
- Result for the organization (which process and data issues do we need to solve to make the goal achievable?)
This way you prevent “not ready for production” from automatically being read as “worthless”.
Expectation management in AI-PoC: The invisible budget killer
The quickest way to ruin a PoC is a creeping expectation gap. You can tell by the fact that in the end everyone is disappointed, even though everyone has worked hard. This often happens when “PoC” means something different in the minds of the stakeholders than in the minds of the data science side.
A PoC is proof of concept. It proves whether an idea is fundamentally technically feasible or not, and what requirements apply to it. It is not a promise of a finished product that can go live immediately. If this distinction is not crystal clear, two typical effects arise:
- The scope grows through special requests and detailed discussions (“special curls”) until time and budget no longer fit.
- The evaluation becomes unfair: one judges against “perfect”, the other against “better than before”.
The problem is particularly exacerbated with generative AI because success often cannot be described with a single number. You need additional evaluation, domain expertise and clear criteria for what “good” means in context.
Practical rule: Before starting the project, specify in writing:
- Goal of the PoC (what is proven?)
- Success criteria (quantitative or qualitative, but verifiable)
- Which is explicitly not part of the PoC
- Decision logic at the end (Go, Pivot, Stop)
That sounds like a formality, but in reality it saves most discussions.
Invest time and budget correctly in AI-PoC: Fast orientation beats perfect analysis
When teams talk about PoCs, many think about model training and algorithms. In practice, a large part of the time goes into communication, data acquisition, data understanding and preparation. This is not a flaw, but the reality of corporate data.
A robust rule of thumb based on project experience: In a PoC with, for example, 20 working days, around half is wasted until the data is there, the first loops have been completed and obvious problems are visible. This is followed by a phase in which you prepare the data so that models can be trained sensibly. The actual model training is often surprisingly short, especially with modern AutoML approaches and proven standard methods.
The mistake many teams make is that they invest too late in what will later hit them hardest, namely data quality and process stability. Or they get lost in details too early before it is clear whether the use case is fundamentally viable.
Practical rule: Build your PoC plan so that you have a rough baseline and an initial model result by the first week at the latest. At the same time, you start communication in order to provide missing data or process information early. Because it almost always takes longer for the customer to deliver data than expected.
Domain knowledge as a success factor: Without the context of its creation, data becomes a trap
Domain knowledge is not a “nice to have” but rather a protection against false conclusions. Especially in industrial and sensor setups, a data science team without context can achieve good numbers in the short term, but the system later fails because conditions change or measured values are temporarily incorrect.
It is crucial that you understand how data is created:
- Which sensors measure what, in what condition, with what maintenance?
- Which process steps influence the values?
- What changes have been made in systems or collection intervals?
- Which anomalies are “real” and which are measurement artifacts?
If this knowledge only belongs to a few people, you have to actively involve them. Otherwise you are optimizing for a data image that is only randomly related to reality.
Practical rule: Secure binding, short time slots with domain experts for the PoC, for example 15 minutes twice a week. Not at some point, but from week 1. A PoC without available contacts wastes time in the wrong places.
Stumbling blocks in AI PoC: Silent errors will hit you later if you ignore them today
An underestimated risk are “silent errors”, i.e. errors that are not technically noticeable but devalue the results from a technical point of view. Examples:
- Data is syntactically correct but semantically incorrect (signs, units, limits).
- Data is available in the PoC, but not in later use.
- Collection intervals or systems change without the project team knowing.
- Preprocessing steps propagate small errors throughout the pipeline.
The solution is not “more care”, but rather simple validations that raise the alarm early: uniqueness of keys, value ranges, plausibility rules, checks when new data deliveries are received. This takes little time and saves massive debugging effort that would otherwise run backwards through the pipeline.
Practical rule: Set up light data validation checks in the PoC. You don't have to achieve a perfect production standard, but you do need guardrails so that you don't waste days troubleshooting because a trivial detail goes wrong somewhere.
Termination criteria for the AI-PoC: How to make a clean decision without drama
A PoC cancellation only seems like a failure if you define it as “we didn’t build a model”. If you define it as “we have created clarity”, it becomes a control instrument.
Use this decision matrix:
Stop (abort) if:
- The first model after a few iterations is far from the target and no realistic levers are visible.
- Critical data cannot be obtained within the planned period.
- The use case cannot be structurally evaluated (missing criteria, lack of experts, no priority).
Pivot (recrop) if:
- The use case makes sense in principle, but data frequency, data collection or process steps need to be adjusted.
- Success criteria should be adjusted because the original expectations do not match reality.
- A sub-problem is clearly solvable and quickly delivers business value.
Go (continue) if:
- You have a baseline and the model is already measurably better than the current state.
- Stakeholders can translate results into business impact.
- Domain experts are available and the project is given the necessary priority.
What you should do differently from tomorrow (specifically)
- Define a baseline in 60 minutes, even if it is rough: duration, costs, error rate, manual steps, dependencies.
- Force a first number early: A first model or evaluation score in week 1 to give you direction.
- Write success criteria on a page: Goal, Measurement, Scope, Non-Scope, Decision Logic.
- Plan domain expertise as a mandatory appointment: short, regular slots instead of sporadic queries.
- Incorporate minimal data validation so that silent errors don't eat up your iteration time.
- After a maximum of three iterations, make an interim decision: Go, Pivot or Stop, and justify it with the baseline and distance to the goal.
Why this matters: Once you have an initial number, you can translate it into a business case. Then “AI would be cool” turns into a concrete conversation about benefits, risks and next steps.
