How can I tell whether an AI project has real value or is just hype?
Judge an AI proposal on five points: a frequent and costly problem, usable data, a meaningful advantage over rules or conventional software, failures that can be controlled, and a small pilot that can measure a business result. If a proposal cannot explain two or more of these, it is not ready for full implementation.
Model names, parameter counts, the number of agents, and a smooth demonstration are not business value. Replacing a reliable search, rule, or form workflow with a language model may add cost and uncertainty. AI is most useful for natural language, images, long documents, and complex long-tail judgments; deterministic calculation, authorization, and consequential actions should usually remain in conventional software.
When defining model, data, and production boundaries, also compare Does AI make sense for a small, non-technology company with limited data? and Will Wavesteam make a simple requirement expensive just to use AI or a new technology?; the linked guidance adds context that should be considered in the same decision.
Evidence of value versus signs of hype
| Test | Credible project | Typical hype | Evidence to request |
|---|---|---|---|
| Business problem | Identifies who spends how much time on which step and the errors it causes | “Make the whole business intelligent” | Current volume, labour, error, and loss baseline |
| Need for AI | Explains why rules, search, or an existing SaaS product is insufficient | “Large models are the future” | Cost and outcome of a non-AI alternative |
| Data | Lawful, accessible samples represent the real distribution | “We will collect data later” | Sample inventory, quality findings, and accountable owner |
| Failure boundary | Defines refusal, approval, rollback, and excluded actions | “The model will keep getting smarter” | Failure list, permission matrix, and fallback process |
| Validation | Gives the PoC and pilot pass gate and stop condition | Shows only a few good responses | Held-out test results and operating metrics |
Start with today's baseline
Record monthly task volume, human minutes per item, wait time, rework, severe errors, and direct cost. Then estimate the complete AI cost: model usage, vector or compute infrastructure, integration, human review, content maintenance, security, and operations. Token fees cannot be compared with salaries in isolation, and every saved minute is not automatically removable headcount. Benefits may instead be faster response, greater capacity, or more expert attention for difficult cases.
Compare against at least one non-AI baseline. Test rules or conventional machine learning for order classification, keyword search for knowledge lookup, and a deterministic process for approval. If AI merely modernizes the interface without improving task success or cost per successful task, it should not become the core system.
Where value is easier to prove
Good first candidates are frequent, contain unstructured input, already have an understood human method, produce checkable results, and can escalate safely. Examples include invoice or order extraction, long-document triage, internal knowledge retrieval, service classification, and response drafts. Automated payment, medical conclusions, high-impact risk decisions, or complete replacement of experts require much stronger evidence and accountability; they should not begin fully autonomous.
Use representative historical samples for a PoC and stratify common, long-tail, and high-risk cases. Measure human takeover, P95 latency, cost per successful task, and severe failures in addition to accuracy. A limited live pilot should show whether people continue using the result, handling time changes, mistakes remain discoverable, and performance survives source and model updates. A report with no failed examples usually reflects incomplete testing.
A practical investment gate
| Decision | Condition | Next step |
|---|---|---|
| Continue | Core metrics pass, severe risks are controlled, and projected benefit covers total cost | Run a limited pilot, then expand in stages |
| Narrow the scope | Some tasks work but long-tail cases or write actions remain risky | Keep the strong tasks and add human confirmation |
| Pause | Data is unavailable, a non-AI solution is better, value is unmeasurable, or failure has no safe fallback | Improve data or the underlying process first |
Wavesteam compares AI, rules, search, existing SaaS, and custom software on the same business outcome. We recommend buying a mature product when it already fits. Custom AI is justified when unstructured understanding, knowledge retrieval, or a genuinely different workflow produces measurable advantage. Our experience can shape the experiment, but no previous case study substitutes for evidence from the current client's samples.
References
- The NIST AI Risk Management Framework organizes governance, mapping, measurement, and management of AI risk; it does not prove commercial return.
- OpenAI's evaluation best practices supports task-specific, stratified, continuing evaluation without requiring a project to remain with one provider.
- The FinOps FOCUS specification helps normalize technology cost data for more complete cloud and AI cost analysis.
Any savings claim must state its sample, period, and calculation. Before a real pilot, a universal ROI promise is not credible.