Use a simple evidence scale
A number without evidence is decoration. For every score, link the representative cases, observation, measurement, or stakeholder decision that supports it. Add a confidence note because five carefully selected cases and fifty representative cases do not carry the same weight.
| Score | Meaning | Required note |
|---|---|---|
| 0 | Not tested or contradicted | What evidence is missing or what failed |
| 1 | Weak or highly conditional | Where it worked and why confidence stays low |
| 2 | Promising with known gaps | What must change before broader use |
| 3 | Strong for the tested scope | The evidence set and boundary of the result |
Score these twelve dimensions
1. User and business value
Does the workflow solve a costly, frequent, risky, or strategically important problem for a clear user? Is the improvement meaningful enough to justify product and operating change?
2. Output quality
Across representative cases, is the result accurate, complete, relevant, and appropriately formatted? Separate harmless variation from failures that change decisions or actions.
3. Failure safety
Can the system abstain, request more context, cite evidence, escalate, or require approval when confidence is weak or impact is high?
4. Data fitness and rights
Are the necessary inputs available, current, representative, permitted, and maintainable? Can sensitive data be minimized and access controlled?
5. Workflow fit
Does the capability fit the actual sequence of work, including roles, handoffs, exceptions, review, and recovery, or does it create a new isolated task?
6. User trust and control
Can people understand what the system did, inspect sources or reasoning cues where useful, correct the result, and remain responsible for important outcomes?
7. Integration feasibility
Can the capability connect to the required systems, permissions, events, and data contracts without creating an unsafe or brittle operating model?
8. Latency
Is response time acceptable for the real workflow, including retrieval, tool calls, retries, and human review rather than only the fastest demonstration?
9. Unit economics
Are model, retrieval, infrastructure, review, and exception costs proportionate to the value of the completed task at realistic volume?
10. Observability
Can the team see inputs, outputs, model or prompt version, retrieval context, errors, latency, cost, and outcome signals well enough to diagnose drift?
11. Security and compliance path
Are identity, access, secrets, data retention, provider terms, audit needs, and regulatory constraints understood for the intended use?
12. Ownership and production path
Is there a credible team, architecture, evaluation process, support model, and rollout route after the proof, or would the POC become an orphaned demo?
Turn the score into a recommendation
Do not treat the total as an automatic answer. First review any stop condition: weak value, unacceptable harm, unavailable or unpermitted data, no accountable owner, or a unit cost that cannot fit the outcome.
- Proceed when the critical feasibility question is answered and remaining gaps have a credible production route.
- Adjust when value remains plausible but scope, workflow, data, evaluation, model approach, or human control needs a bounded change.
- Pause when a dependency, owner decision, data right, integration, or operating condition must change before more technical work creates useful evidence.
- Stop when the tested workflow does not create enough value, cannot be made acceptably safe, or requires disproportionate cost and complexity.
Run a review that preserves context
The scorecard becomes a bridge into production planning. It shows which areas are already supported, which remain hypotheses, and which gaps could invalidate the next investment.
- Freeze the tested scope and representative case set.
- Have product, domain, engineering, and operational owners score independently before discussion.
- Review disagreements as evidence gaps, not as votes to average.
- Record conditions attached to every proceed or adjust decision.
- Store the scorecard with the POC code, evaluation cases, and next-scope brief.
Sources and further reading
- NIST: AI Risk Management Framework Core : Deployment-like measurement, documented risk, safe failure, and evidence-led go or no-go decisions.
- OpenAI: Evaluation best practices : Task-specific tests, real-world distributions, human calibration, and explicit thresholds.
- NIST: AI RMF Manage Playbook : Deciding whether deployment should proceed and monitoring risk through the lifecycle.