Skip to content
For fTECHNOLOGY & DESIGN
For f

Generative AI / PoC / Commissioning guide

Five decision criteria before taking a generative AI PoC into production

Published by:For f Inc.

Five decision criteria before taking a generative AI PoC into production

Assess answer quality, business value, information handling, cost and speed, and operations before deciding whether to take an AI PoC into production.

Before moving from a generative AI PoC to production development, assess answer quality, operational impact, information controls, cost and latency, and ongoing operations under consistent usage conditions.A successful demo response and a system people can rely on in daily work require separate evaluation.

For companies considering generative AI for internal support or document search, this article presents For f's pre-engagement considerations. These five areas are a discussion framework, not an official certification or universal pass mark.

A PoC should inform the next decision

A PoC tests feasibility and impact within a limited scope. For document AI, record not only whether questions are answered but which sources are used, which questions cannot be answered and how much human review is needed. This helps define production scope.

Google Cloud's evaluation documentation describes preparing task-relevant datasets, selecting metrics and inspecting individual responses and aggregate results. The useful principle is evaluation against your own workflows as well as public benchmarks.Source: Google Cloud Gen AI evaluation service overview

Five criteria before moving to production development

A starting point for criteria agreed by the client and development team
Evaluation areaWhat to verify in the PoCDeliverables to retain
1. Answer qualityAre answers grounded in evidence? Are unanswerable questions handled appropriately?Questions, expected and actual answers, and reasons for each judgment
2. Operational impactDoes the workflow improve when review and correction are included?Current and AI-assisted procedures with comparison records
3. Information controlsAre user access scopes and confidential information requirements respected?Data flows, permission-based tests and unresolved issues
4. Cost and latencyAre cost and waiting time acceptable at expected usage levels?Measurement conditions, usage, cost breakdown and response times
5. OperationsCan quality be checked again after documents or models change?Update procedures, evaluation datasets and human fallback procedures

1. Answer quality: inspect errors as well as accuracy

Include common questions, questions with no answer in the documents, conflicts between versions and questions missing necessary context. Agree who evaluates them and what constitutes a correct answer.

NIST's Generative AI Profile identifies the risk of confidently generated false content. It recommends avoiding broad performance claims from narrow, unsystematic evaluations and checking the sources of generated content.Source: NIST AI 600-1, section 2.2, MS-2.5-001 and MS-2.5-003

Useful example categories include ready to use, usable after human correction, should abstain and escalate, and incorrect. Adapt these to the workflow to make improvement discussions more concrete.

2. Operational impact: include human review time

A fast AI answer may not reduce the overall workload if checking and rewriting take time. Compare input, review, correction and follow-up in the current and AI-assisted workflows.

For support inquiries, record drafting time, source verification, departmental consultation and final revisions. Agree what can be measured and how, rather than promising a reduction percentage first.

3. Information controls: test the same question with different permissions

For internal documents, inventory who may access each source. Test the same questions as different user roles, checking quotations and links as well as answers. Also consider input storage, log access and retention periods.

Separate answer correctness from whether the user is authorized to see it. Before testing real confidential material, agree sharing scope and handling conditions internally and with the implementation partner.

4. Cost and latency: align usage and measurement conditions

Separate model charges, search, storage, application hosting and operational work. Compare response times with recorded conditions such as short and long questions and low or concurrent usage.

Do not treat PoC measurements as production costs. State how many people use the system, how often and with what data lengths. Verify current official prices and usage conditions when estimating. This article provides no unverified cost, duration or performance figures.

5. Operations: can you evaluate again after changes?

Retain test questions, expected answers, source versions, models and settings. Aim to rerun the same questions after changes. Assign an owner for reported errors and define how to stop AI processing and switch to people.

Google Cloud describes generative AI evaluation as a continuous activity across development stages. This supports repeating evaluation after changes, rather than assessing only once at the end of a PoC.Source: Google Cloud Blog, How good is your AI? Gen AI evaluation at every stage, explained (June 13, 2025)

Classify the result as proceed, investigate further or stop

For f proposes the following discussion categories. Treating serious unresolved issues separately, rather than averaging everything into one score, helps define the engagement scope.

  • Proceed to production development:The agreed usage scope meets the criteria, and remaining issues and production responsibilities can be explained.
  • Run a narrower additional evaluation:Potential benefits exist, but specific questions about documents, permissions or answer quality remain to be tested.
  • Defer adoption or consider another approach:Business benefits are unverified, or there is no clear path to resolving critical issues.

Completing a PoC does not require proceeding to development. Understanding why not to proceed is also a valid outcome.

Useful preparation before a consultation or estimate

  • Target workflows and current difficulties
  • Actual questions or inputs and expected outputs
  • Source documents and user access scopes
  • Current procedures and any known volumes or processing times
  • Desired timing, budget approach and the decision this trial should inform

Start from your current situation even if documentation is incomplete. Clarifying scope, evaluation, deliverables, exclusions and progression conditions helps separate PoC and production estimates.

Frequently asked questions

Is high accuracy enough to proceed to production?

No. Evaluate error scenarios and consequences, information controls, review workload, cost and operations against actual usage conditions.

What percentage should be the pass mark?

There is no universal percentage. Required criteria depend on the workflow, consequences of errors and human review scope. Separate ordinary answer quality from serious failures such as unauthorized disclosure.

Can we consult you before requirements are defined?

Yes. For f can help scope the work from your current processes and challenges.AI discovery and adoptioncan be your starting point, orAI proof of conceptcan define the evaluation scope, depending on your stage.

AI PoC / NEXT STEP

Let's start with what you need to validate.

For f supports AI discovery, workflow analysis, PoCs, production development and data platforms. We clarify the workflow and decision the trial should support, then propose scope, deliverables and an approach.

For development informed by evaluation results, seeProduction AI systems.

References and verification date

Information checked on September 7, 2026. The public references below inform this article; procurement considerations are For f's explanations. The thumbnail is an AI-generated concept image, not actual evaluation results or a system interface.

  1. Google Cloud:Gen AI evaluation service overview— Evaluation datasets, metrics and results.
  2. NIST AI 600-1:Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile— Generative AI misinformation risks and evaluation considerations. July 2024.
  3. Google Cloud Blog:How good is your AI? Gen AI evaluation at every stage, explained— Evaluation across development stages. June 13, 2025.
Back to insights

LET’S CREATE WHAT’S NEXT

From insight to implementation.

Talk to us about operational challenges related to this article and practical implementation.

Discuss your project