Is Your RFP Asking the Right Questions About AI?

Hospitals evaluating ambient AI know to ask about security, EHR integration, data residency, uptime, and support. Those questions are standard in most health system requests for proposal. But one area often receives far less scrutiny: Does the AI actually deliver outcomes, and what evidence proves it?
With AI, evaluating the platform is only part of the equation. Hospitals also need to evaluate the models behind it:
- What were they built to do?
- How well do they perform against objectives?
- How were that performance and impact validated?
- Do the claims hold up in real clinical environments?
Here’s a framework for making sure your RFP covers all the right bases. Following this guide will set your organization up to evaluate vendor proposals according to the likelihood of achieving your goals with a new hospital operations AI project.
1. Define exactly what each AI capability is expected to do
Broad questions about whether a platform “uses AI” or what AI capabilities it offers leave too much room for interpretation. An RFP should specify what each capability does, what specific technology underlies the capability, and what workflow or decision it is intended to support.
Does the platform detect specific events as they happen? Predict when a procedure will end? Identify when a turnover is running late? Support scheduling or capacity decisions? Is the platform built on its own, dedicated computer vision models and a robust, performant infrastructure?
For each capability, ask vendors to define:
- What the model is designed to detect, predict, or recommend
- The clinical environments and workflows it was designed for
- How frequently the model produces an output
- How that output is used by clinical or operational teams
This creates a clear foundation for evaluating the evidence that follows. You can only meaningfully assess performance when you clearly define the capability and the technology.
Check out our complete AI implementation checklist for ORs.
Download now
2. Require performance metrics for each capability
Avoid broad RFP questions such as “Describe the accuracy of your AI.” Different AI capabilities require different performance measures, and a single accuracy number can obscure important differences between models.
For event-detection models, one useful measure is an F1 score, which accounts for both precision and recall. Predictive models require different measures, such as how often a forecast falls within a meaningful window or how predicted durations compare with actual case durations.
Depending on the capability, an RFP might require:
- Precision, recall, and F1 score for event detection
- Latency between an event occurring and the system detecting it
- Prediction accuracy within windows that are meaningful to operations
- Performance across different events, specialties, or clinical environments
Instead of asking:
“How accurate are your AI models?”
Consider asking:
“Provide the performance metrics used to evaluate each AI capability, including how each metric is calculated and the results achieved.”

3. Require vendors to show how performance was established
A performance metric needs context. Your RFP should require vendors to explain how the technology was evaluated and what evidence supports the results.
For each reported metric, ask vendors to disclose:
- Against what the model was measured and how ground truth was established
- Whether it was evaluated using real-world clinical data, synthetic data, or both
- The number of cases included in the evaluation
- The number of hospitals, procedure rooms, and specialties represented
- Whether the results were internally evaluated, independently validated, or peer-reviewed
Ground truth is important for ambient AI. If a model is evaluated against manually documented EHR timestamps, for example, the reference point itself may be delayed or inconsistent.
The scale of the evidence matters too. Performance demonstrated across many real-world cases, hospitals, and clinical environments provides a stronger indication of how a model performs across the variation it will encounter in practice. Look for evidence that clearly describes the size and scope of the evaluation, along with how performance was measured and validated.
Instead of asking:
“Has your AI been clinically validated?”
Consider asking:
“Provide evidence supporting the performance of each AI capability, including how it was evaluated, the size of the evaluation, how ground truth was established, and any independent or peer-reviewed validation.”
4. Build post-deployment performance into the RFP
Performance at the time of evaluation is only a part of the picture. Clinical environments change after go-live. Equipment moves, room configurations change, workflows evolve, and models encounter scenarios they haven't seen before.
The RFP should establish how performance will be measured and maintained once the technology is deployed. Ask vendors to describe:
- How model performance is monitored after go-live
- How uncertain or unusual cases are identified
- How ground truth is established when a case requires additional review
- Who is responsible for reviewing data quality and model performance
- How model updates and changes are managed
- What performance reporting is available to the hospital
At Apella, for example, a dedicated data quality team reviews model outputs against verified ground truth, identifies cases that warrant additional review, and uses them to maintain and improve model performance over time.
Hospitals should evaluate that process before implementation, rather than discovering how model quality is managed after go-live.
5. Ask for evidence that the technology is used and creates measurable impact
Strong model performance should translate into technology that teams use and measurable improvements in hospital operations. An RFP should ask vendors for evidence of both.
Ask who is actually using the technology today. Are charge nurses, surgeons, anesthesiologists, and perioperative leaders incorporating it into their workflows? What does user feedback show? How consistently is the technology being used?
Then require evidence of measurable operational impact. Depending on the use case, that might include:
- Turnover time
- Case volume
- Schedule accuracy
- Utilization
- On-time starts
- Late-running rooms
- Available capacity
The RFP should also ask vendors to explain how they measured those outcomes, over what period, and across how many sites. This context matters because results demonstrated at a single site may not generalize to other hospital environments.
Make the evidence part of the requirement
Security, integration, infrastructure, and support will always be critical parts of healthcare technology evaluation, but AI adds another category of diligence to that process.
RFPs should require vendors to support their AI claims with clear evidence. That means showing how performance was measured and validated, how it is maintained over time, and what results have been achieved in real hospital environments.
A strong AI RFP gives hospitals the information they need to evaluate the technology with the same rigor they apply to security, integration, and infrastructure.
Check out our complete AI implementation checklist for ORs.
Download now

