Always Learning: Why Healthcare AI Models Need Human Expertise

Always Learning: Why Healthcare AI Models Need Human Expertise
Always Learning: Why Healthcare AI Models Need Human Expertise
Nathan Abercrombie
LinkedIn Logo
Sr. Machine Learning Engineer
September 3, 2026

There’s a common misconception about AI: once a model is trained and deployed, the hard work is done. In reality, deployment is when some of the most important learning begins. That’s especially true in procedure rooms, like the operating room. ORs, cath labs, and interventional radiology are all dynamic environments where workflows change, rooms are rearranged, equipment differs, and, occasionally, something happens that a model hasn’t seen before.

For AI to produce data that every stakeholder and team across the hospital can trust, there needs to be a way to identify when the real world challenges the model, understand what actually happened, and use those examples to improve performance over time. 

At Apella, human judgment plays a targeted role in this process, helping our models learn from the moments where additional context matters most. In my last post, I dove into why and how we built our computer vision models for hospitals' most dynamic environments, eschewing the real limitations created by relying on synthetic data. Today, I share the essential role of data quality teams in making those models continuously improve under the most complex conditions.

Learn how Houston Methodist built team trust and cut costs with data accuracy.
Read the case study

Why real-world environments create edge cases

The more an AI model, whether computer vision, machine learning, or natural language processing, operates in real-world environments, the more variation it encounters. Workflows don’t always unfold the same way, and the environment itself can change.

A piece of equipment might temporarily block a camera’s view. A team might complete familiar steps in an unusual sequence. A hospital might introduce new equipment or change a procedure room's configuration.

These situations are inevitable in real healthcare environments. Taking a closer look at them helps establish what actually happened and, when needed, strengthen the model over time.

How Apella identifies what needs a closer look

At Apella, signals from our models and systems help identify healthcare data that may warrant a closer look. Lower-confidence detections, unusual patterns, or other signals can indicate that the model has encountered something outside the norm.

That doesn’t mean someone is sitting behind a screen manually reviewing every second of every procedure in every anesthesia room in every hospital — that would require a team of thousands of humans.

AI can process information at a scale that would be impossible to review manually. Human expertise can then focus on the moments where additional judgment matters most.

Where Apella’s data quality team adds value

Human review is a critical part of maintaining high-quality data, particularly when a case warrants additional scrutiny. Some AI platforms rely exclusively on their systems, but at Apella, we believe human expertise remains essential.

It's why we have a dedicated in-house data quality team that reviews cases that warrant additional review. In these instances, the team can establish the ground truth: what actually happened in the OR or other procedure room.

Finding an error solves an immediate problem, but understanding that error can solve a much bigger one. If the same type of error appears again, there may be a pattern worth investigating and a need to improve the model’s precision.

  • Why did the model struggle?
  • Is it encountering a new scenario it hasn’t seen enough?
  • Is an existing feature missing something important?
  • Is the issue unrelated to the model altogether?

Looking more closely at those questions turns a single correction into a learning opportunity.

The examples that challenge an AI model can be particularly valuable because they reveal where there is still room to improve. Those difficult cases can become valuable examples that inform future training and AI model development.

This creates an important feedback loop. As a computer vision or machine learning model operates in real hospital environments, signals can surface moments that warrant a closer look. The data quality team establishes what actually happened, and those difficult cases can inform future training and model development.

Over time, the real-world environment itself becomes an important source of learning. ORs and other procedure rooms are constantly evolving as hospitals rearrange rooms, introduce new equipment, and adjust workflows. Even relatively small changes can affect what a model encounters and introduce scenarios that weren’t well represented in its original training data.

Ongoing review helps teams recognize when those changes are affecting performance, understand what’s happening, and incorporate newer examples into future training. As the hospital evolves, the dataset can evolve with it.

Why this matters for hospitals

Progress in AI is often associated with reducing the need for human involvement. In healthcare, though, human expertise still plays an important role, particularly in moments when context, judgment, and a deeper understanding of the environment matter most.

At Apella, the goal has always been to improve human actionability, prioritization, and effective decision-making through AI-driven data collection, analytics, and forecasting, while using automation to address the time burdens typically placed on overworked clinicians and hospital staff.

AI can process far more information than a person could reasonably review manually. Apella’s data quality team can then focus on difficult cases, helping establish ground truth and identify opportunities to improve future models so that healthcare workers can continue to execute that data with confidence and efficiency.

Healthcare AI doesn’t operate in a controlled environment. ORs, cath labs, and IRs change, workflows evolve, and new edge cases emerge. Human judgment helps us recognize those changes, understand what the model is encountering, and use those insights to keep improving real-world performance and the reliability of the data hospitals use to make operational decisions.

Learn how Houston Methodist built team trust and cut costs with data accuracy.
Read the case study

Always Learning: Why Healthcare AI Models Need Human Expertise

As a senior machine learning engineer at Apella, Nathan builds the computer vision systems that ensure the quality, reliability, and scalability of the company’s AI models in real-world clinical environments. His work spans computer vision, event detection, and predictive modeling, with a strong emphasis on measurement, evaluation, and continuous improvement. He develops evaluation pipelines and realtime tooling that enable high-quality training data at scale, helping ensure Apella’s AI systems deliver transparent, trustworthy performance in the operating room.