Apella set out to build computer vision for ambient AI in the operating room, later expanding into procedure rooms like cath labs and interventional radiology. From the start, we faced a choice: adapt existing models or build specifically for the environments we needed to understand.
General-purpose models are powerful, and more off-the-shelf options are available today than ever before. But the problem we needed to solve wasn’t simply recognizing what appears in an image. We needed to understand what was happening inside any and every procedure room, accurately and consistently, across all the variation that comes with real anesthesia and surgical environments. That required a different approach.
The OR is not a generic environment
Operating rooms and other procedure rooms are complex, dynamic environments. Teams move in and out, room configurations vary wildly, and equipment is constantly repositioned. Physician workflows and perioperative processes differ across hospitals, specialties, and even individual rooms, and events don’t always happen in the sequence you would expect.
Consider something as simple as detecting when a patient enters the OR. Most of the time, the order of operations is straightforward. But what happens when a patient enters the room, is wheeled back out, and returns five minutes later? A machine learning model may recognize a patient or stretcher in the room without understanding what that unusual sequence means, operationally.
Those exceptions matter. If healthcare teams are going to use ambient AI data to understand utilization, identify delays, read forecasts, or make decisions about surgical capacity, the computer vision model needs to perform when workflows get messy, not just when everything goes according to plan.

Built for the environment where it operates
Rather than adapting a general-purpose computer vision model to the OR and other procedure rooms, Apella first built models specifically for perioperative workflows. The models are trained to recognize the operational events that matter, as well as their context, using data that reflects how each OR and procedure room actually functions, rather than an imaginary, over-idealized version.
Importantly, that training is grounded in real hospital environments rather than relying on synthetic representations of what a room or workflow should look like. Synthetic data can be useful, but real hospitals introduce a level of variation that is difficult to replicate, from different room configurations and equipment to unexpected sequences of events, and situations where the camera’s view is partially obstructed. Training across that real-world variation helps models learn not just what an OR or other procedure room looks like, but also how they actually operate.
Learn how Houston Methodist built team trust and cut costs with data accuracy.
Read the case study
Real-world data makes the models stronger
When Apella enters a new environment, we evaluate how our models perform across the specific workflows they encounter. If a model struggles with a particular scenario, that becomes valuable information. It signals what the model is missing and where the next version can improve. It also means that the deep learning capabilities built into Apella’s neural network are working as they were designed and able to rapidly adapt to newly relevant details.
Over time, that creates a growing body of knowledge and insight about how anesthesia rooms operate across different hospitals, specialties, room configurations, and workflows. And that knowledge influences how we build and hone our models as well as what to expect, in terms of additional context and training that needs to be gained, as Apella is adopted by new hospitals. The more real-world variation Apella’s computer vision models encounter, the better we can apply what we’ve learned to new hospitals and adapt when unfamiliar scenarios arise.
Owning the models means owning the ability to improve them
Building our own computer vision models also gives Apella control over what the models are designed to do and how they evolve.
We can measure performance against the events that matter to our customers, investigate where models are less confident or where unexpected sequences occur, and focus development on the scenarios that most affect the quality of the data our customers ultimately use. We can also evaluate that performance rigorously: in a peer-reviewed study of more than 100,000 surgical cases, Apella’s computer vision models achieved F1 scores above 0.98 across all reported perioperative events.
That matters because our goal isn’t computer vision for its own sake. It’s to accurately understand what’s happening in the OR or other procedure room, quickly enough to make that information useful and actionable.
Sometimes that means improving how a computer vision model handles an unusual workflow. Sometimes it means adapting to changes in the environment. And sometimes it means being selective about what we train our models to detect, focusing on the events that provide meaningful operational value.
That focus also makes models more efficient. General-purpose models use computing power to interpret a vast range of objects and scenarios, many of which aren’t relevant in a procedure room. Because Apella’s models are built specifically for this, they can concentrate computing resources on what actually matters. As a result, they can run continuously in real time at a small fraction of the cost of general-purpose models, while keeping latency low.
Owning the underlying models allows us to make those decisions based on what matters in the specific hospital context, rather than what a general-purpose system happens to be designed to detect.

Purpose-built for the OR and other procedure rooms
General-purpose computer vision models are powerful, but a model’s size or breadth doesn’t necessarily determine how well it will perform on a highly specialized, context-relevant problem.
In the OR and other procedure rooms like cath and IR, performance depends on understanding both the workflows that happen every day and the unusual moments that inevitably come up. That understanding provides the context needed to translate what’s happening in the OR or other procedure room into useful operational data.
That’s why Apella built its own computer vision models. A purpose-built approach gave us the depth of understanding needed to turn what actually happens in hospitals into accurate, trustworthy, timely data that is precise enough to be used for predictions and to be considered future-proofed against any amount of change that may happen, either as a result of the constant pace of change around health systems or, especially, process improvements that come from impactful ambient intelligence. And, with every new hospital, workflow, and real-world scenario Apella’s models encounter, that understanding just gets stronger and stronger.
Learn how Houston Methodist built team trust and cut costs with data accuracy.
Read the case study

