Explore our Topics:

What happens when AI diagnostic tools start hearing and seeing patients?

Google is integrating video analysis into its AI medical system to add audio and visual cues to its diagnostic processes.
By admin
Aug 17, 2026, 10:57 AM

Human communication is deeply nuanced, with a constant stream of subtle changes in tone, shifts in posture, facial expressions, and other non-verbal signals supplementing the semantic meaning of the words we say out loud.

The ability to identify and interpret these signals is a key part of the diagnostic process. And it’s something that’s been largely missing from the AI clinical decision support environment up until now.

Large language models that rely on text-based input have achieved remarkable degrees of accuracy, setting the stage for AI-assisted diagnostics to eventually enter clinical practice.

These tools can analyze huge volumes of medical records, academic literature, and best practices to come up with suggestions – as long as all of that data has already been converted into words.

But what happens when an AI tool gets eyes and ears so it can actually listen to a cough or see a patient hunched over in pain without relying on another entity to interpret that information for it?

That removes a critically important layer of “telephone” between the AI engine and the actual patient, which could enhance the way AI diagnostic tools operate as clinical assistants.

Researchers at Google are taking on this challenge by adding audio-visual capabilities to AMIE, its research medical AI system. AMIE (which stands for Articulate Medical Intelligence Explorer) has already produced notable successes as a text-based diagnostic tool, demonstrating expert-level performance across multiple clinical dimensions.

Now, in a research paper published on arXiv, the team behind the tool explains how synchronous clinical video consultations can help the model perceive non-verbal clinical cues and further enhance its capabilities, potentially opening up a very different role for AI in the diagnostic process.

A multi-agent system for asynchronously analyzing input

AMIE Video is actually three AI agents working in concert to process different inputs, the team explained. A “talker” agent handles live conversation and prioritizes speed of response, while the “planner” takes charge of understanding the evolving clinical information, developing differential diagnoses, and managing the goals of the encounter. Meanwhile, the “perception” agent continuously analyzes the audiovisual input from the video stream for clinically relevant observations, and keeps a running memory of the results.

The multi-agent architecture allows the different components of the system to operate asynchronously, so that the “talker” can be providing immediate responses while the other components do more computationally intensive work in the background. The strategy helps maintain a natural flow of conversation with the user instead of forcing the system to pause or delay while processing every new piece of information it gathers.

Testing the system in clinical simulations

To test the system, the researchers asked patient actors to interact with one of three entities: the text-only version of AMIE, AMIE Video, and real primary care physicians conducing a telehealth visit with the patient.

A separate board of primary care physicians evaluated the encounters on a 20-point rubric, covering history taking, clinical reasoning, treatment planning and communication, as well as the accuracy of differential diagnoses, the reasoning behind the diagnoses, and the system’s ability to guide a physical exam.

The results indicated that the addition of audiovisual input could add accuracy and value to the AI diagnostic process, although it’s important to emphasize that with staged encounters by paid actors, the results might not translate quite as cleanly into real-world practice.

AMIE Video achieved an overall case-specific score of 83%, compared to 68% for the human PCPs. Its first-choice diagnosis matched the reference diagnosis in 91% of cases, versus just 77% for PCPs. And AMIE Video topped human clinicians in clinical reasoning, treatment planning, and history taking.

The biggest difference was in the ability to use the audiovisual information to collect relevant information and guide next steps in the exam. AMIE Video scored 77% (versus 51% for physicians) in using live video for information gathering and 72% versus 39% in guiding patients through physical examinations.

In one case, for example, AMIE Video was able to ask a patient with neck and shoulder pain to rotate his head. The “perception” agent registered the patient’s restricted movement along with a facial expression indicating discomfort. The tool then asked the patient to raise his arms, but observed that the patient didn’t perform the movement as instructed due to restricted range of motion. The observations were fed back into the clinical reasoning model and supported a diagnosis of musculoskeletal strain.

Moving upstream to collect information autonomously, not just interpret it

The methodology and performance of this particular model is intriguing – and even though it has its limitations, it points to a potential shift in the way clinical AI tools may eventually operate in a real-world setting.

Instead of primarily interpreting information that has already been gathered by other entities, AI tools that include audiovisual capabilities may take on a larger role in collecting primary data directly.

And by adding capabilities that allow AI tools to actually guide the clinical encounter to collect information it deems important to the diagnostic process, AI might be able to actively tell a physician in real time not only that it doesn’t have enough data yet to make an informed decision, but what specifically to do in order to gather the information that’s missing.

Enabling AI to participate in the iterative process of generating information would be a major shift in the (still very limited) role it currently occupies in clinical decision making, and could be a bridge to the type of real-time, comprehensive, multimodal data analysis that many AI enthusiasts envision.

AMIE Video provides one clue about how such a system could be designed. Future models could leverage additional agents to more directly collect and incorporate more diverse data streams, including waveform data, lab data, medical images, respiratory data, and outputs from wearables and other devices.

By adding these data sources to language-based reasoning models that are already maturing quickly, AI developers could begin to create high-value tools for augmenting clinical decision-making in real time.


Jennifer Bresnick is a journalist and freelance content creator with a decade of experience in the health IT industry.  Her work has focused on leveraging innovative technology tools to create value, improve health equity, and achieve the promises of the learning health system.  She can be reached at [email protected].


Show Your Support

Subscribe

Newsletter Logo

Subscribe to our topic-centric newsletters to get the latest insights delivered to your inbox weekly.

Enter your information below

By submitting this form, you are agreeing to DHI’s Privacy Policy and Terms of Use.