Curated summary
Advancing AMIE towards expert-level audio-visual clinical consultations
AMIE (Video) is Google’s real-time audiovisual medical AI system, designed to overcome the limitations of text-only clinical conversations. Built on Gemini and Project Astra, it observes visual and auditory cues, guides patients through virtual examinations, and performs diagnostic reasoning during live consultations. In a randomized study involving 300 simulated consultations, the system was evaluated against text-only AMIE and board-certified primary care physicians.
Why Audio-Visual Consultation Matters
- Traditional text-based systems lose important clinical information, including:
- Gait and visible physical symptoms
- Breathing patterns and signs of distress
- Vocal and auditory cues
- Patient responses during physical examination maneuvers
- Requiring patients to describe symptoms in writing can reduce diagnostic accuracy, particularly for people with limited digital or health literacy.
- Audiovisual interaction may also improve trust, communication, and access to medical expertise.
AMIE’s Broader Development
- Earlier versions of AMIE demonstrated expert-level performance in:
- Text-based diagnostic dialogue
- Differential diagnosis support
- Disease treatment and longitudinal management
- Specialist evaluations in oncology, cardiology, and ophthalmology
- Reasoning over medical images and clinical documents
- Google has also explored physician oversight and real-world clinical feasibility studies.
Asynchronous Multi-Agent Architecture
AMIE (Video) divides the consultation among three agents operating in parallel:
Talker agent
- Maintains natural, low-latency spoken conversation.
- Incorporates information and recommendations from the other agents.
Planner agent
- Performs deeper clinical reasoning in the background.
- Updates differential diagnoses and management plans.
- Identifies missing information and reprioritizes clinical objectives.
Perception agent
- Continuously analyzes audio and video.
- Detects non-verbal findings such as visible distress, physical signs, and auditory abnormalities.
- Interprets observations in the context of the conversation.
This separation allows AMIE to reason deeply without creating long conversational pauses. Automated tests indicated that the agents contributed to improvements in history-taking, clinical reasoning, treatment recommendations, communication quality, and response latency.
Automated Evaluation Framework
- Google created a taxonomy of audiovisual clinical competencies based on medical literature.
- The taxonomy covered:
- Non-verbal visual cues
- Auditory signals
- Physical examination maneuvers
- The evaluation suite included:
- Single-turn tests targeting specific perception and reasoning abilities
- Multi-turn simulated consultations assessing complete conversational performance
- Simulations injected visual findings as textual descriptions, such as a patient holding handwriting samples up to the camera.
- These tests helped identify capabilities and failure modes before human evaluation.
Randomized Video Study
- The study used a synchronous video consultation interface and an Objective Structured Clinical Examination format.
- It included:
- 100 clinical scenarios
- Five body systems: cardiopulmonary, abdominal, HEENT, neurological/psychiatric, and musculoskeletal
- 15 trained patient actors
- 300 standardized consultations
- Three study arms were compared:
- AMIE (Video): Real-time audiovisual consultations
- AMIE (Text): Text-only AMIE used to isolate the value of audiovisual capabilities
- PCP (Video): Board-certified primary care physicians using the same video interface
- An independent panel of 20 experienced primary care physicians assessed the consultations using established clinical rubrics.
AMIE (Video) represents a move from text-based medical dialogue toward interactive, multimodal consultations. Its multi-agent design and audiovisual perception are intended to preserve conversational responsiveness while supporting richer clinical reasoning, though the reported findings come from simulated consultations and require further validation in real-world clinical care.
Related reading
Continue with another curated summary.
SymptomAI: Towards a conversational AI agent for everyday symptom assessment
Read originalHow AI tools can redefine universal design to increase accessibility
Read originalSmall models, big results: Achieving superior intent extraction through decomposition
Read originalA New Era of Innovation: Google Research at I/O 2026
Read original