google3 min read

Curated summary

Exploring the feasibility of conversational diagnostic AI in a real-world clinical study

Read original(opens in new tab)

The study evaluated Google’s conversational medical AI, AMIE, in a real-world primary care workflow rather than simulated cases. In a prospective, IRB-approved study at Beth Israel Deaconess Medical Center, AMIE conducted supervised pre-visit history-taking with 100 patients. Results suggested that supervised deployment was feasible and conversationally safe, while AMIE’s diagnostic and management-plan quality was broadly comparable to that of primary care physicians, with physicians performing better on practicality and cost effectiveness.

Study Design and Clinical Workflow

  • Patients with new, non-emergency, episodic complaints used AMIE through a secure web link before an in-person or telehealth appointment.
  • A physician supervised each AI-patient interaction through live video and screen-sharing.
  • AMIE produced a transcript and summary for the patient’s primary care physician.
  • Independent clinical evaluators assessed:
    • The quality of the AMIE conversation
    • AMIE’s differential diagnoses
    • AMIE’s management plans
    • Comparable outputs from physicians
  • The study was prospective, single-center, single-arm, pre-registered, and IRB approved.

Participants

  • 100 adults completed the AMIE interaction.
  • 98 attended their scheduled primary care appointments.
  • Participants represented varied ages, racial and ethnic groups, health literacy, technology literacy, and prior chatbot experience.
  • Compared with all 1,452 urgent care visits during the study period, participants tended to be younger, although the sample reflected the broader population’s female and white demographic skew.

Safety Oversight

  • Human supervisors could stop an interaction if they observed:
    • Immediate risk of harm to the patient or others
    • Significant emotional distress related to the AI interaction
    • Potential clinical harm
    • A patient’s explicit request to end the session
  • No safety stops were required across the study.
  • The authors interpret this as evidence that supervised AMIE interactions were conversationally safe in this setting.

Clinical Reasoning Performance

  • Three independent clinical evaluators reviewed each case using blinded, randomized assessments.
  • AMIE and physicians showed similar overall quality for:
    • Differential diagnoses
    • Management plans
    • Management-plan appropriateness and safety
  • Physicians performed better on the practicality and cost effectiveness of management plans.
  • AMIE’s differential-diagnosis accuracy was reported as high, including cases where the final diagnosis was confirmed through diagnostic testing.

Patient and Clinician Experience

  • The study measured trust, perceptions, and acceptance among both patients and clinicians.
  • Patient trust in AI increased after interacting with AMIE.
  • Overall findings indicated that the system was well received within the supervised pre-visit workflow.

The study supports cautious, supervised testing of conversational diagnostic AI in clinical environments. It does not establish that AMIE can independently replace clinicians; rather, it suggests that pre-visit information gathering may be a practical early use case, provided rigorous oversight, safety protocols, and further evaluation in larger and more diverse settings.

Continue with another curated summary.