Guidance for patient use of AI is in the works
Lask week, Microsoft and Amazon made waves by announcing healthcare artificial intelligence agents, joining Anthropic and OpenAI in an increasingly crowded market. At the moment, though, utilization is largely limited. As of mid-March, Anthropic’s Claude requires a subscription, Amazon’s tool is limited to Prime members, and Copilot Health and ChatGPT Health are waitlist-only.
In the meantime, patients will continue to turn to general-purpose large language models (LLMs) for healthcare help – imperfect as it is. Such activity “represents a major gap in health policy,” a multidisciplinary and multinational research team recently wrote in Nature, adding that “users are increasingly adopting these tools for health purposes despite this lack of oversight.”
In the absence of regulation and governance, the authors are developing The Health Chatbot Users’ Guide so “patients and the public use these tools safely.” Co-author Andrea Downing of the patient advocacy group The Light Collective described the guide in a LinkedIn post as an effort “to create practical, evidence-based guidance for how the public can use AI for health” and “actively challenge old power dynamics in digital health.”
Lots of demand, little guidance
Technology companies with forthcoming healthcare-specific AI tools have indicated there’s clear demand. Microsoft’s announcement noted its general-purpose Copilot products respond to 50 million consumer health questions per day. Meanwhile, OpenAI said 230 million global users ask health and wellness questions on ChatGPT every week.
But the learning curve for LLMs is steep – from entering prompts to interpreting outputs – and AI model efficacy is still being tested. Two recent studies show why there’s reason for caution.
- LLMs may be good at identifying health conditions and providing next-best action recommendations, but humans struggle to make sense of what AI chatbots say. Part of it is user error, as adding new information partway through a conversation alters the output. Part of it is human nature, such as forgetting or discounting bits of information.
- LLMs have a startling tendency to accept false health claims, especially if they appear in what sounds like a real clinical note and are written with a hint of authenticity. The real concern is when neither the AI model nor human reviewers spot a myth and it spreads from a note to a discharge summary to a care plan.
Healthcare-specific models represent a step forward. They’ve been fine-tuned to answer medical questions, and they keep healthcare data in a walled garden. Still, they show signs of immaturity: A recent study indicted ChatGPT Heath can handle “textbook emergencies” but struggled with nuanced cases, often suggesting a wait-and-see approach when human clinicians would recommend an urgent care or even ED visit.
A tool to reduce LLM risk
That’s where The Health Chatbot Users’ Guide enters the picture. The group behind the guide – a collaboration among advocacy groups, clinicians, technologists, and ethicists, according to Downing – aims to agree upon the best and worst use cases for AI in healthcare. From there, the goal is to survey experts, approve the guide’s final advice, and share it in accessible format later this year.
Though the groups’ work is still in the discovery phase, a few “red lines” have been identified: Entering personally identifiable information, using chatbots for medical emergencies, and calculating medication doses. The best ways to use AI tools? Asking for plain-language recommendations, refining questions for upcoming appointments, and checking facts.
Brian Eastwood is a Boston-based writer with more than 10 years of experience covering healthcare IT and healthcare delivery. He also writes about enterprise IT, consumer technology, and corporate leadership.