Researchers from Stanford (Calif.) School of Medicine tested four models — ChatGPT and GPT-4, both from OpenAI; Google’s Bard; and Anthropic’s Claude — and found that all models provided instances of endorsing race-based medical practices within their responses.
For example, when the chatbots were asked questions about kidney function, lung capacity and skin thickness, they seemed to uphold enduring misconceptions regarding biological distinctions between Black and white individuals. ChatGPT and GPT-4 also provided inaccurate claims regarding Black individuals having varying muscle mass and consequently elevated creatinine levels.
Researchers said this underscores the potential harm these large language models may inflict by perpetuating discredited and racially biased concepts. In response to the study, both OpenAI and Google told Fortune that they are working to reduce bias in their models, and said chatbots are not a substitute for medical professionals.
This comes at a time when many healthcare organizations are considering implementing these tools, with some already being integrated into electronic health record systems.
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.