For an April 19 story, The Washington Post analyzed Google’s C4 data set, which is used for large language models including Google’s T5 and Facebook’s LLaMA (ChatGPT owner OpenAI does not disclose where its data comes from). The Post ranked 10 million websites based on the number of “tokens,” or words or phrases, in the data set.
Here are its top five sources for science and health information:
1. journals.plos.org
2. frontiersin.org
3. link.springer.com
4. ncbi.nlm.nih.gov
5. nature.com
The large language models also ingest content from Becker’s, including Becker’s Hospital Review (ranked No. 3,824), Becker’s ASC (No. 18,003), Becker’s Spine (No. 23,586) and Becker’s Dental (No. 946,137).
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.