The RxRx1 dataset takes up approximately 300 gigabytes of space and comprises more than 125,000 images taken through a microscope. Each image shows one of four cell types — sourced from umbilical veins, retinas, liver cancer and bone cancer — exposed to a piece of RNA that had been modified to suppress one of 1,000 genes.
Recursion’s open-source data aims to help other biotech companies using AI to identify specific molecules that can then be targeted with new drugs, since programming those machine learning models requires massive amounts of data that can be difficult, costly and time-consuming to obtain.
“The best models we can train are still data-limited — they’re still hungry for more data,” Jason Yosinski, PhD, a machine learning adviser at Recursion, told STAT. “By training on more images, models will be able to learn more subtle features.”
More articles about AI:
How AI helped 2 health systems automate case reviews & reduce claim denials
Two studies use AI-generated images to stimulate specific brain neurons
Study: Deep learning can predict reactions to thrombolysis
At the Becker's 11th Annual IT + Revenue Cycle Conference: The Future of AI & Digital Health, taking place September 14–17 in Chicago, healthcare executives and digital leaders from across the country will come together to explore how AI, interoperability, cybersecurity, and revenue cycle innovation are transforming care delivery, strengthening financial performance, and driving the next era of digital health. Apply for complimentary registration now.