arxivq-bio.QMcs.AIcs.LG2026-07-09
TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, et al.
The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-…