Why biological data matters more in AI drug discovery
artificialintelligence-news.com Aug 3, 2026

Why biological data matters more in AI drug discovery

AI-summarised brief · reviewed before publication

GlaxoSmithKline has expanded its research collaboration with Relation Therapeutics, signing an agreement worth up to $110 million to advance AI-assisted drug discovery. Relation will generate large-scale datasets measuring human cellular responses to genetic changes and drug interventions. This biological data will train AI models, including those in Relation’s MORGAN platform, to identify potential drug targets. The partnership builds on previous work targeting fibrotic diseases and osteoarthritis, which utilized Relation’s Lab-in-the-Loop platform. This approach combines computational analysis with laboratory experiments, integrating human genetics, single-cell multi-omics, and functional assays. While public repositories offer vast single-cell data, technical challenges like noise and dataset overlap complicate model training. Recent studies indicate that simply increasing biological data volume does not guarantee better AI performance, suggesting that data quality and model capacity balance are critical for robust single-cell foundation models in pharmaceutical research.

💡 Why It Matters

  • · The deal underscores a strategic pivot from chasing data volume to prioritizing high-quality, experimentally validated biological inputs for AI training.
  • · It challenges the prevailing assumption that larger datasets automatically yield superior drug discovery models, emphasizing instead the necessity of rigorous data curation and balanced computational resources.