跳到主要导航 跳到搜索 跳到主要内容

Bridging imaging and genomics: Domain knowledge guided spatial transcriptomics analysis

  • Wei Zhang
  • , Xinci Liu
  • , Tong Chen
  • , Wenxin Xu
  • , Collin Sakal
  • , Ximing Nie
  • , Long Wang
  • , Xinyue Li*
  • *此作品的通讯作者
  • City University of Hong Kong
  • Capital Medical University
  • Tsinghua University
  • University of Science and Technology Beijing

科研成果: 期刊稿件文章同行评审

摘要

Spatial Transcriptomics (ST) provides spatially resolved gene expression distributions mapped onto high-resolution Whole Slide Images (WSIs), revealing the association between cellular morphology and gene expression profiles. However, the high costs and equipment constraints associated with ST data collection have led to a scarcity of ST datasets. Moreover, existing ST datasets often exhibit sparse gene expression distributions, which limit the accuracy and generalizability of gene expression prediction models derived from WSIs. To address these challenges, we propose DomainST (Domain knowledge-guided Spatial Transcriptomics analysis), a novel framework that leverages domain knowledge through Large Language Models (LLMs) to extract effective gene representations and utilizes foundation models to obtain robust image features for enhanced spatial gene expression prediction. Specifically, we utilize public gene reference databases to retrieve comprehensive gene summaries and employ LLMs to refine gene descriptions and generate informative gene embeddings. Concurrently, we apply medical visual-language foundation models to distill robust image representations at multiple scales, capturing the spatial context of WSIs. We further design a multimodal mixture of experts fusion module to effectively integrate multimodal data, leveraging complementary information across modalities. Extensive experiments conducted on three public ST datasets indicate that our method consistently outperforms state-of-the-art (SOTA) methods, with increases ranging from 6.7 % to 13.7 % in PCC@50 across all datasets compared to the SOTA, demonstrating the effectiveness of combining foundation models and LLM-derived domain knowledge for gene expression prediction. Our code and gene features are available at https://github.com/coffeeNtv/DomainST.

源语言英语
文章编号103746
期刊Information Fusion
127
DOI
出版状态已出版 - 3月 2026

指纹

探究 'Bridging imaging and genomics: Domain knowledge guided spatial transcriptomics analysis' 的科研主题。它们共同构成独一无二的指纹。

引用此