Breast ultrasound diagnosis based on cross-modal generation and cross-attention(PDF)
《中国医学物理学杂志》[ISSN:1005-202X/CN:44-1351/R]
- Issue:
- 2026年第7期
- Page:
- 889-897
- Research Field:
- 医学影像物理
- Publishing date:
Info
- Title:
- Breast ultrasound diagnosis based on cross-modal generation and cross-attention
- Author(s):
- FU Lijia1; LIN Yanping1; LI Na2; JIA Chao3; 4; LI Gang3; 4; DU Lianfang3; 4; LI Fan4; 5
- 1. School of Mechanical Engineering, Shanghai Jiao Tong University, Shanghai 200240, China 2. School of Health Science and Engineering, University of Shanghai for Science and Technology, Shanghai 200093, China 3. Department of Ultrasound, Shanghai General Hospital, Shanghai 201620, China 4. School of Medicine, Shanghai Jiao Tong University, Shanghai 200025, China 5. Department of Ultrasound, Shanghai Chest Hospital, Shanghai 200030, China
- Keywords:
- Keywords: breast cancer breast ultrasound contrast-enhanced ultrasound auxiliary diagnosis generative adversarial network cross-modal generation cross-attention feature fusion
- PACS:
- R318;TP391.41
- DOI:
- DOI:10.3969/j.issn.1005-202X.2026.07.008
- Abstract:
- Abstract: Objective To address the limited accuracy of conventional B-mode ultrasound (US) in breast cancer screening, resulting from poor inter-reader agreement and atypical malignant signs of tiny lesions, and resolve the clinical dilemma where contrast-enhanced ultrasound (CEUS) is hard to popularize due to its high cost and operational complexity, this study proposes a three-stage cascaded diagnostic framework of "segmentation-generation-classification" to achieve high-precision diagnosis of benign and malignant breast tumors relying solely on conventional US images. Methods A CNN-Transformer framework-based segmentation network was first employed to extract tumor regions from US images. Subsequently, a cross-modal generation model based on a conditional generative adversarial network was constructed, utilizing a pre-trained ResNet-34 encoder to synthesize corresponding CEUS images from the segmented US images. Finally, a dual-branch cross-attention network with Inception-v3 as the backbone was designed to achieve deep interaction between the real US and the generated CEUS information through dynamic feature fusion, thereby accomplishing accurate benign and malignant tumor classification. Results Validated on a breast ultrasound dataset including 677 patients, the proposed framework achieved an area under the receiver operating characteristic curve (AUC) of 0.930, significantly outperforming several single-modal US baseline models which reached a maximum AUC of 0.868. Ablation experiments verified the essential contributions of lesion segmentation, cross-modal generation, and cross-attention feature fusion modules in enhancing diagnostic performance. Conclusion This study validates the feasibility and effectiveness of generating and utilizing multimodal ultrasound information with deep learning for high-precision diagnosis. The proposed framework provides multimodal diagnostic benefits in a cost-effective and highly efficient way, holding great promise to provide robust auxiliary decision support for clinical breast cancer screening and reduce the risks of missed diagnoses and misdiagnoses.
Last Update: 2026-07-22