Abstract:
Objective: To investigate the effect of standardized descriptions of chest imaging terminologies on the diagnostic efficacy of general large language models (DeepSeek-R1). Methods: One hundred cases of multiple and varied pulmonary lesions from CT scans were examined (23 pneumonia, 22 tuberculosis, 20 fibrosis, 18 pulmonary edema, 7 allergic pneumonia, and 10 rare cases). The input text was classified into Condition A (imaging description only) and Condition B (imaging + medical history/laboratory examinations), with each group including both original and standardized descriptions. Based on clinical diagnosis as the benchmark, the accuracy rate for the top result as well as the Top-3 and Top-5 inclusion rates were calculated. The paired McNemar test was performed to compare differences between groups, while the Wilcoxon test was conducted to assess the performance improvement. Results: The standardized descriptions improved the overall Top-1 accuracy rate for Condition A from 32% to 70% (Δ38%, P < 0.001), while multimodal input (B) increased the Top-1 accuracy rate by 24% relative to the original descriptions. The Top-5 inclusion rate for all diseases in the standardized B group reached 100%, whereas the Top-1 accuracy rate was 95%. Conclusion: Standardized terminology significantly enhances the model’s ability to interpret imaging features. Multimodal input can compensate for the deficiencies of nonstandard descriptions, while the combined use of both can yield optimal diagnostic efficacy.