Extraction of clinical data on major pulmonary diseases from unstructured radiologic reports using a large language model

Citations

WEB OF SCIENCE

0

초록

Despite advances in big data technology, extracting information from unstructured clinical data remains challenging. This study aims to investigate the usefulness of large language models (LLMs) in extracting clinical data from unstructured radiologic reports.In this retrospective study at three university hospitals, a total of 1800 radiologic reports were analyzed. Pulmonologists evaluated the reports for major pulmonary diseases, and their findings were compared with assessments by Gemini, GPT-3.5, and GPT-4.0 to evaluate sensitivity, specificity, and accuracy.The presence of seven pulmonary diseases, including active tuberculosis, emphysema, interstitial lung disease, lung cancer, pleural effusion, pneumonia, and pulmonary edema, was evaluated in the radiologic reports. For the determination of presence of pulmonary diseases from chest X-rays (n=900) by GPT-4.0, the sensitivity ranged from 0.606 to 0.996, specificity from 0.897 to 0.993, and accuracy from 0.864 to 0.985. (Figure 1) In identifying pulmonary diseases using chest computed tomography scans (n=900), GPT-4.0 showed a sensitivity between 0.909 and 0.990, a specificity between 0.850 and 0.984, and an accuracy ranging from 0.873 to 0.982.The LLM’s capability to classify unstructured radiologic data showed excellent accuracy. We could suggest the potential usefulness of LLMs as a substitute for manual chart reviews by clinicians.

제목
Extraction of clinical data on major pulmonary diseases from unstructured radiologic reports using a large language model
저자
Park, Hyung Jun; Huh, Jin-Young; Chae, Ganghee; Choi, Myeong Geun
DOI
10.1183/13993003.congress-2024.PA1653
발행일
2024-09
유형
Meeting Abstract
저널명
European Respiratory Journal
권
64
호
S68