MedicineComputer Science

Wing S Kwok, Geraldine Wallbank, Philip Hodgson, Thomas Schrader, Le-Wen Shao, Mark R Elkins, J. Fandim, J. Scott, Catherine Sherrington, Adrian C. Traeger

2026.4.30JOURNAL OF CLINICAL EPIDEMIOLOGY

DOI: 10.1016/j.jclinepi.2026.112309

Abstract

OBJECTIVES: To compare accuracy, precision, recall, F1, and time spent using commercial tools to identify physiotherapy trials based on title and abstract, compared with a human approach. STUDY DESIGN: This study compared two approaches for title and abstract screening of 10,793 newly published records. In the reference standard human approach, two reviewers independently screened records using prespecified rules to assess relevance to physiotherapy. A third person resolved disagreements. We evaluated three large language models (LLMs) (gpt-4o, gpt-4.5, and gpt-4-turbo) within two commercial, web-based tools (ChatGPT and Copilot). Outcomes were accuracy (proportion of records that model correctly identified as relevant or irrelevant), precision (proportion of records identified as relevant that were considered as relevant by human approach), recall (the proportion of all actual relevant records that the model successfully identified), F1 (harmonic mean of precision and recall), and time spent. Exploratory analyses compared the performance of the commercial tools with local approaches, including local LLMs implementation, machine learning, and natural language processing. RESULTS: Commercial tools showed comparable performance across all metrics (ChatGPT vs Copilot: accuracy: 83% vs 86%; precision: 44% vs 48%; recall: 88% vs 87%; F1: 59% vs 62%). The total time spent using commercial tools with a labeled dataset was equivalent to 37% of the time required for the human-only screening process. Exploratory analysis showed that the Application Programming Interface-based implementation has comparable performance (accuracy: 82%; precision: 42%; recall: 93%; F1: 58%). Yet, LLM-based models demonstrated lower performance compared with other local, custom-adapted automation approaches such as machine learning and natural language processing. CONCLUSION: This proof-of-concept study demonstrates that commercial web-based LLMs may have sufficient accuracy to support title and abstract screening and substantially reduce the time to identify field-specific trials. However, alternative approaches, including machine learning or natural language processing, could achieve screening performance similar to or slightly higher than that of commercial tools, yet they require a series of preprocessing steps for implementation.

Citation format

KWOK, Wing S, et al. Automated approaches to identifying clinical trials based on title and abstract in the field of physiotherapy: A comparative analysis. JOURNAL OF CLINICAL EPIDEMIOLOGY, 2026, 196: 112309.