M. Cuevas-Nunez, Cosimo Galletti, J. Flores-Fraile, Cosimo Galletti, Shokoufeh Shahrabi Farahani, Wilmer Rodrigo Díaz-Castañeda, D. Portelli, L. Fiorillo, V. Mehta, M. Fernandez-Figueras
2026.1.1Oral Oncology Reports
Abstract
Artificial intelligence (AI) is increasingly integrated into pathology, but its accuracy in immunohistochemical (IHC) marker selection remains underexplored. This study evaluated LeChat’s ability to recommend IHC markers for benign and malignant salivary gland tumors, focusing on accuracy, completeness, relevance, consistency, and marker-level errors across tumor types and subtypes. A total of 21 tumor types were selected and classified by behavior (benign vs. malignant) and histologic subtype. Expert-derived reference panels served as the gold standard. For each tumor, LeChat was queried three times using standardized prompts. Recommendations were scored across three domains: accuracy (inclusion of essential markers), completeness (inclusion of secondary markers), and relevance (absence of irrelevant markers). Composite scores (range: 3–9) and intra-tumor variability were calculated. Mann–Whitney U and Kruskal–Wallis tests assessed differences by tumor category and subtype. Marker-level analysis evaluated over- and under-recommendations. Across 63 total prompts, LeChat achieved a mean accuracy of 1.56, completeness of 1.59, and relevance of 2.35. Only 3.2% of prompts included all essential markers, and 17.5% achieved composite scores ≥7. No responses included both primary and secondary markers. Performance did not differ significantly between benign and malignant tumors (p = 0.405), nor across histologic subtypes (p = 0.988). Tumor-level variability was highest for basaloid squamous cell carcinoma and lowest for pleomorphic adenoma. Marker-level analysis identified 14 false positives (e.g., CD20, IgG4) and 9 false negatives (e.g., SOX10, β-catenin). A moderate correlation was observed between marker frequency and accuracy (r = 0.48, p < 0.01). LeChat’s IHC recommendations were generally relevant but frequently incomplete and inconsistent, particularly for malignant and histologically complex tumors. Despite its potential for assisting in straightforward cases, current performance remains inferior to expert standards. These findings highlight the limitations of general-purpose LLMs in pathology workflows and emphasize the need for domain-specific refinement before clinical integration. • LeChat’s accuracy in IHC marker selection for salivary gland tumors was evaluated. • An expert consensus panel served as the gold standard for performance comparison. • LeChat showed moderate accuracy but low completeness in marker recommendations. • AI performance was most consistent in benign and myoepithelial salivary tumors. • Findings highlight AI’s potential and current limits in diagnostic histopathology.
Citation format
CUEVAS-NUNEZ, M., et al. Lechat in oral pathology: Assessment of its immunohistochemical marker selection accuracy in salivary gland tumors. Oral Oncology Reports, 2026, 17: 100781.