Cite as: Archiv EuroMedica. 2026. 16; 4. DOI 10.35630/2026/16/Iss.4.30
Artificial intelligence, particularly deep learning based on convolutional neural networks, is increasingly applied to dental imaging, including bitewing, periapical, panoramic, cone beam computed tomography, and cephalometric examinations .
To synthesise evidence published between 2021 and 2026, alongside one methodologically relevant study published in 2020, on the applications, reported performance, and limitations of artificial intelligence in dental radiological diagnostics.
This narrative review was prepared using SANRA as a methodological reference. PubMed and Google Scholar were searched for publications issued between January 2021 and May 2026. The last search was performed on 21 May 2026. Data on study type, imaging modality, diagnostic task, model architecture, dataset size, performance measures, validation strategy, and reported limitations were synthesised narratively.
Artificial intelligence has been applied to caries detection, periapical lesion detection, tooth detection and numbering, tooth segmentation, cephalometric landmark localisation, periodontal bone loss assessment, and mandibular canal segmentation. Reported results were favourable for several clearly defined tasks, particularly tooth detection and numbering, but varied substantially across diagnostic targets, imaging modalities, datasets, reference standards, and validation methods. Caries detection showed wide ranges of sensitivity and accuracy, whereas tooth detection and numbering reached F measures of approximately 0.96. Recurrent limitations included retrospective single institution datasets, heterogeneous annotation procedures, limited external validation, and poor comparability between task specific performance measures.
Current evidence supports the use of artificial intelligence as an auxiliary decision support tool rather than as a replacement for clinical and radiological judgement. Reader studies suggest that artificial intelligence assistance may improve selected diagnostic measures, particularly among less experienced observers, but this benefit has not been demonstrated consistently across all tasks or clinical settings. Further clinical translation requires external validation, more consistent reference standards and reporting, and resolution of ethical, regulatory, and professional liability issues.
Keywords: artificial intelligence; deep learning; dental radiography; cone beam computed tomography; caries detection; periapical lesion; cephalometry; dentistry.
Dental imaging underpins diagnosis and treatment planning in dentistry. However, image interpretation is operator dependent and susceptible to interobserver variability, particularly given the high volume of examinations encountered in routine clinical practice [4,38]. The detection of subtle radiographic findings may also depend on examiner experience [4].
Within this context, artificial intelligence (AI) refers to computational methods capable of performing tasks that ordinarily require human intelligence [4,43]. In dental imaging, these tasks include image classification, object detection, segmentation, and anatomical landmark localisation. Machine learning (ML) is a subfield of AI in which computational models identify patterns in training data and use them to generate predictions or classifications for new data. Deep learning (DL) employs multilayer artificial neural networks that learn representations directly from imaging data. Convolutional neural networks (CNNs) are the architectures most frequently used for automated dental image analysis [2,19]. Narrative and bibliometric analyses document a rapid expansion of AI applications in dentistry over the past decade [5,6], and recent systematic reviews have mapped these applications across dental specialties and summarised their reported diagnostic performance [1,41].
The clinical relevance of AI in dental imaging is illustrated by studies of approximal caries and periapical lesion detection. Radiographic detection of approximal caries depends on examiner experience, and less experienced dentists may perform less accurately than experienced clinicians [11]. Operator dependence has also been reported for periapical lesion detection, although differences between panoramic radiography and CBCT additionally reflect the inherent diagnostic capabilities of the respective imaging modalities [26]. AI is therefore being investigated as an auxiliary tool for supporting the interpretation of dental images [4,38].
The available evidence is distributed across different imaging modalities, diagnostic tasks, model architectures, dataset compositions, annotation protocols, validation strategies, and performance measures. Classification, detection, segmentation, and landmark localisation are distinct computational tasks and cannot be evaluated using identical metrics or interpreted as equivalent forms of diagnostic performance. Moreover, high performance on an internal test dataset does not establish external generalisability or clinical utility. This heterogeneity limits direct comparison between studies and complicates assessment of the extent to which reported technical performance can be translated into routine clinical practice. A structured synthesis is therefore needed to integrate reported performance with the methodological, translational, ethical, and medico legal limitations of current AI systems in dental imaging.
Aim: This narrative review aims to synthesise recent evidence on the application of artificial intelligence in dental radiological diagnostics.
Objectives: The objectives of this review are to identify the diagnostic tasks and imaging modalities in which AI is currently applied, summarise the reported technical and diagnostic performance, and assess the remaining methodological, translational, ethical, and medico-legal barriers.
This study was conducted as a narrative review. The Scale for the Assessment of Narrative Review Articles, SANRA, was used as a methodological reference for the preparation and self-assessment of the review. A literature search was performed in PubMed and Google Scholar for publications issued between January 2021 and May 2026. The search used combinations of the terms “artificial intelligence”, “deep learning”, and “convolutional neural network” with “dental radiography”, “bitewing”, “panoramic”, “CBCT”, “cephalometric”, “caries”, “periapical”, “periodontal bone loss”, and “mandibular canal”. The reference lists of relevant reviews were also examined to identify additional publications. The last search was performed on 21 May 2026.
Inclusion criteria were:
Exclusion criteria were publication before January 2021, except for the predefined methodological exception, absence of relevant dental imaging content, conference abstracts without a full publication, retracted articles, and publications for which the full text could not be retrieved. For studies included in the diagnostic performance synthesis, the absence of extractable quantitative data was an additional exclusion criterion.
The search identified 1,284 records, including 612 from PubMed and 672 from Google Scholar. After removal of 241 duplicates, 1,043 records were screened by title and abstract. Of these, 887 records were excluded because they were not relevant to artificial intelligence in dental imaging. The full texts of 156 publications were assessed. A total of 112 publications were excluded because they lacked relevant dental imaging content, were published before January 2021, did not provide quantitative data required for the diagnostic performance synthesis, were available only as conference abstracts, could not be retrieved in full text, duplicated an already included dataset, or had been retracted. Forty-four publications were included in the narrative synthesis.
Title and abstract screening and full-text assessment were performed independently by two authors, I.M. and J.P. Disagreements were resolved through discussion. When agreement could not be reached, a third author, Z.J., reviewed the publication.
Data were extracted into a structured spreadsheet by one author, B.S., and independently checked against the original publications by a second author, I.M. The extracted information included study type, imaging modality, diagnostic task, model architecture, dataset size, reported performance measures, validation strategy, and limitations reported by the authors. Performance measures included accuracy, sensitivity, specificity, F1-score, intersection over union, Dice coefficient, mean average precision, and area under the receiver operating characteristic curve, when available.
The findings were synthesised narratively and organised according to diagnostic task. Quantitative results were presented descriptively, without statistical pooling. No meta-analysis or independent formal assessment of risk of bias was performed. Risk-of-bias and quality assessments are reported only when they were provided by the original systematic reviews or meta-analyses.
Caries detection on bitewing radiographs is the most extensively studied radiological application. A systematic review of 23 studies published before March 2025 found that the most common architectures were ResNet and YOLO, with dataset sizes ranging from 112 to 8,539 images and diagnostic accuracy from 70% to over 99%; several models matched or surpassed human experts, particularly for advanced lesions [9]. A complementary synthesis of twenty studies (periapical, bitewing, and orthopantomographic radiographs; datasets of 15-2,900 examinations) reported sensitivity of 0.44-0.86, specificity of 0.85-0.98, accuracy of 0.73-0.98, and AUC of 0.84-0.98, with most studies at low risk of bias under QUADAS-2 [2]. Sensitivity and specificity ranges of 0.63-0.95 and 0.86-0.99, respectively, were corroborated in a further analysis that emphasised the influence of annotation quality and reference standard [13].
Primary studies illustrate both the promise and the limitations of these methods. A U-Net model trained on 304 bitewings and tested on 50 achieved a precision of 63.29%, recall of 65.02%, and F1 of 64.14%; importantly, when three dentists used the model output as a reference, their sensitivity improved significantly, demonstrating value as a second opinion [8]. Comparative studies of newer detectors (YOLOv9c, YOLOv11) reported high performance for enamel and dentin caries, in some cases exceeding that of dentists [11]; an architecture comparison found AlexNet reaching 95.56% accuracy for proximal-caries classification, whereas an Inception-based classifier reached 73.3% on a small dataset [12]. A YOLOv11 model trained on 730 bitewing radiographs containing 1,115 annotated lesions and evaluated within the ICCMS framework reported strong annotation reliability (interrater IoU 0.82; Dice similarity coefficient 0.85), illustrating the importance of standardised lesion labelling [10]. An ensemble of object-detection CNNs performed at least as well as experienced dentists, although the detection of incipient lesions remained challenging [14].
AI for periapical-lesion detection spans intraoral periapical radiographs, panoramic radiographs (OPGs), and CBCT [26]. On panoramic radiographs, a two-stage CNN (detector plus classifier) trained on 18,618 periapical root areas from 713 radiographs achieved a detector average precision of 74.95% and, in combination, an accuracy of 84.6%, sensitivity of 72.2%, and specificity of 85.6% [25]. A narrative review of thirteen OPG-based studies reported a detection mAP between 0.832 and 0.953 and accuracy between 0.673 and 0.812, with RetinaNet performing best; it further noted that experienced clinicians reading OPGs exhibit high specificity (~95.8%) but low sensitivity (~34.2%) relative to CBCT, underscoring the scope for AI assistance [26]. A commercial system (Diagnocat), evaluated on 616 teeth from 357 panoramic radiographs, was assessed for sensitivity, specificity, and predictive values, with particular attention to the confounding effect of the palatoglossal air space [24]. Deep-learning detection on panoramic radiographs was found to localise periapical lesions reliably and to support both clinicians and the wider dental healthcare system [23].
Automated tooth detection and FDI/universal numbering on panoramic radiographs is a mature application. A Faster R-CNN Inception-v2 system trained on 2,482 panoramic radiographs achieved sensitivity of 0.9559, precision of 0.9652, and an F-measure of 0.9606 on a 249-image test set [34]. A Mask R-CNN with a ResNet-101 backbone trained on 8,000 panoramic images reached 93.83% accuracy for numbering [33], whereas a YOLOv4 model applied to 500 radiographs reported 88.5% accuracy and an F1 of 93.44%, with detection in roughly 20 ms per image [35]. More recent work applied YOLOv10 to mixed-dentition radiographs [37] and tailored CNNs to combined numbering and cavity detection [36].
On CBCT, segmentation has advanced from whole teeth to fine anatomical structures. A deep CNN that segmented and classified teeth bearing orthodontic brackets achieved an intersection-over-union of 0.99, a classification recall of 99.9%, and a precision of 99%, processing a full scan in approximately 14 seconds [16]. A two-stage pipeline combining 3D TransUNet, nnU-Netv2, and a 3D DenseNet169 segmented teeth and generated auxiliary diagnostic reports from 450 CBCT datasets [17]. A SpatialConfiguration-Net combined with U-Net localised teeth with 97.3% accuracy on 144 CBCT images and detected lesions with sensitivity of 0.97 and specificity of 0.88 [18]. An AI tool validated for primary-tooth segmentation on 402 teeth from 37 paediatric CBCT scans demonstrated expert-level, time-efficient, and consistent performance [15]. A broader methodological review surveyed CNN- and Transformer-based segmentation across panoramic, CBCT, and intraoral-scan data [19].
Cephalometric landmarking has been assessed in several meta-analyses. A systematic review and meta-analysis of 19 studies (2017–2020), all employing CNNs and predominantly using 2-D lateral radiographs, found a landmark-prediction error centred around a 2-mm threshold, with the proportion of landmarks detected within 2 mm at 0.799 (95% CI 0.770-0.824); the body of evidence was consistent but at high risk of bias, and the authors emphasised the need to demonstrate generalisability [27]. Subsequent systematic reviews and meta-analyses extended this assessment to 2-D landmark detection and prediction [28] and to automated 3-D cephalometric landmarking versus manual tracing [29], applying QUADAS-2 for quality appraisal. Across these analyses, DL achieved consistent and largely high accuracy, although the 3-D evidence base remained sparser than the 2-D.
For periodontal diagnosis, a systematic review and critical appraisal of thirteen studies reported accuracies from 73.0% to 98.6% depending on task and architecture, with CNNs and hybrid models predominating; one CNN reached 92% accuracy for classifying periodontal bone loss on panoramic radiographs, and a CNN-based system reached 81.0% accuracy for premolars and 76.7% for molars in diagnosing periodontally compromised teeth [30]. A retrospective study employing a VGG-16 CNN on 1,724 intraoral periapical images reported 73.0% accuracy for normal-versus-disease classification and 59% for grading severity [31]. A deep-learning ensemble (YOLOv5 with VGG-16 and U-Net) applied to 8,000 periapical radiographs (27,964 teeth) reached approximately 90% accuracy in detecting tooth position, shape, and radiographic bone loss [32].
Accurate localisation of the mandibular (inferior alveolar) canal on CBCT is critical for avoiding nerve injury during implant placement, extractions, and orthognathic surgery [20]. Deep-learning segmentation of the canal and the surrounding mandibular bone has been pursued using U-Net variants and, more recently, YOLOv8-seg, which was reported as a clinically applicable and computationally efficient tool for delineating both the inferior alveolar canal and the alveolar bone for pre-operative planning [21]. A systematic review and meta-analysis of U-Net-based CBCT segmentation for implant planning reported a Dice similarity coefficient of 0.91 for unilateral edentulous spans versus 0.73 for bilateral cases, illustrating how anatomical complexity affects performance [22]. These applications extend AI from interpretive detection towards quantitative surgical-planning support.
The main applications of AI in dental radiology and their reported performance are summarised in Table 1.
Table 1. Representative AI applications in dental radiology by task, modality, and reported performance (2021–2026).
| Task | Modality | Model / synthesis | Ref. | Accuracy | Sensitivity | Specificity | F1 | IoU / Dice | mAP / AUC |
| Caries detection | BW | ResNet, YOLO (SR, 23 studies) | [9] | 70%–>99% | — | — | — | — | — |
| Caries detection | PA, BW, OPG | ANN, CNN, DCNN (SR, 20 studies) | [2] | 0.73–0.98 | 0.44–0.86 | 0.85–0.98 | — | — | AUC 0.84–0.98 |
| Caries detection | BW | Annotation-quality analysis | [13] | — | 0.63–0.95 | 0.86–0.99 | — | — | — |
| Caries detection | BW | U-Net | [8] | — | 65.02% | — | 64.14% | — | — |
| Caries detection | BW | AlexNet / Inception | [12] | 95.56% / 73.3% | — | — | — | — | — |
| Caries detection | BW | YOLOv11 (ICCMS) | [10] | — | — | — | — | IoU 0.82; Dice 0.85 a | — |
| Periapical lesions | OPG | Two-stage CNN (detector + classifier) | [25] | 84.6% | 72.2% | 85.6% | — | — | AP 74.95% |
| Periapical lesions | OPG | RetinaNet best (NR, 13 OPG studies) | [26] | 0.673–0.812 | — | — | — | — | mAP 0.832–0.953 |
| Tooth detection / numbering | OPG | Faster R-CNN Inception-v2 | [34] | — | 0.9559 | — | 0.9606 | — | — |
| Tooth detection / numbering | OPG | Mask R-CNN + ResNet-101 | [33] | 93.83% | — | — | — | — | — |
| Tooth detection / numbering | OPG | YOLOv4 | [35] | 88.5% | — | — | 93.44% | — | — |
| Tooth segmentation | CBCT | DCNN (teeth with brackets) | [16] | — | 99.9% b | — | — | IoU 0.99 | — |
| Tooth segmentation / lesions | CBCT | SCN + U-Net | [18] | 97.3% c | 0.97 | 0.88 | — | — | — |
| Cephalometric landmarks | Lateral ceph (2-D) | CNN (SR/MA, 19 studies) | [27] | — | — | — | — | — | 0.799 within 2 mm |
| Periodontal bone loss | OPG, PA | CNN, hybrid (SR, 13 studies) | [30] | 73.0%–98.6% | — | — | — | — | — |
| Periodontal bone loss | PA | VGG-16 | [31] | 73.0% d; 59% e | — | — | — | — | — |
| Periodontal bone loss | PA | YOLOv5 + VGG-16 + U-Net | [32] | ~90% | — | — | — | — | — |
| Mandibular canal | CBCT | U-Net variants (SR/MA) | [22] | — | — | — | — | Dice 0.91 f; 0.73 g | — |
Abbreviations: AI — artificial intelligence; ANN — artificial neural network; AP — average precision; AUC — area under the receiver-operating-characteristic curve; BW — bitewing radiograph; CBCT — cone-beam computed tomography; ceph — cephalogram; CNN — convolutional neural network; DCNN — deep convolutional neural network; Dice — Dice similarity coefficient; F1 — F1-score (F-measure); ICCMS — International Caries Classification and Management System; IoU — intersection over union; MA — meta-analysis; mAP — mean average precision; NR — narrative review; OPG — orthopantomogram (panoramic radiograph); PA — periapical radiograph; ResNet — residual network; SCN — SpatialConfiguration-Net; SR — systematic review; VGG — Visual Geometry Group network; YOLO — You Only Look Once. — indicates that the metric was not reported in the cited source. Footnotes: a annotation (interrater) agreement, not model performance; b classification recall (precision 99%); c tooth-localisation accuracy, sensitivity/specificity refer to lesion detection; d normal-versus-disease classification; e severity grading; f unilateral edentulous span; g bilateral edentulous span.
As summarised in Table 1, the strongest and most consistent quantitative evidence concerns caries and tooth detection/numbering, where multiple architectures reach accuracies or F-measures above 0.90; periapical, periodontal, and implant-planning tasks show high but more variable performance, and cephalometric landmarking clusters around a clinically meaningful 2-mm error.
Clinical interpretation of the evidence. Across the tasks reviewed, three patterns emerge. First, CNN based deep learning has shown promising but variable performance across dental imaging tasks. The highest reported performance was observed mainly in tooth detection and numbering, whereas caries detection showed a wider range of results. Performance was also more variable in tasks involving subtle or poorly demarcated findings, including incipient lesions and severity grading, as summarised in Table 1. Second, available reader studies suggest that clinical benefit may arise from interaction between the model and the clinician rather than from stand alone model performance alone. Clinicians’ sensitivity increased when model output was available as a reference [8], and a synthesis of augmented intelligence studies reported the largest gains among less experienced observers [42]. Third, reported performance is strongly influenced by dataset size, case mix, and labelling protocol, which helps explain the wide ranges observed even within apparently similar tasks [2,13]. The central question is therefore whether performance reported under controlled study conditions can be retained in external and routine clinical settings.
It should be stated explicitly that, on the current evidence, AI in dental radiology constitutes a decision-support tool and does not replace the clinical judgement of the practitioner. Model output represents an additional source of information to be weighed against the clinical examination, the patient’s history, and the operator’s own radiographic assessment; final diagnostic and treatment decisions, together with the accompanying professional responsibility, remain with the treating clinician [3,42].
Methodological limitations of the evidence base. The reported performance of AI models should be interpreted in light of recurring methodological limitations. Many studies rely on retrospective datasets from a single institution, with limited sample sizes, heterogeneous annotation procedures, and selected images that may not fully represent routine clinical practice [35,40]. These features may introduce spectrum bias and lead to optimistic estimates during internal validation. External validation remains uncommon, and performance may decline when models are evaluated using data from different institutions, scanners, or patient populations [38,39]. Cephalometric meta analyses have also reported a high risk of bias despite generally favourable accuracy estimates [27]. In addition, differences in imaging modality, reference standards, dataset composition, annotation procedures, and performance metrics limit direct comparison between studies.
Ethical and medico-legal considerations. Clinical adoption of AI in dental imaging depends not only on diagnostic performance but also on several ethical, legal, and organisational factors. These include protection of patient data used for model development and clinical application, the risk of algorithmic bias when training datasets do not adequately represent the target population, uncertainty regarding liability for AI-assisted diagnostic errors, and automation bias, in which clinicians may place excessive reliance on model output [7,40,44]. Qualitative evidence identifies limited explainability and concern about unrecognised model failure as important barriers to clinical acceptance [44]. Regulatory clearance alone does not establish clinical effectiveness, and a review of FDA-cleared dental imaging systems found that some commercially available solutions lacked independent peer-reviewed clinical validation [40].
This narrative review has several limitations. The literature search was restricted to PubMed and Google Scholar, and relevant publications indexed exclusively in other databases may have been missed. The search was limited to English language publications, which may have excluded relevant evidence published in other languages. Although study selection and data extraction were independently checked, no independent formal assessment of methodological quality or risk of bias was performed. Quality assessments were reported only when they were available in the original systematic reviews or meta analyses. The review included both primary studies and review articles, which may have resulted in overlap between some of the evidence summarised from different sources. Because the findings were synthesised narratively, the reported performance estimates were interpreted descriptively and were not statistically pooled.
One deliberate exception was made to the predefined publication period. Reference 14, published in 2020, compared a deep learning model with individual dentists in detecting caries lesions of different radiographic extension on bitewing radiographs. It was retained because of its methodological relevance to direct comparisons between model and clinician performance and its frequent citation in more recent syntheses included in this review. The remaining 43 included publications were published within the predefined period from January 2021 to May 2026.
Artificial intelligence is being applied across multiple areas of dental imaging, including caries detection, periapical lesion detection, tooth detection and numbering, tooth segmentation, cephalometric landmark localisation, periodontal bone loss assessment, and mandibular canal segmentation. Reported results are favourable for several clearly defined tasks, particularly tooth detection and numbering, but vary according to the diagnostic target, imaging modality, dataset composition, reference standard, and validation method. Performance measures used for classification, detection, segmentation, and landmark localisation are not directly comparable. High performance obtained using retrospective and internally validated datasets should not be interpreted as evidence of routine clinical effectiveness.
Current evidence supports the use of AI as an auxiliary decision support tool rather than as a replacement for clinical and radiological judgement. Reader studies suggest that AI assistance may improve selected diagnostic measures, particularly among less experienced observers, but this benefit has not been demonstrated consistently across all diagnostic tasks, imaging modalities, or clinical settings. Improvements in sensitivity should be considered together with specificity, false positive and false negative findings, and the potential clinical consequences of diagnostic errors.
Further clinical translation requires external and preferably multicentre validation using representative patient populations, different imaging systems, and images reflecting routine clinical quality. More consistent reference standards, annotation procedures, and reporting of task appropriate performance measures are also needed. External generalisability, transparency, appropriate clinician oversight, data protection, regulatory requirements, and professional liability remain important considerations. Until these issues are addressed, AI output should be treated as supplementary information within a clinician led diagnostic process and interpreted together with image quality, imaging modality, clinical examination, and patient history.
Conceptualization: Izabela Migdał, Justyna Polko, Maja Witek. Methodology: Izabela Migdał, Zuzanna Jeziorska, Maja Witek. Data collection and literature selection: Izabela Migdał, Justyna Polko, Zuzanna Jeziorska, Bartosz Szymajda. Analysis and interpretation: Izabela Migdał, Maja Witek, Jakub Artur Czapkiewicz, Weronika Kwaśnica, Tomasz Horodniczy. Writing original draft: Izabela Migdał, Maja Witek, Zuzanna Jeziorska. Writing review and editing: Izabela Migdał, Zuzanna Galicka, Katarzyna Raczek, Jakub Artur Czapkiewicz. Supervision: Weronika Kwaśnica, Tomasz Horodniczy.
Final approval of the manuscript: all authors.
No external funding was received for this manuscript.
The authors declare no conflict of interest.
Artificial intelligence was used only as an auxiliary tool to support literature search and language editing. All references, scientific interpretation, conclusions and the final version of the manuscript were independently checked and approved by the authors. The authors independently verified the scientific content, references, interpretation of evidence and final wording, and accept full responsibility for the submitted manuscript.