Deep learning-enabled 3D multimodal fusion of cone-beam CT and intraoral mesh scans for clinically applicable tooth-bone reconstruction

Planning orthodontic treatment and dental implant surgery demands precise 3D models of both tooth crowns and their underlying root and bone structure. Cone-beam CT (CBCT) captures the bone and root anatomy but has limited surface resolution and requires manual, time-intensive segmentation. Intraoral optical scanners produce high-fidelity digital impressions of the visible tooth surfaces but cannot image beneath the gum line. This paper bridges that gap with a deep learning framework for multimodal 3D fusion. A learned registration network rigidly aligns the CBCT volume with the intraoral mesh, after which a multimodal segmentation network jointly processes both data sources to produce individual tooth-root and alveolar bone labels. The framework eliminates manual segmentation for the majority of clinical cases and produces reconstructions with accuracy validated against expert annotations on a large retrospective CBCT dataset. Published in Patterns (Cell Press, 2023), a high-impact data science journal, the work directly addresses a significant clinical bottleneck in digital dentistry and provides an open, reproducible benchmark for dental CBCT segmentation. The combined tooth-bone model is immediately usable in surgical planning software, enabling same-day patient consultations with accurate 3D visualisation.
Problem setting
Accurate 3D reconstruction of teeth and surrounding bone is essential for orthodontic and implant surgery planning, but current workflows require laborious manual segmentation of cone-beam CT (CBCT) images, which is time-consuming and operator-dependent. Intraoral optical scanners provide high-resolution surface geometry of visible dental structures but cannot capture subsurface bone. This work presents a deep learning framework that fuses CBCT volumetric data with intraoral mesh scans, combining the complementary strengths of both modalities to generate complete tooth-bone 3D reconstructions.
The figures below collect representative visual evidence from Patterns (Cell Press), 4(9).
Method and visual evidence
The visuals focus on multimodal registration between CBCT volumes and intraoral mesh scans, the fusion network, and reconstructed tooth-bone models for clinical planning.

Method overview.

Representation and setup.
Results and impact
The evaluation reported in Patterns (Cell Press), 4(9) is summarized through the figures above.