2026 MICCAI Society

Five Papers, Two Challenges, and an Award: The xAILab Bamberg at MICCAI 2026 in Strasbourg

From September 27th to October 1st, the xAILab Bamberg traveled to Strasbourg, France, to participate in the 29th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI 2026). This year's edition received a record 4,601 full paper submissions and accepted 1,167 papers from 48 countries. We contributed to this global forum with five papers across the main conference and four workshops, three of them presented as orals, as well as two challenge entries. Our paper on statistical color matching received the Best Paper Award Runner-Up at the Workshop on Data Engineering in Medical Imaging (DEMI).

Our Contributions at a Glance

Our group was represented in Strasbourg by PhD candidates Sebastian Doerrich, Francesco Di Salvo, Shyam Nandan Rai, and Hanh Huyen My Nguyen, together with Professor Christian Ledig. The week opened on Sunday, September 27, with two workshop orals, one at the Workshop on Endoluminal Intervention Navigation and Autonomy (EndoLINA) and one at the Workshop on Uncertainty for Safe Utilization of Machine Learning in Medical Imaging (UNSURE). On the same day, Hanh Huyen My Nguyen represented us in the RARE26 challenge. On Monday, September 28, we presented our main conference poster. The final conference day, Thursday, October 1, brought two further workshop contributions, a poster at the Workshop on Efficient Medical AI (EMA4MICCAI) and an oral at DEMI, as well as the results announcement of the ORena SAVE FOCUS challenge.

Polyp Detection Under Real Clinical Conditions at EndoLINA

At EndoLINA, Sebastian Doerrich opened our week with an oral presentation of "TRUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection". Computer-aided detection systems for colonoscopy are meant to help physicians find polyps, the precursors of colorectal cancer, that would otherwise be missed. Most of these systems are developed on curated video clips that center on the polyps themselves and leave out the long stretches of a real examination in which no polyp is visible. The paper introduces TRUE-Colon, a standardized benchmarking protocol that measures how a detector behaves in practice: how accurately it localizes polyps, how many false alarms it raises, how quickly it reacts once a polyp appears, and how consistently it keeps detecting it. We evaluated four real-time detectors on two curated benchmarks and on 60 complete, unedited colonoscopy procedures.

The results show a clear asymmetry. Detectors trained on curated clips lost most of their accuracy on full procedures; for YOLOv11, the localization score (mAP50) dropped from 0.724 on the curated SUN benchmark to 0.164 on full procedures. Detectors trained on full procedures, in turn, kept their accuracy when tested on curated clips, reaching mAP50 values between 0.705 and 0.717 on SUN. Our conclusion is that both training and evaluation of these systems should move toward complete procedures.

You can explore the full paper here.

Choosing the Right Layers for Out-of-Distribution Detection at UNSURE

Later the same day, Shyam Nandan Rai gave a long oral presentation at UNSURE on "Layer Selection in VLMs for Zero-Shot OOD Detection via Multi-Resolution Entropy Estimation". Medical AI systems need to recognize when an image differs from the data they were built on, for example because it comes from a different hospital, scanner, or patient population. Vision-language models can flag such out-of-distribution images without any additional training, but existing methods rely almost exclusively on the model's final layer. The work shows that intermediate layers carry complementary information and that the most useful depth depends on the imaging modality: shifts in histopathology are detected best in early layers, shifts in brain MRI in middle to final layers. Earlier approaches select layers with an entropy estimate computed at a single resolution, and the paper shows that their detection performance (AUROC) varies by up to 19.3 percentage points depending on this one setting. The proposed method combines entropy estimates across several resolutions, which keeps the layer selection stable and requires only in-distribution data. On histopathology and brain MRI benchmarks with two medical vision-language models, it achieved the best or competitive results in nearly all configurations.

You can explore the full paper here.

Variance-Aware Out-of-Distribution Detection in the Main Conference

On September 28, Francesco Di Salvo presented our main conference poster "MaRS: Robust Out-of-Distribution Detection via Mahalanobis Residual Scoring". Detecting images that fall outside a model's training distribution is essential for its safe use in the clinic. One family of methods learns to reconstruct the features of familiar images and flags an image as unfamiliar when its reconstruction error is large. Such methods have produced inconsistent results, and the paper traces this back to how the reconstruction error is summarized into a single score: the standard approach weights every direction of the error equally, although the errors of familiar images vary much more in some directions than in others. MaRS learns the structure of familiar images with a lightweight autoencoder on top of a frozen foundation model and scores the reconstruction error with a Mahalanobis distance, which accounts for this uneven variance. Across three imaging modalities and several types of distribution shift, MaRS achieved the best average AUROC of all compared methods (86.01, ahead of 85.34 for the strongest baseline) and the lowest average false positive rate at 95 percent true positive rate (42.72, compared to 43.59), a statistically significant improvement over that baseline. The method requires no labels and can be added to an existing model after training.

You can explore the full paper here.

One Model for Many Classification Tasks at EMA4MICCAI

On October 1, Sebastian Doerrich presented the poster "MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification" at the 2nd MICCAI Workshop on Efficient Medical AI. Large foundation models are usually adapted to a medical task by training a small set of additional parameters while the original network stays frozen. This parameter-efficient fine-tuning works well with limited data, but it requires a separate adapter for every task, and simply merging the adapters into one network can make the tasks interfere with each other. MoPET combines several small adapters, so-called experts, inside one frozen foundation model and learns a router that sends each image through only a few of them. This lets related tasks share capacity while keeping conflicting ones apart. On the MedMNIST benchmark, the paper first confirms that parameter-efficient fine-tuning outperforms updating the full network, raising average accuracy from 86.50 to 88.97 percent. A single MoPET model trained on four heterogeneous datasets then reached an average accuracy of 93.46 percent, compared to 92.83 percent for the best separately trained adapters. Training together with auxiliary datasets also improved the average accuracy on data-scarce target tasks from 81.58 to 83.58 percent.

You can explore the full paper here.

Best Paper Award Runner-Up for Statistical Color Matching at DEMI

Our week closed with an oral presentation by Sebastian Doerrich at the 4th Workshop on Data Engineering in Medical Imaging on "Simple, Safe, and Overlooked: Reclaiming Sustainable Domain Generalization with Statistical Color Matching". Medical image classifiers often fail after deployment because images from a new scanner, a different staining protocol, or a different patient population look different in color from the training data. Generating such variations during training is a common remedy, yet simple color changes add too little variety, while deep generative style transfer models are computationally expensive and can alter or invent anatomical structures. The paper revisits classical statistical color matching and turns it into Colorist, a data augmentation that transfers the mean and standard deviation of each color channel from one training image to another. Colorist requires no training and no neural network, and because it changes only color values, the anatomy of the image remains exactly intact. Compared with thirteen deep style transfer models, Colorist preserved image structure best while matching colors most closely. Across 12 in-distribution and 7 shifted out-of-distribution datasets covering histopathology, blood cell, dermatology, and retinal images, it improved balanced accuracy by up to 9 percentage points over leading domain generalization methods and by 13 percentage points over training without augmentation. On the same hardware, it was faster than every compared style transfer method and consumed between 5.4 and 10,601 times less energy per image. The DEMI organizers recognized the work with the Best Paper Award Runner-Up.

You can explore the full paper here.

Two Challenge Entries: RARE26 and ORena SAVE FOCUS

Beyond the papers, our group took part in two MICCAI challenges. Hanh Huyen My Nguyen competed in RARE26, the challenge on Recognition of Abnormalities in low prevalence cancer that is part of the EndoVis challenge series. It addresses the detection of early cancer in patients with Barrett's esophagus during routine endoscopic surveillance, a setting in which early neoplasia is found in typically less than 1 percent of patients and the few relevant images are easily outnumbered by normal findings. Our method, MILES, adapts a gastrointestinal foundation model pretrained on endoscopy images from multiple centers and examines each image region by region, so that small, localized changes in an otherwise normal-looking image can still be detected. With this approach, our team placed 6th out of 34 teams.

In the FRAME track of the ORena SAVE FOCUS challenge, Shyam Nandan Rai addressed the question of whether vision-language models can answer clinically relevant questions about a single laparoscopic image. The questions focus on foreign objects such as sponges, needles, and clips, which must not be left in the patient's body after minimally invasive surgery. Our approach combines a general-purpose and a surgery-specific image encoder within a vision-language model. It placed 16th out of 26 teams and outperformed the baseline provided by the organizers. The challenge results were announced at MICCAI on October 1.

Science, Networking, and the Strasbourg Experience

MICCAI 2026 took place at the Strasbourg Convention Center. The city had originally been selected to host MICCAI in 2021, when the pandemic moved the conference online, and this year it stepped in again after the conference had to be relocated from Abu Dhabi. With a record number of submissions, an acceptance rate of approximately 27 percent, and more than 3,900 reviewers from 71 countries involved in the review process, the conference reflected the strong growth of the field. Many of the themes we encountered across the main conference and the workshops, from robustness under distribution shift and out-of-distribution detection to efficient adaptation of foundation models and the careful evaluation of clinical AI systems, closely match the research focus of our group. MICCAI is also the place where we reliably meet many of our friends and collaborators, and Strasbourg was no exception: we were glad to see colleagues from FAU Erlangen-Nürnberg, Imperial College London, the Technical University of Munich, and many other institutions again, and we enjoyed great conversations with them throughout the week. The host city added its own scientific heritage: the University of Strasbourg, founded in the heart of the Grande Île, has been teaching since 1538, and Louis Pasteur once taught there.

Looking Forward

With five papers, three orals, two challenge entries, and an award, MICCAI 2026 was a week full of exchange for our group. We are particularly proud that two of this year's papers grew out of master's theses written at our lab. Master's students Andreas Schwab and Daniel Würtinger brought their thesis work into TRUE-Colon and MoPET, respectively, and each is a co-first author of the resulting paper. We sincerely thank both of them for their dedication and effort. Their work shows that a thesis at the xAILab Bamberg can lead to a publication at one of the leading conferences in medical image analysis, and we warmly invite motivated students who would like to contribute to research of this kind to get in touch with us about a thesis. The questions raised at our posters and after our talks will shape the next steps of these projects. We return from MICCAI 2026 with valuable feedback and look forward to tackling the next challenges.