Back to portfolioModel weights

RESEARCH CASE STUDY

Microglomeruli Segmentation

Fine-tuning a foundation model to count synaptic boutons in a fly brain, and proving it against the state of the art

M.SC. THESIS · RWTH AACHEN · 3D INSTANCE SEGMENTATION · MICROSAM / NNU-NET V2 / SWINUNETR

A reproducible deep-learning pipeline that fine-tunes a foundation model for 3D instance segmentation of synaptic boutons in confocal Z-stacks of the Drosophila mushroom body calyx, then benchmarks it head-to-head against three SOTA 3D architectures — nnU-Net v2, Cellpose, SwinUNETR — under a physically-calibrated, instance-matching evaluation protocol.

Supervised by Prof. Dr. Abigail Morrison (Software Engineering Group, RWTH Aachen), with Prof. Dr.-Ing. Johannes Stegmaier as second examiner. Built for the Tavosanis lab, who counted these structures by hand — and who are using the tool now.

THESIS DATASHEETSTATUS
Work
M.Sc. thesis, RWTH Aachen. Supervised by Prof. Dr. Abigail Morrison, second examiner Prof. Dr.-Ing. Johannes Stegmaier. Built for the Tavosanis lab, who counted these structures by hand.
Shipped
Task
3D instance segmentation of synaptic boutons in confocal Z-stacks of the Drosophila mushroom body calyx — how many, and how big each one is.
Shipped
Dataset
13 volumetric stacks, hand-annotated from scratch in napari across seven brain preparations and two acquisition systems. No public benchmark for this structure existed.
Measured
Benchmark
4 architectures × 4 preprocessing variants under one evaluation protocol — MicroSAM against nnU-Net v2, Cellpose 3D and SwinUNETR.
Measured
Headline result
Fine-tuned MicroSAM Large recovered every bouton — recall 1.000, matched mIoU 0.787, the best of any model evaluated.
Measured
Baseline beaten
The lab’s current semi-manual Imaris workflow scores recall 0.735, mIoU 0.505 on the same test set.
Measured
In production
BoutonViewer is in active use by biologists today, on their own acquisitions rather than on the held-out volumes. A napari desktop app that turns a TIFF stack into a per-bouton table in µm³ and µm²; checkpoints published on Hugging Face with the rejected runs left out.
In use
Known weakness
That perfect recall is bought with 16 false positives, several of them one large bouton split into parts. The model card says predictions should be reviewed before unsupervised quantification.
Measured
Not verified
The scored test set is 2 images and 33 ground-truth boutons, from two specimens. Performance under substantially different imaging or labelling has never been formally measured — daily use is adoption evidence, not benchmark evidence, and the two are not interchangeable.
Unverified
Computer vision
3D instance segmentationFoundation-model fine-tuningMicroSAM / SAMnnU-Net v2SwinUNETRCellpose 3D
Evaluation
Instance matchingVolume-weighted metricsRecall-weighted panoptic qualityHeld-out specimen splits
Imaging data
Hand annotation in napariAnisotropic voxel calibrationRichardson-Lucy deconvolutionDifference-of-Gaussians
Engineering
PyTorchMONAIKorniaHTCondor cluster jobsnapari pluginHugging Face Hub
01

The whole experiment, in one figure

The claim on this page is not that a model worked. It is that a fine-tuned foundation model beat three SOTA 3D backbones and the lab’s existing tool under one protocol, and that claim is only worth making because sixteen runs sit behind it. So the grid is the centre of the figure, and everything else is what feeds it or what came out.

two instruments, two preprocessing paths — one shared path would deconvolve already-deconvolved dataCONFOCAL LSM0.3 × 0.0709 × 0.0709 µmrolling-ball + Richardson-LucyAIRYSCAN0.3 × 0.0425 × 0.0425 µmalready deconvolved — light normalise13 HAND-ANNOTATED VOLUMES7 brain preparations · 9 train / 2 validation / 2 testno public benchmarkexisted — the groundtruth had to be madebefore anything ran16 TRAINING RUNSrawDoGPSFallMicroSAMCellpose 3DnnU-Net v2SwinUNETRfour architectures × fourpreprocessing variants is agrid, not a run — launched asHTCondor submit files on thecluster, not started by handgreen = the row that shippedINSTANCE MATCHING, NOT PIXEL OVERLAPweighted by physical voxel volume, so coarse Z spacingcannot quietly inflate or deflate a scorea pixel score can’tsay whether oneobject became tworecall leads:a missed boutonis permanentCHECKPOINTSon Hugging Face, with therejected runs left outBoutonViewera biologist loads a TIFF stackand gets a table in µm³A checkpoint on a cluster helps nobody — which is why the napari app is part of this project, and why biologists use it now.
Two constraints shape the whole thing. Every training example was annotated by hand, which puts a hard ceiling on how much data can ever exist — that is what makes a model with a strong microscopy prior a practical choice rather than a fashionable one. And the two microscopes are genuinely different instruments, so they take different preprocessing paths; averaging them into one would mean deconvolving data the microscope had already deconvolved.
02

What it looks like

Raw confocal microscopy volume of the Drosophila mushroom body calyx, rotating in 3D. Boutons appear as faint overlapping blobs against background signal.
Raw input
What the microscope gives you.
Hand-annotated ground-truth instance labels for the same volume, each bouton rendered as a distinctly coloured 3D object.
Ground truth
Hand-annotated in napari, one label per bouton.
Predicted instance segmentation from the fine-tuned MicroSAM model on a held-out volume, closely matching the ground-truth labels.
Prediction
Fine-tuned MicroSAM (vit_l_lm), held-out volume.
One held-out volume, rotating. Left is what the microscope produces. Middle is what a human said the answer was. Right is what the fine-tuned model said, having never seen this volume. Each colour is one bouton instance.
03

Why this is hard

NO CANONICAL BOUTON
Two boutons in the same image can look nothing alike

Shape and intensity both vary, and they vary within a single volume rather than only between specimens: there is no reference size to filter on and no reference brightness to threshold at. That is exactly the assumption classical detection rests on — a threshold, a blob filter, a size prior all expect the thing you are looking for to look roughly the same wherever it turns up — which is why those methods fall over here. And the failure cuts both ways. Noise and background signal throw up structures that pass the same tests a real bouton passes, so a pipeline tuned far enough down to catch the faint ones starts reporting boutons that were never there. Beating that trade is what a learned model with a strong prior is actually being asked to do.

3D, AND ANISOTROPIC
A confocal Z-stack is not a cube of equal voxels

Z resolution is coarser than XY. A model that reasons in voxels rather than micrometres will systematically misjudge volume along one axis, so evaluation has to happen in physical units. Instance matching here accounts for voxel volume rather than counting voxels as if they were isotropic.

NO DATASET EXISTED
The annotations had to be made before the models could be trained

There is no public benchmark for this structure at this acquisition setting. The ground truth was hand-annotated from scratch in napari, which caps how much data there is and makes the choice of a foundation model with a strong prior a practical decision rather than a fashionable one.

THE END USER IS NOT AN ENGINEER
A checkpoint on a cluster helps nobody

The people who need these counts are biologists, not PyTorch users. If the deliverable had stopped at a weights file and a table of scores, the thesis would be complete and the work would be worthless. That is why BoutonViewer is part of this project rather than a follow-up to it.

04

The benchmark, and why each model is in it

Four architectures, one dataset, one evaluation protocol. The dataset is 13 volumetric stacks hand-annotated from scratch in napari across seven Drosophila brain preparations and two acquisition systems (confocal LSM and Airyscan), split nine train / two validation / two test. The point of including nnU-Net v2 and Cellpose 3D is not to have them lose. It is that a fine-tuned foundation model beating a self-configuring U-Net and the community default is a claim worth making, and beating nothing is not.

MicroSAM (vit_l_lm)Foundation model, fine-tunedselected

A microscopy-specialised SAM variant. Strong prior, so it can be fine-tuned on a small hand-annotated set. Trained here in 2D with Kornia augmentations and custom post-processing to assemble instances.

Cellpose 3DGeneralist cell segmentation

The default in the microscopy community, and the honest baseline to beat. If a general-purpose tool already solves this, the rest of the thesis is unnecessary.

nnU-Net v2Self-configuring U-Net

The standard against which medical and biological segmentation is measured, precisely because it removes architecture tuning as an excuse. Included so that a win is a real win.

SwinUNETRTransformer encoder, 3D (MONAI)

Tests whether native 3D attention beats a 2D foundation model with a strong prior, on a dataset this size. A genuinely open question at small n.

Each trained across four preprocessing variants

originalRaw

Unprocessed acquisition. The control.

dogDifference-of-Gaussians

Classical blob enhancement. Tests whether a hand-designed prior still buys anything once a foundation model is doing the work.

psfPSF-deconvolved

Richardson-Lucy against the measured point-spread function. Physically motivated, and expensive.

allCombined

All variants pooled into one training set.

Four architectures times four variants is a grid, not a single run, which is why training is launched through HTCondor submit files on the LFB cluster rather than by hand.

05

Results

Recall is the headline metric, not accuracy or Dice, and that is deliberate: a missed bouton is a permanent counting error, while a false positive can be filtered downstream. Evaluation prioritises recall, matched-instance mIoU, and a recall-weighted panoptic quality score. Test set: matched two-image test set · 33 ground-truth boutons.

Bouton instance-segmentation benchmark: recall, matched mIoU, recall-weighted panoptic quality and false-positive count across models.
ModelRecallMatched mIoURWPQFalse pos.
MicroSAM Largevit_l_lm · All-PSF1.0000.7870.78716
MicroSAM Basevit_b_lm · All-PSF1.0000.7710.7716
CellposeSAM2D XYZ · All-DoG0.9470.7680.7263
Swin UNETRAll-PSF0.9470.7320.6931
nnU-NetAll-PSF0.8950.6780.6092
Imarissemi-manual · current lab tool0.7350.5050.3715

The fine-tuned MicroSAM Large checkpoint recovered every bouton — perfect recall — and posted the highest matched-instance mIoU of any model evaluated, ahead of three SOTA 3D backbones and far ahead of the lab's current Imaris workflow. The Base variant matches that recall and gives up only 0.016 mIoU while cutting false positives from 16 to 6, which is why the card recommends it when GPU memory is tight.

The honest limitation, stated on the card: MicroSAM Large trades that perfect recall for a higher false-positive count (16 on the deconvolved condition), several of them large boutons split into multiple predicted components rather than spurious detections in empty regions. The card is explicit that predictions should be reviewed before unsupervised quantification — which is exactly why BoutonViewer surfaces oversized and merged predictions instead of silently filtering them. The model was fine-tuned on a comparatively small annotated set drawn from two brain specimens, so performance under substantially different imaging or labelling has not been verified.

RWPQ = recall-weighted panoptic quality (PQ_rec). Figures are the deconvolved All-PSF condition; the full evaluation set spans 61 ground-truth instances across the two held-out specimens.

Full model card

Evaluation is instance matching, not per-pixel overlap, and it weights by physical voxel volume so that anisotropic Z spacing does not quietly inflate or deflate a score. The code is in tools/evaluate_segmentation.py.

06

Two decisions worth defending

DECISION 1 · A 2D FOUNDATION MODEL FOR A 3D PROBLEM
Prior beats dimensionality when the dataset is hand-made

The obvious move on a 3D problem is a native 3D architecture, which is why SwinUNETR and nnU-Net v2 are in the benchmark. But every training example here was annotated by hand, which puts a hard ceiling on n. A model carrying a strong microscopy prior can be fine-tuned into that regime; a 3D transformer trained from a much weaker starting point has to learn more from less. Running both was the only honest way to find out which effect dominates, rather than asserting it.

DECISION 2 · TWO ACQUISITION PIPELINES, NOT ONE AVERAGE
LSM and Airyscan are different instruments and get different preprocessing

BoutonViewer runs rolling-ball background subtraction plus Richardson-Lucy deconvolution for confocal LSM stacks, and lightweight normalisation for Airyscan, which is already deconvolved by the microscope. Collapsing both into one path would have been less code and quietly wrong: you would be deconvolving already-deconvolved data. Voxel calibration is auto-derived from acquisition type and image size (LSM uses the confocal pitch 0.3 × 0.0709 × 0.0709 µm; a native super-resolution Airyscan image gets the finer 0.3 × 0.0425 × 0.0425 µm), and the stats are live-recomputed if a user overrides those fields, because the failure mode of getting this wrong is a plausible number in the wrong units rather than a crash. The Base variant on LSM also skips deconvolution entirely and takes the lighter path, so a GPU-limited run stays fast without a separate code branch.

07

Shipping it: BoutonViewer

A napari desktop application that runs the pipeline on a confocal or Airyscan TIFF stack, shows the raw channels and predicted labels in 3D, and reports per-bouton volume and surface area in µm³ and µm². A biologist loads a file and gets a table. It is in active use in the lab today, on their own acquisitions — which is the part of this project that a benchmark table cannot show.

  • Interactive stats table: hover or click a bouton in the viewer to inspect it, or delete a false positive without touching the model
  • Prediction caching, so changing a display setting does not re-run inference
  • Oversized and merged predictions are deliberately not auto-removed. A silent filter would hide exactly the failure mode that matters, so the tool surfaces them and lets a human decide
  • Model and data notes shipped alongside the tool: what it was trained on, at which voxel sizes, and where it should not be trusted
Model & data notes
08

Reproducibility

The thesis claim is reproducibility, so the repo has to earn it. Training, inference and evaluation are separate scripts with a --dataset flag selecting the preprocessing variant, cluster jobs are committed as submit files rather than remembered, and the selected checkpoints are published on Hugging Face with the rejected experiment runs left out. Two environment definitions exist on purpose: the annotation machine runs napari tooling under mamba, the training machine runs pip, because micro_sam is not reliably installable from PyPI and pretending otherwise would break the first person who tried to reproduce this.