TRUE
MARGIN.
Medical-imaging research on a narrow question: when an image-registration method reports uncertainty, does that uncertainty actually identify where the registration is wrong?
Calibration is not the same as knowing where a registration is wrong.
An uncertainty estimator can look reasonable when results are averaged across a dataset and still be unhelpful at the individual locations where an error matters. TrueMargin is testing that distinction directly.
The real prostate-imaging dataset does not provide independently verified pointwise T2-to-DCE correspondences. Real-data error measurements are therefore treated as a reference proxy, not as ground truth. Synthetic known-deformation experiments are reserved for settings where true registration error can actually be known.
The first task was to test whether the original uncertainty mechanism was doing meaningful work.
An audit of the historical ensemble found that it added Gaussian noise with a fixed raw-intensity standard deviation of 0.02 to DICOM images whose intensity scales were often in the hundreds or thousands. That made the perturbation likely too small to probe meaningful registration sensitivity.
A scale-aware replacement was defined before result-bearing runs: noise magnitude was tied to each image's own intensity standard deviation. Five patients and six frozen settings were then evaluated across 150 registrations.
The corrected intensity-perturbation ensemble did not pass the predefined gate.
None of the tested nonzero perturbation levels satisfied the predeclared combination of repeatability, non-inertness, and error-degradation criteria. Small perturbations were often difficult to distinguish from baseline repeatability, while larger perturbations increasingly produced registration failures or large displacement-field divergence.
The failed gate is part of the result. The experiment was not extended with additional perturbation values after the fact to search for a more favorable outcome.
Before testing another uncertainty mechanism, the registration itself has to be stable enough.
Gate A produced enough failures and divergence to raise a more basic question: are the uncertainty experiments measuring sensitivity to meaningful perturbations, or partly measuring optimizer instability?
The current experiment varies only the registration iteration budget while keeping the rest of the mesh-3 registration regime fixed. Budgets of 15, 30, 60, and 100 iterations are tested with five repeated registrations for each of five patients. The 100-iteration regime is the reference, giving a frozen 100-registration study.
Proxy error is recorded descriptively, but it is not allowed to choose the iteration budget.
Designed around prospective decisions and reproducible evidence.
Patient-level T2/DCE registration experiments with explicit data provenance.
Repeated registrations and displacement-field comparisons under frozen settings.
Aggregate behavior is kept separate from pointwise informativeness.
Ground-truth language is reserved for experiments where the deformation is known.
The next uncertainty experiment depends on the convergence result.
If a stable registration regime is established, the next uncertainty mechanism will be specified before its results are observed. Candidate directions include controlled initialization or registration-parameter perturbations, followed by sensitivity checks, synthetic known-deformation experiments, final cohort analysis, and manuscript reconciliation.
TrueMargin remains private while the methodology is still being tested. The project currently makes no clinical-performance, diagnostic, or deployment claims.