Python: match the MATLAB MSSIM formula in Measure_Quality - #783
Merged
Merged
Conversation
Member
|
Same as before: remove test, we can merge |
The MSSIM branch used K1*d**2 and K2*d**2 in the denominators of the luminance and contrast terms, and (K2*d**2)/2 in the structure term, while the numerators of the same expressions already use (K1*d)**2 and (K2*d)**2. MSSIM.m squares the whole product in every one of those six places. K2 was also 0.02 here against 0.03 in MSSIM.m, and the standard deviations came from np.std, which normalises by N, while delta on the next line and MATLAB's std both normalise by N-1. Squaring only K*d leaves the stabilisation term 1/K1 = 100 times too large in l; in c and s, with the K2 difference on top, 0.02*d**2 is about 22 times the (0.03*d)**2 of MSSIM.m. So l, c and s never reach 1: on master, MSSIM(x, x) * N is 0.79 for uniform noise and 0.57 for a sparse phantom instead of 1, and the ceiling depends on the image, so values are not comparable across datasets. Fixes CERN#782
Dev-next-gen
force-pushed
the
fix/mssim-matlab-parity
branch
from
September 29, 2026 11:55
02e7b7f to
918122b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #782
I was comparing the Python and MATLAB quality measures and noticed that MSSIM does not give the same number on the two sides. Looking at the
MSSIMbranch ofPython/tigre/utilities/Measure_Quality.pyagainstMATLAB/Utilities/Quality_measures/MSSIM.m, four things differ.The denominators of the luminance, contrast and structure terms use
K1 * d ** 2,K2 * d ** 2and(K2 * d ** 2) / 2, whileMSSIM.msquares the whole product in all three:(K1*d)^2,(K2*d)^2and((K2*d)^2)/2. The numerators of those same three expressions already use(K1 * d) ** 2and(K2 * d) ** 2, so the file was inconsistent with itself. On top of thatK2was0.02against0.03inMSSIM.m(0.03 is also the value used in the original SSIM paper, Wang et al. 2004), and the standard deviations came fromnp.std(), which normalises byN, whiledelta, computed a few lines below, and MATLAB'sstdboth normalise byN - 1.Squaring only
K*dleaves the stabilisation term1/K1 = 100times too large inl; incands, with theK2difference on top, it is0.02 * d ** 2whereMSSIM.mhas(0.03 * d) ** 2, about 22 times too large. That is enough to keep l, c and s away from 1 even when the two images are identical, and the size of the gap depends on the image, so the numbers are not comparable from one dataset to the next. On master,Measure_Quality(x, x, ["MSSIM"]) * x.size— which should be 1 — gives 0.785 for uniform noise in [0, 1], 0.788 for the same noise scaled to [0, 255], 0.574 for a sparse phantom and 0.799 for a low contrast volume.The change is confined to the
MSSIMbranch of that one function;RMSE,nRMSE,CC,UQIandSSDare untouched, and no signature changes. It does move the MSSIM numbers that users get, but in the direction of the MATLAB reference this file is a port of: for the same two inputs, the Python function now returns whatMSSIM.mreturns. That is parity at the level of the function only; the iterative algorithms do not all passres_prevandresin the same order on both sides, anddis taken from the first argument.I checked the change with three tests: the self-similarity identity above, which needs no reference implementation to be convincing, and two comparisons against a literal transcription of
MSSIM.m, one with a wide dynamic range and one withd = 1. They were inPython/tests/test_measure_quality_mssim.pyin the first version of this PR and were removed on request, so the PR now only touchesMeasure_Quality.py. All three fail on master and pass with this change. The other Python tests behave the same with and without the change; the only other failure I see in my environment istest_gpuids_default.py::test_awmin_tv_accepts_none, which failed in 1 of 4 full-suite runs with this change and in 1 of 4 on master, and passes on its own. Three files do not collect there either, for reasons unrelated to this change:algorithm_test.pyandtest_config_gen.pybecausetigre.demosis not in the installed wheel, andstdout_test.pybecause it imports the Python 2StringIOmodule.I have not touched
UQI, but it has the samenp.var()against MATLABvar()normalisation difference, so it is probably worth a separate look.AI Usage
Found by a defect-hunting pipeline I build and run (Dev-next-gen), using Claude Code with Anthropic's Claude Opus 5.
MSSIM.m, wrote the patch toPython/tigre/utilities/Measure_Quality.pyand the test used to check it (not part of the PR), and drafted this description.Nothing this pipeline produces is submitted without human approval. The entire pull request — the code change, the tests and this description — was reviewed by a human before the pipeline was allowed to submit it.
(@Dev-next-gen)