SNaP: One-Step Posterior Sampling
for Noisy Inverse Problems

Geometry-aware one-step transport from measurement to posterior

Shirin Shoushtari1*, Edward P. Chandler2*, Xiao Shi2, Ulugbek S. Kamilov2
1WashU 2UW-Madison *Equal contribution
Preprint
Left: LPIPS versus wall-clock time per sample on CelebA deblurring. Right: SNaP draws and their averages for 4x super-resolution on AFHQ-Cat.

Left: LPIPS against wall-clock time per sample for CelebA Gaussian deblurring—SNaP reaches the lowest LPIPS at a fraction of the cost of iterative samplers. Right: $4\times$ super-resolution on AFHQ-Cat: one SNaP draw, then averages of $M = 4, 16, 100$ draws, and the ground truth. Every draw costs one network evaluation; averaging trades fine detail for lower pixel error. Click the figure to enlarge.

1 NFE
per posterior sample—one source solve and one network evaluation
>30×
faster per sample than iterative posterior samplers on CelebA
12 ms
per 128×128 sample

Abstract

Diffusion and flow-matching models can produce high-quality posterior samples for inverse problems, but typically require tens to thousands of network evaluations per draw. MeanFlow enables one-step generation, yet applying it to inverse problems leaves no intermediate steps at which to enforce measurement consistency. We introduce SNaP, a one-step MeanFlow posterior sampler for linear inverse problems with Gaussian noise. Its central innovation is a measurement-adapted source: a Gaussian distribution whose mean and anisotropic covariance are determined by the measurement operator, observation, and noise level. The source anchors well-measured directions while preserving variation where the measurements are weak or uninformative. We show that the exact conditional flow transports this source to the true posterior. Across natural-image restoration and multi-coil MRI, SNaP produces diverse, high-quality samples with one network evaluation per draw, 30 to 2250× faster than iterative samplers.

Method

Why one step is hard

Iterative posterior samplers pull toward the measurement $y = Ax_1 + n$ at every step of the trajectory. MeanFlow collapses the trajectory into a single jump $x_1 = x_0 + u(x_0, 0, 1)$, so there is no intermediate state left to correct. The obvious fix—condition the MeanFlow on $y$ and keep the isotropic $\mathcal{N}(0, I)$ source—fails in a specific way: the draws become nearly identical, the sampler degenerates into a point estimator, and averaging them gains nothing (calibration ratio 0.00). SNaP changes where the transport starts instead.

Schematic: likelihood tubes around the least-squares set; an isotropic source versus the SNaP source
Geometry-aware one-step transport. (a) Likelihood superlevel sets are tubes around the least-squares set $\mathcal{S}_y$: narrow in well-measured directions (width $\propto \sigma_n/s_i$), wider in weakly measured ones, unbounded along $\mathrm{null}(A)$. (b) An isotropic source spreads samples in every direction, producing long source-to-posterior displacements. (c) The SNaP source is squeezed where the measurement is informative and left wide where it is not, so displacements are short.
ONE POSTERIOR SAMPLE  =  1 SOLVE + 1 NETWORK EVALUATION measurement y = A x₁ + n operator & noise level A,  σₙ,  λ = σₙ² / τ² auxiliary randomness ε ~ N(0, τ²I) n′ ~ N(0, σₙ²I) perturb-and-solve source x₀ = ε + Aᵀ(AAᵀ + λI)⁻¹ · (y − Aε − n′) no matrix square root—one solve: diagonal, 2 FFTs, or a short CG (MRI) measurement-adapted source x₀ | y ~ N(mᵧ, τ²W) variance along vᵢ: τ²wᵢ, with wᵢ = σₙ² / (σₙ² + τ²sᵢ²) tight if measured · τ² on null(A) one MeanFlow step x̂₁ = x₀ + uθ(x₀, 0, 1 | y, σₙ) conditional average velocity over the whole interval [0, 1] no guidance · no projection posterior sample x̂₁ ≈ p(x₁ | y) redraw (ε, n′) with y fixed → M independent draws for M NFEs average them for distortion (PSNR / SSIM)  ·  read their spread as uncertainty

The source is a design variable

Along the straight path from a source sample $x_0$ to a posterior sample $x_1$, the conditional velocity is $x_1 - x_0$. SNaP keeps the path and the endpoint coupling fixed and designs the source, choosing a family with a linear, measurement-dependent center $x_0 = Ky + \xi$, where $\xi$ has covariance $\Sigma$. The expected displacement splits into a center term and a spread term:

$$ \mathbb{E}\big[\lVert x_1 - x_0 \rVert_2^2\big] \;=\; \operatorname{tr} R(K) \;+\; \operatorname{tr}\Sigma, \qquad R(K) := \mathbb{E}\big[(x_1 - Ky)(x_1 - Ky)^\top\big]. $$

Proposition 1 shows the center term is uniquely minimized by the linear MMSE estimator $K_C = CA^\top(ACA^\top + \sigma_n^2 I)^{-1}$, for any image distribution with covariance $C$. With the isotropic working covariance $C = \tau^2 I$ this is the Tikhonov estimate, and matching the spread to its error covariance gives the source

$$ q_\tau(x \mid y) = \mathcal{N}\big(x;\, m_y,\, \tau^2 W\big), \qquad \lambda = \sigma_n^2/\tau^2, $$ $$ m_y = A^\top(AA^\top + \lambda I)^{-1} y, \qquad W = I - A^\top(AA^\top + \lambda I)^{-1} A. $$

In the right singular basis of $A$, the source variance along $v_i$ is $\tau^2 w_i$ with $w_i = \sigma_n^2 / (\sigma_n^2 + \tau^2 s_i^2)$: small in well-measured directions, approaching $\sigma_n^2/s_i^2$ when the signal dominates, and equal to $\tau^2$ on $\mathrm{null}(A)$. The Tikhonov center also avoids amplifying noise in weakly measured directions, unlike the least-squares anchor $A^\dagger y$. The working model is used only to build the source; it does not assume the images are Gaussian. As $\sigma_n \to 0$ the source tends to $\mathcal{N}(A^\dagger y, \tau^2 P)$, which recovers the noiseless NullFlow source.

The target is still the true posterior

Proposition 2. Let $x_0 \sim q_\tau(\cdot \mid y)$ and $x_1 \sim p(\cdot \mid y)$ be conditionally independent and joined by straight lines. If the posterior has finite second moments and the marginal flow has a terminal limit, the exact one-step map

$$ T_y(x) = x + u(x, 0, 1 \mid y) \quad\text{satisfies}\quad (T_y)_\# \, q_\tau(\cdot \mid y) = p(\cdot \mid y) \quad \text{for every } \tau > 0. $$

The source changes the transport, not its destination; the trained network $u_\theta$ approximates $u$, so SNaP samples inherit the guarantee up to that approximation error. A companion result (Lemma 2) shows that the irreducible velocity-regression error—the part no network can remove—is smaller under $q_\tau$ than under any isotropic source $\mathcal{N}(0, \sigma^2 I)$ with $\sigma \ge \tau$.

Training and sampling

Sampling $q_\tau$ directly would need $W^{1/2}$. Perturb-and-solve avoids it: with independent $\epsilon \sim \mathcal{N}(0, \tau^2 I)$ and $n' \sim \mathcal{N}(0, \sigma_n^2 I)$,

$$ x_0 = \epsilon + A^\top (AA^\top + \lambda I)^{-1} \big(y - A\epsilon - n'\big) \;\sim\; \mathcal{N}(m_y, \tau^2 W) \text{ exactly.} $$

The solve is elementwise for inpainting and super-resolution, two FFTs for deblurring, and 10–20 conjugate-gradient iterations for multi-coil MRI. The network is trained with the conditional MeanFlow identity (one Jacobian–vector product per step) and the adaptive-weight loss of MeanFlow. At test time each sample is one source draw and one network evaluation; redrawing $(\epsilon, n')$ with $y$ fixed produces conditionally independent samples.

Algorithms

SampleSource—perturb-and-solve require: y, A, σₙ, τ
λ ← σₙ² / τ²
ε ~ N(0, τ²I),   n′ ~ N(0, σₙ²I)
r ← y − Aε − n′
x₀ ← ε + Aᵀ(AAᵀ + λI)⁻¹ r
return x₀
 
x₀ | y ~ N(mᵧ, τ²W) exactly (Lemma 1);
no square root of W is formed
Training step require: A, σₙ, τ, pratio = 0.5
x₁ ~ p,   y ← A x₁ + n
t ~ logit-normal
r ← t w.p. pratio, else r ~ U(0, t)
x₀ ← SampleSource(y, A, σₙ, τ)
zᵣ ← (1−r) x₀ + r x₁,   v ← x₁ − x₀
u, dᵣu ← JVP(uθ; (zᵣ, r, t), (v, 1, 0))
utgt ← v + (t − r) sg[dᵣu]
ℓ ← ‖u − utgt‖² / n
θ ← AdamW(θ, ∇θ sg[ω(ℓ)] ℓ)
Sampling—M posterior draws require: y, A, σₙ, trained uθ, M
for j = 1 … M do
  x₀⁽ʲ⁾ ← SampleSource(y, A, σₙ, τ)
  u ← uθ(x₀⁽ʲ⁾, 0, 1 | y, σₙ)
  x̂₁⁽ʲ⁾ ← x₀⁽ʲ⁾ + u
return {x̂₁⁽ʲ⁾}
M draws = M NFEs, one batched pass;
mean → distortion, spread → uncertainty

Qualitative Results

Comparison with DPS, DiffPIR and Flower on CelebA super-resolution and AFHQ-Cat box inpainting and deblurring
CelebA super-resolution, AFHQ-Cat box inpainting and deblurring. Values report PSNR/LPIPS for the displayed sample. SNaP uses one network evaluation, DiffPIR 100, and DPS and Flower 1000; SNaP attains the best LPIPS on all three.
Multi-coil fastMRI brain reconstructions at 8x acceleration and 20 dB input SNR
fastMRI brain, multi-coil, ×8 acceleration, 20 dB input SNR. Reconstructions with PSNR/SSIM, error magnitude and a zoomed region. SNaP is competitive with one NFE; the baselines use 1000.
Box inpainting reconstructions on CelebA
Box inpainting on CelebA (centered 40×40 mask, $\sigma_n = 0.05$).
Box inpainting reconstructions on AFHQ-Cat
Box inpainting on AFHQ-Cat (centered 80×80 mask, $\sigma_n = 0.05$).

Diverse Posterior Draws

Holding the measurement fixed and redrawing only the source noise gives independent samples. Each figure shows six draws for one CelebA box-inpainting measurement from DPS (1000 NFEs per draw), Flower1-OT (100) and SNaP (1). Variation should appear inside the masked box, where the measurement says nothing, and vanish where the pixels are observed.

Six box-inpainting draws each from DPS, Flower1-OT and SNaP
Six independent draws per method for one box-inpainting measurement. Rows: DPS, Flower1-OT, SNaP.
Six box-inpainting draws each from DPS, Flower1-OT and SNaP
Six draws per method for a second measurement. Rows: DPS, Flower1-OT, SNaP.

Inference Cost

A SNaP sample costs one network evaluation and one source draw. The source solve with $(AA^\top + \lambda I)^{-1}$ is diagonal or Fourier-diagonal for the natural-image operators and a short conjugate-gradient run for multi-coil MRI, so it adds little to the network call. That makes SNaP 30 to 2250× faster per sample than the iterative samplers on CelebA deblurring and 335 to 563× faster at ×8 MRI acceleration.

CelebA, Gaussian deblurring
MetricDPSOT-ODEFlowerDDRMSNaP
NFE ↓1000180100201
Time / sample (s) ↓27.0146.5752.6140.3600.012
Slowdown vs. SNaP2251×548×218×30×–
fastMRI brain, multi-coil ×8
MetricCSGMDAPSPnP-DMDPSSNaP
NFE ↓10001000100010001
Time / sample (s) ↓127.7896.2892.0276.070.227
Slowdown vs. SNaP563×424×405×335×–

NFE and single-batch wall-clock time per posterior sample. Best in each row.

Quantitative Results

100 test measurements per setting. Natural-image degradations follow PnP-Flow and Flower with $\sigma_n = 0.05$ (0.01 for random inpainting): a $61\times61$ Gaussian blur, $\times2$ (CelebA) or $\times4$ (AFHQ-Cat) strided super-resolution, 70% random inpainting, and a centered $40\times40$ or $80\times80$ box. MRI uses $A = MFS$ with ESPIRiT sensitivities at $\times4$ and $\times8$ Cartesian acceleration. For SNaP we report a single draw ($M = 1$) and the average of $M$ draws, which costs $M$ NFEs.

A single SNaP draw is competitive in LPIPS at one NFE; averaging draws attains the best SSIM on every CelebA task and the best PSNR on three of four. SNaP also reaches the best FID on two of three CelebA operators and the calibration ratio closest to 1 on all three.

DeblurringSuper-resolutionRandom inpaintingBox inpainting
MethodNFEPSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓
Degraded–27.260.8380.21411.670.1820.85911.980.1981.07022.330.7540.217
MSE regressor135.740.9530.02934.170.9450.03334.420.9550.022–––
PnP-GS2333.970.9240.04131.230.8900.06529.170.8740.066–––
DDRM2035.020.9460.02732.240.9270.02732.400.9420.029–––
DiffPIR10034.830.9380.02731.870.8920.03032.450.9240.02430.480.9180.028
OT-ODE18032.960.9200.02931.340.9030.02728.690.8710.04929.370.9190.038
PnP-Flow550034.800.9410.04631.440.9050.05633.980.9530.02131.060.9390.043
Flower5-OT50035.650.9540.03133.030.9310.03933.970.9520.01931.850.9520.023
DPS100034.270.9360.01932.230.9170.02333.430.9510.01330.260.9420.022
SNaP (M=1)132.900.9190.01831.270.9030.02331.970.9270.01830.730.9340.021
SNaP (M=4)434.980.9470.02233.260.9340.02333.970.9510.01532.780.9540.018
SNaP (M=16)1635.680.9540.02833.950.9430.02934.640.9570.01833.430.9590.020
SNaP (M=100)10035.910.9560.03034.160.9460.03134.840.9590.01933.690.9610.022

100 CelebA test images. Best and second best among all methods (the degraded reference is excluded). – marks a setting where a baseline does not apply. Scroll the table sideways on narrow screens.

DeblurringSuper-resolutionRandom inpaintingBox inpainting
MethodNFEPSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓
Degraded–24.390.5320.54311.980.2120.90013.250.2141.09121.570.7360.214
MSE regressor129.820.8060.28929.160.8170.20133.230.9150.066–––
PnP-GS2328.390.7870.38724.440.6390.41129.890.8410.123–––
PnP-Flow5500–250029.010.7850.31228.010.7910.16734.470.9330.04427.140.9000.127
DDRM2029.010.7840.19427.090.7810.18632.200.9070.064–––
DiffPIR10028.890.7730.18723.160.6350.29131.760.8810.05327.710.8800.059
OT-ODE18027.820.7350.12626.710.7370.10829.980.8490.08624.470.8730.095
Flower5-OT500–250029.730.8010.26427.090.7670.27234.410.9320.04727.320.9220.067
DPS100025.900.6610.14825.980.6940.14432.060.8960.03626.780.9000.054
SNaP (M=1)126.250.6740.15926.070.7030.13030.480.8530.06626.460.8910.054
SNaP (M=4)428.330.7500.15728.070.7770.12032.380.8960.05228.490.9140.047
SNaP (M=16)1629.090.7790.23128.770.8020.15733.040.9090.05629.320.9220.056
SNaP (M=100)10029.320.7880.32028.990.8100.18733.250.9130.05829.530.9250.069

100 AFHQ-Cat test images. PnP-Flow and Flower step counts depend on the operator, hence the NFE range. The MSE regressor is omitted on box inpainting, where its output is the blurred conditional mean of the masked region. Best and second best.

PSNR, SSIM and LPIPS do not say whether the spread of the draws is the right size. The calibration ratio $\Delta_{\text{cal}} = V/B$ divides the variance $V$ of the draws by the squared error $B$ of their mean; for an exact posterior sampler the two are equal. Values below 1 mean draws collapsed toward a point estimate, above 1 over-dispersed. It is a second-moment check, not a coverage certificate.

Gaussian deblurringSuper-resolutionBox inpainting
MethodNFEFID↓Δcal(→ 1)FID↓Δcal(→ 1)FID↓Δcal(→ 1)
MSE regressor122.06–24.76–––
DDRM2020.260.3118.940.19––
Flower1-OT10017.860.2323.790.2418.720.24
DPS100018.250.6419.940.6715.422.15
SNaP (M=1)115.801.0316.800.9915.531.08

CelebA. Every SNaP entry is a single draw; FID on 2048 reconstructions, $\Delta_{\text{cal}}$ with $M = 100$. Best and second best ($\Delta_{\text{cal}}$ ranked by distance to 1).

×4 · 20 dB×4 · 30 dB×8 · 20 dB×8 · 30 dB
MethodNFEPSNR↑SSIM↑PSNR↑SSIM↑PSNR↑SSIM↑PSNR↑SSIM↑
Zero-filled–25.350.79125.370.79321.980.69421.990.694
Wavelet+ℓ1–26.570.67227.750.76823.020.55123.810.643
TV–26.190.79826.270.79922.620.68922.660.691
MSE regressor133.800.90933.850.91230.710.87530.730.878
DiffPIR10029.840.70930.010.71926.140.64226.310.655
CSGM100027.740.73432.390.80227.740.73328.330.754
DPS100028.810.64629.830.66924.890.57525.170.574
DAPS100031.550.81832.440.78227.310.74928.010.714
PnP-DM100032.160.77833.390.78627.180.70129.000.712
SNaP (M=1)132.060.88932.890.89928.100.81829.140.845
SNaP (M=4)432.680.90433.800.91929.060.84329.880.864
SNaP (M=16)1632.850.90834.050.92429.340.85030.090.869

100 multi-coil fastMRI brain (AXT2) test slices; PSNR/SSIM on magnitude images. Best and second best among all methods. Among the posterior samplers, SNaP is competitive with a single NFE.

Ablations and Robustness

Same network, conditioning and schedule; only the source differs. With either isotropic source the calibration ratio falls to 0.00: the draws are near-identical, the sampler has become a point estimator with a random input, and averaging 100 draws changes nothing. The measurement-adapted source keeps the draws dispersed. Its single-draw advantage is perceptual (LPIPS 0.018 vs. 0.030), and its average overtakes the isotropic sources in PSNR.

SourceM=1 · PSNR ↑LPIPS ↓M=100 · PSNR ↑LPIPS ↓Δcal (→ 1)
N(0, I)35.850.03035.860.0300.00
N(0, τ²I)35.830.03035.840.0300.00
N(my, τ²W)—ours32.900.01835.910.0301.03

CelebA Gaussian deblurring, $\sigma_n = 0.05$.

Six draws and per-pixel standard deviation under the three sources
Six draws per source for one measurement, with the per-pixel standard deviation on the same ×20 scale. Under the SNaP source the draws differ in the hair, skin texture and the region behind the sunglasses, and the spread concentrates on edges the blur suppressed. Under either isotropic source the draws are indistinguishable and the spread map is black—raising $\tau$ from 0.15 to 1.0 does not restore it, so it is the shape of the source, not its size, that sets the spread.
PSNR and LPIPS against the number of averaged draws for the three sources
PSNR (left) and LPIPS (right) against the number of averaged draws $M$. Only the SNaP source responds to $M$; the isotropic curves are flat because their draws carry almost no variance to average away.

$\sigma_n$ enters the source in closed form, so a wrong value changes how the source allocates uncertainty rather than breaking the method. A model trained at $\sigma_n = 0.05$ and tested on cleaner measurements (over-estimating the noise by up to 5×) degrades gracefully: PSNR still rises as the measurements get cleaner and LPIPS stays within 0.017–0.020. Under-estimation is not tested.

Test noise σnPSNR ↑SSIM ↑LPIPS ↓Input PSNRGain
0.0134.290.9400.02029.96+4.33
0.0234.150.9390.01929.62+4.53
0.0333.890.9360.01829.12+4.77
0.0433.500.9300.01728.52+4.98
0.05 † (training noise)32.900.9190.01827.81+5.09

CelebA Gaussian deblurring; source and model fixed at $\sigma_n = 0.05$.

Proposition 2 says $\tau$ shapes the transport but not its target, so it is a design parameter rather than a modelling choice. Distortion is indistinguishable between $\tau = 0.15$ and $0.5$ at every $M$ and drops about 0.4 dB at $\tau = 1.0$; the relative spread of the draws (5.49%, 5.48%, 5.72%) barely moves. $\tau$ does not need precise tuning.

τ = 0.15τ = 0.5τ = 1.0
DrawsPSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓
M = 132.940.9200.01932.920.9190.01932.510.9060.024
M = 434.970.9460.02134.950.9460.01834.590.9390.020
M = 1635.690.9540.02835.660.9530.02335.320.9480.025
M = 10035.910.9560.03035.880.9550.02635.550.9500.027

CelebA Gaussian deblurring, 100 test measurements; $M$ = number of averaged draws.

Composing $u_{r,t}$ over a uniform partition of $[0,1]$ gives $k$-step sampling without retraining. More steps are not monotonically better: every metric peaks at $k = 2$ and declines afterwards, and $k = 8$ is worse than one step for $M \ge 16$. The model is trained to jump source-to-data, so evaluating it at intermediate $(r,t)$ compounds its error. The gain at $k = 2$ is small, so SNaP uses $k = 1$.

k = 1k = 2k = 4k = 8
DrawsPSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓
M = 132.940.9200.01933.080.9210.01733.040.9200.01732.990.9200.017
M = 434.970.9460.02135.040.9470.02135.000.9470.02134.940.9460.021
M = 1635.690.9540.02835.730.9540.02835.690.9540.02835.620.9530.028
M = 10035.910.9560.03035.940.9560.03035.900.9560.03035.830.9550.031

CelebA Gaussian deblurring, $\tau = 0.15$.

Acknowledgments

This work was supported in part by the National Science Foundation under CAREER award CCF-2625643 and award CCF-2622128.

BibTeX

@article{shoushtari2026snap,
  title   = {SNaP: One-Step Posterior Sampling for Noisy Inverse Problems},
  author  = {Shoushtari, Shirin and Chandler, Edward P. and Shi, Xiao
             and Kamilov, Ulugbek S.},
  journal = {arXiv preprint arXiv:2609.34071},
  year    = {2026}
}