Diffusion and flow-matching models can produce high-quality posterior samples for inverse problems, but typically require tens to thousands of network evaluations per draw. MeanFlow enables one-step generation, yet applying it to inverse problems leaves no intermediate steps at which to enforce measurement consistency. We introduce SNaP, a one-step MeanFlow posterior sampler for linear inverse problems with Gaussian noise. Its central innovation is a measurement-adapted source: a Gaussian distribution whose mean and anisotropic covariance are determined by the measurement operator, observation, and noise level. The source anchors well-measured directions while preserving variation where the measurements are weak or uninformative. We show that the exact conditional flow transports this source to the true posterior. Across natural-image restoration and multi-coil MRI, SNaP produces diverse, high-quality samples with one network evaluation per draw, 30 to 2250× faster than iterative samplers.
Iterative posterior samplers pull toward the measurement $y = Ax_1 + n$ at every step of the trajectory. MeanFlow collapses the trajectory into a single jump $x_1 = x_0 + u(x_0, 0, 1)$, so there is no intermediate state left to correct. The obvious fix—condition the MeanFlow on $y$ and keep the isotropic $\mathcal{N}(0, I)$ source—fails in a specific way: the draws become nearly identical, the sampler degenerates into a point estimator, and averaging them gains nothing (calibration ratio 0.00). SNaP changes where the transport starts instead.
Along the straight path from a source sample $x_0$ to a posterior sample $x_1$, the conditional velocity is $x_1 - x_0$. SNaP keeps the path and the endpoint coupling fixed and designs the source, choosing a family with a linear, measurement-dependent center $x_0 = Ky + \xi$, where $\xi$ has covariance $\Sigma$. The expected displacement splits into a center term and a spread term:
Proposition 1 shows the center term is uniquely minimized by the linear MMSE estimator $K_C = CA^\top(ACA^\top + \sigma_n^2 I)^{-1}$, for any image distribution with covariance $C$. With the isotropic working covariance $C = \tau^2 I$ this is the Tikhonov estimate, and matching the spread to its error covariance gives the source
In the right singular basis of $A$, the source variance along $v_i$ is $\tau^2 w_i$ with $w_i = \sigma_n^2 / (\sigma_n^2 + \tau^2 s_i^2)$: small in well-measured directions, approaching $\sigma_n^2/s_i^2$ when the signal dominates, and equal to $\tau^2$ on $\mathrm{null}(A)$. The Tikhonov center also avoids amplifying noise in weakly measured directions, unlike the least-squares anchor $A^\dagger y$. The working model is used only to build the source; it does not assume the images are Gaussian. As $\sigma_n \to 0$ the source tends to $\mathcal{N}(A^\dagger y, \tau^2 P)$, which recovers the noiseless NullFlow source.
Proposition 2. Let $x_0 \sim q_\tau(\cdot \mid y)$ and $x_1 \sim p(\cdot \mid y)$ be conditionally independent and joined by straight lines. If the posterior has finite second moments and the marginal flow has a terminal limit, the exact one-step map
The source changes the transport, not its destination; the trained network $u_\theta$ approximates $u$, so SNaP samples inherit the guarantee up to that approximation error. A companion result (Lemma 2) shows that the irreducible velocity-regression error—the part no network can remove—is smaller under $q_\tau$ than under any isotropic source $\mathcal{N}(0, \sigma^2 I)$ with $\sigma \ge \tau$.
Sampling $q_\tau$ directly would need $W^{1/2}$. Perturb-and-solve avoids it: with independent $\epsilon \sim \mathcal{N}(0, \tau^2 I)$ and $n' \sim \mathcal{N}(0, \sigma_n^2 I)$,
The solve is elementwise for inpainting and super-resolution, two FFTs for deblurring, and 10–20 conjugate-gradient iterations for multi-coil MRI. The network is trained with the conditional MeanFlow identity (one Jacobian–vector product per step) and the adaptive-weight loss of MeanFlow. At test time each sample is one source draw and one network evaluation; redrawing $(\epsilon, n')$ with $y$ fixed produces conditionally independent samples.
require: y, A, σₙ, τ
λ ← σₙ² / τ²
ε ~ N(0, τ²I), n′ ~ N(0, σₙ²I)
r ← y − Aε − n′
x₀ ← ε + Aᵀ(AAᵀ + λI)⁻¹ r
return x₀
x₀ | y ~ N(mᵧ, τ²W) exactly (Lemma 1);
no square root of W is formed
require: A, σₙ, τ, pratio = 0.5
x₁ ~ p, y ← A x₁ + n
t ~ logit-normal
r ← t w.p. pratio, else r ~ U(0, t)
x₀ ← SampleSource(y, A, σₙ, τ)
zᵣ ← (1−r) x₀ + r x₁, v ← x₁ − x₀
u, dᵣu ← JVP(uθ; (zᵣ, r, t), (v, 1, 0))
utgt ← v + (t − r) sg[dᵣu]
ℓ ← ‖u − utgt‖² / n
θ ← AdamW(θ, ∇θ sg[ω(ℓ)] ℓ)
require: y, A, σₙ, trained uθ, M
for j = 1 … M do
x₀⁽ʲ⁾ ← SampleSource(y, A, σₙ, τ)
u ← uθ(x₀⁽ʲ⁾, 0, 1 | y, σₙ)
x̂₁⁽ʲ⁾ ← x₀⁽ʲ⁾ + u
return {x̂₁⁽ʲ⁾}
M draws = M NFEs, one batched pass;
mean → distortion, spread → uncertainty
Holding the measurement fixed and redrawing only the source noise gives independent samples. Each figure shows six draws for one CelebA box-inpainting measurement from DPS (1000 NFEs per draw), Flower1-OT (100) and SNaP (1). Variation should appear inside the masked box, where the measurement says nothing, and vanish where the pixels are observed.
A SNaP sample costs one network evaluation and one source draw. The source solve with $(AA^\top + \lambda I)^{-1}$ is diagonal or Fourier-diagonal for the natural-image operators and a short conjugate-gradient run for multi-coil MRI, so it adds little to the network call. That makes SNaP 30 to 2250× faster per sample than the iterative samplers on CelebA deblurring and 335 to 563× faster at ×8 MRI acceleration.
| CelebA, Gaussian deblurring | |||||
|---|---|---|---|---|---|
| Metric | DPS | OT-ODE | Flower | DDRM | SNaP |
| NFE ↓ | 1000 | 180 | 100 | 20 | 1 |
| Time / sample (s) ↓ | 27.014 | 6.575 | 2.614 | 0.360 | 0.012 |
| Slowdown vs. SNaP | 2251× | 548× | 218× | 30× | – |
| fastMRI brain, multi-coil ×8 | |||||
|---|---|---|---|---|---|
| Metric | CSGM | DAPS | PnP-DM | DPS | SNaP |
| NFE ↓ | 1000 | 1000 | 1000 | 1000 | 1 |
| Time / sample (s) ↓ | 127.78 | 96.28 | 92.02 | 76.07 | 0.227 |
| Slowdown vs. SNaP | 563× | 424× | 405× | 335× | – |
NFE and single-batch wall-clock time per posterior sample. Best in each row.
100 test measurements per setting. Natural-image degradations follow PnP-Flow and Flower with $\sigma_n = 0.05$ (0.01 for random inpainting): a $61\times61$ Gaussian blur, $\times2$ (CelebA) or $\times4$ (AFHQ-Cat) strided super-resolution, 70% random inpainting, and a centered $40\times40$ or $80\times80$ box. MRI uses $A = MFS$ with ESPIRiT sensitivities at $\times4$ and $\times8$ Cartesian acceleration. For SNaP we report a single draw ($M = 1$) and the average of $M$ draws, which costs $M$ NFEs.
A single SNaP draw is competitive in LPIPS at one NFE; averaging draws attains the best SSIM on every CelebA task and the best PSNR on three of four. SNaP also reaches the best FID on two of three CelebA operators and the calibration ratio closest to 1 on all three.
| Deblurring | Super-resolution | Random inpainting | Box inpainting | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | NFE | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ |
| Degraded | – | 27.26 | 0.838 | 0.214 | 11.67 | 0.182 | 0.859 | 11.98 | 0.198 | 1.070 | 22.33 | 0.754 | 0.217 |
| MSE regressor | 1 | 35.74 | 0.953 | 0.029 | 34.17 | 0.945 | 0.033 | 34.42 | 0.955 | 0.022 | – | – | – |
| PnP-GS | 23 | 33.97 | 0.924 | 0.041 | 31.23 | 0.890 | 0.065 | 29.17 | 0.874 | 0.066 | – | – | – |
| DDRM | 20 | 35.02 | 0.946 | 0.027 | 32.24 | 0.927 | 0.027 | 32.40 | 0.942 | 0.029 | – | – | – |
| DiffPIR | 100 | 34.83 | 0.938 | 0.027 | 31.87 | 0.892 | 0.030 | 32.45 | 0.924 | 0.024 | 30.48 | 0.918 | 0.028 |
| OT-ODE | 180 | 32.96 | 0.920 | 0.029 | 31.34 | 0.903 | 0.027 | 28.69 | 0.871 | 0.049 | 29.37 | 0.919 | 0.038 |
| PnP-Flow5 | 500 | 34.80 | 0.941 | 0.046 | 31.44 | 0.905 | 0.056 | 33.98 | 0.953 | 0.021 | 31.06 | 0.939 | 0.043 |
| Flower5-OT | 500 | 35.65 | 0.954 | 0.031 | 33.03 | 0.931 | 0.039 | 33.97 | 0.952 | 0.019 | 31.85 | 0.952 | 0.023 |
| DPS | 1000 | 34.27 | 0.936 | 0.019 | 32.23 | 0.917 | 0.023 | 33.43 | 0.951 | 0.013 | 30.26 | 0.942 | 0.022 |
| SNaP (M=1) | 1 | 32.90 | 0.919 | 0.018 | 31.27 | 0.903 | 0.023 | 31.97 | 0.927 | 0.018 | 30.73 | 0.934 | 0.021 |
| SNaP (M=4) | 4 | 34.98 | 0.947 | 0.022 | 33.26 | 0.934 | 0.023 | 33.97 | 0.951 | 0.015 | 32.78 | 0.954 | 0.018 |
| SNaP (M=16) | 16 | 35.68 | 0.954 | 0.028 | 33.95 | 0.943 | 0.029 | 34.64 | 0.957 | 0.018 | 33.43 | 0.959 | 0.020 |
| SNaP (M=100) | 100 | 35.91 | 0.956 | 0.030 | 34.16 | 0.946 | 0.031 | 34.84 | 0.959 | 0.019 | 33.69 | 0.961 | 0.022 |
100 CelebA test images. Best and second best among all methods (the degraded reference is excluded). – marks a setting where a baseline does not apply. Scroll the table sideways on narrow screens.
| Deblurring | Super-resolution | Random inpainting | Box inpainting | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | NFE | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ |
| Degraded | – | 24.39 | 0.532 | 0.543 | 11.98 | 0.212 | 0.900 | 13.25 | 0.214 | 1.091 | 21.57 | 0.736 | 0.214 |
| MSE regressor | 1 | 29.82 | 0.806 | 0.289 | 29.16 | 0.817 | 0.201 | 33.23 | 0.915 | 0.066 | – | – | – |
| PnP-GS | 23 | 28.39 | 0.787 | 0.387 | 24.44 | 0.639 | 0.411 | 29.89 | 0.841 | 0.123 | – | – | – |
| PnP-Flow5 | 500–2500 | 29.01 | 0.785 | 0.312 | 28.01 | 0.791 | 0.167 | 34.47 | 0.933 | 0.044 | 27.14 | 0.900 | 0.127 |
| DDRM | 20 | 29.01 | 0.784 | 0.194 | 27.09 | 0.781 | 0.186 | 32.20 | 0.907 | 0.064 | – | – | – |
| DiffPIR | 100 | 28.89 | 0.773 | 0.187 | 23.16 | 0.635 | 0.291 | 31.76 | 0.881 | 0.053 | 27.71 | 0.880 | 0.059 |
| OT-ODE | 180 | 27.82 | 0.735 | 0.126 | 26.71 | 0.737 | 0.108 | 29.98 | 0.849 | 0.086 | 24.47 | 0.873 | 0.095 |
| Flower5-OT | 500–2500 | 29.73 | 0.801 | 0.264 | 27.09 | 0.767 | 0.272 | 34.41 | 0.932 | 0.047 | 27.32 | 0.922 | 0.067 |
| DPS | 1000 | 25.90 | 0.661 | 0.148 | 25.98 | 0.694 | 0.144 | 32.06 | 0.896 | 0.036 | 26.78 | 0.900 | 0.054 |
| SNaP (M=1) | 1 | 26.25 | 0.674 | 0.159 | 26.07 | 0.703 | 0.130 | 30.48 | 0.853 | 0.066 | 26.46 | 0.891 | 0.054 |
| SNaP (M=4) | 4 | 28.33 | 0.750 | 0.157 | 28.07 | 0.777 | 0.120 | 32.38 | 0.896 | 0.052 | 28.49 | 0.914 | 0.047 |
| SNaP (M=16) | 16 | 29.09 | 0.779 | 0.231 | 28.77 | 0.802 | 0.157 | 33.04 | 0.909 | 0.056 | 29.32 | 0.922 | 0.056 |
| SNaP (M=100) | 100 | 29.32 | 0.788 | 0.320 | 28.99 | 0.810 | 0.187 | 33.25 | 0.913 | 0.058 | 29.53 | 0.925 | 0.069 |
100 AFHQ-Cat test images. PnP-Flow and Flower step counts depend on the operator, hence the NFE range. The MSE regressor is omitted on box inpainting, where its output is the blurred conditional mean of the masked region. Best and second best.
PSNR, SSIM and LPIPS do not say whether the spread of the draws is the right size. The calibration ratio $\Delta_{\text{cal}} = V/B$ divides the variance $V$ of the draws by the squared error $B$ of their mean; for an exact posterior sampler the two are equal. Values below 1 mean draws collapsed toward a point estimate, above 1 over-dispersed. It is a second-moment check, not a coverage certificate.
| Gaussian deblurring | Super-resolution | Box inpainting | |||||
|---|---|---|---|---|---|---|---|
| Method | NFE | FID↓ | Δcal(→ 1) | FID↓ | Δcal(→ 1) | FID↓ | Δcal(→ 1) |
| MSE regressor | 1 | 22.06 | – | 24.76 | – | – | – |
| DDRM | 20 | 20.26 | 0.31 | 18.94 | 0.19 | – | – |
| Flower1-OT | 100 | 17.86 | 0.23 | 23.79 | 0.24 | 18.72 | 0.24 |
| DPS | 1000 | 18.25 | 0.64 | 19.94 | 0.67 | 15.42 | 2.15 |
| SNaP (M=1) | 1 | 15.80 | 1.03 | 16.80 | 0.99 | 15.53 | 1.08 |
CelebA. Every SNaP entry is a single draw; FID on 2048 reconstructions, $\Delta_{\text{cal}}$ with $M = 100$. Best and second best ($\Delta_{\text{cal}}$ ranked by distance to 1).
| ×4 · 20 dB | ×4 · 30 dB | ×8 · 20 dB | ×8 · 30 dB | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | NFE | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ |
| Zero-filled | – | 25.35 | 0.791 | 25.37 | 0.793 | 21.98 | 0.694 | 21.99 | 0.694 |
| Wavelet+ℓ1 | – | 26.57 | 0.672 | 27.75 | 0.768 | 23.02 | 0.551 | 23.81 | 0.643 |
| TV | – | 26.19 | 0.798 | 26.27 | 0.799 | 22.62 | 0.689 | 22.66 | 0.691 |
| MSE regressor | 1 | 33.80 | 0.909 | 33.85 | 0.912 | 30.71 | 0.875 | 30.73 | 0.878 |
| DiffPIR | 100 | 29.84 | 0.709 | 30.01 | 0.719 | 26.14 | 0.642 | 26.31 | 0.655 |
| CSGM | 1000 | 27.74 | 0.734 | 32.39 | 0.802 | 27.74 | 0.733 | 28.33 | 0.754 |
| DPS | 1000 | 28.81 | 0.646 | 29.83 | 0.669 | 24.89 | 0.575 | 25.17 | 0.574 |
| DAPS | 1000 | 31.55 | 0.818 | 32.44 | 0.782 | 27.31 | 0.749 | 28.01 | 0.714 |
| PnP-DM | 1000 | 32.16 | 0.778 | 33.39 | 0.786 | 27.18 | 0.701 | 29.00 | 0.712 |
| SNaP (M=1) | 1 | 32.06 | 0.889 | 32.89 | 0.899 | 28.10 | 0.818 | 29.14 | 0.845 |
| SNaP (M=4) | 4 | 32.68 | 0.904 | 33.80 | 0.919 | 29.06 | 0.843 | 29.88 | 0.864 |
| SNaP (M=16) | 16 | 32.85 | 0.908 | 34.05 | 0.924 | 29.34 | 0.850 | 30.09 | 0.869 |
100 multi-coil fastMRI brain (AXT2) test slices; PSNR/SSIM on magnitude images. Best and second best among all methods. Among the posterior samplers, SNaP is competitive with a single NFE.
Same network, conditioning and schedule; only the source differs. With either isotropic source the calibration ratio falls to 0.00: the draws are near-identical, the sampler has become a point estimator with a random input, and averaging 100 draws changes nothing. The measurement-adapted source keeps the draws dispersed. Its single-draw advantage is perceptual (LPIPS 0.018 vs. 0.030), and its average overtakes the isotropic sources in PSNR.
| Source | M=1 · PSNR ↑ | LPIPS ↓ | M=100 · PSNR ↑ | LPIPS ↓ | Δcal (→ 1) |
|---|---|---|---|---|---|
| N(0, I) | 35.85 | 0.030 | 35.86 | 0.030 | 0.00 |
| N(0, τ²I) | 35.83 | 0.030 | 35.84 | 0.030 | 0.00 |
| N(my, τ²W)—ours | 32.90 | 0.018 | 35.91 | 0.030 | 1.03 |
CelebA Gaussian deblurring, $\sigma_n = 0.05$.
$\sigma_n$ enters the source in closed form, so a wrong value changes how the source allocates uncertainty rather than breaking the method. A model trained at $\sigma_n = 0.05$ and tested on cleaner measurements (over-estimating the noise by up to 5×) degrades gracefully: PSNR still rises as the measurements get cleaner and LPIPS stays within 0.017–0.020. Under-estimation is not tested.
| Test noise σn | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Input PSNR | Gain |
|---|---|---|---|---|---|
| 0.01 | 34.29 | 0.940 | 0.020 | 29.96 | +4.33 |
| 0.02 | 34.15 | 0.939 | 0.019 | 29.62 | +4.53 |
| 0.03 | 33.89 | 0.936 | 0.018 | 29.12 | +4.77 |
| 0.04 | 33.50 | 0.930 | 0.017 | 28.52 | +4.98 |
| 0.05 † (training noise) | 32.90 | 0.919 | 0.018 | 27.81 | +5.09 |
CelebA Gaussian deblurring; source and model fixed at $\sigma_n = 0.05$.
Proposition 2 says $\tau$ shapes the transport but not its target, so it is a design parameter rather than a modelling choice. Distortion is indistinguishable between $\tau = 0.15$ and $0.5$ at every $M$ and drops about 0.4 dB at $\tau = 1.0$; the relative spread of the draws (5.49%, 5.48%, 5.72%) barely moves. $\tau$ does not need precise tuning.
| τ = 0.15 | τ = 0.5 | τ = 1.0 | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Draws | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ |
| M = 1 | 32.94 | 0.920 | 0.019 | 32.92 | 0.919 | 0.019 | 32.51 | 0.906 | 0.024 |
| M = 4 | 34.97 | 0.946 | 0.021 | 34.95 | 0.946 | 0.018 | 34.59 | 0.939 | 0.020 |
| M = 16 | 35.69 | 0.954 | 0.028 | 35.66 | 0.953 | 0.023 | 35.32 | 0.948 | 0.025 |
| M = 100 | 35.91 | 0.956 | 0.030 | 35.88 | 0.955 | 0.026 | 35.55 | 0.950 | 0.027 |
CelebA Gaussian deblurring, 100 test measurements; $M$ = number of averaged draws.
Composing $u_{r,t}$ over a uniform partition of $[0,1]$ gives $k$-step sampling without retraining. More steps are not monotonically better: every metric peaks at $k = 2$ and declines afterwards, and $k = 8$ is worse than one step for $M \ge 16$. The model is trained to jump source-to-data, so evaluating it at intermediate $(r,t)$ compounds its error. The gain at $k = 2$ is small, so SNaP uses $k = 1$.
| k = 1 | k = 2 | k = 4 | k = 8 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Draws | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ |
| M = 1 | 32.94 | 0.920 | 0.019 | 33.08 | 0.921 | 0.017 | 33.04 | 0.920 | 0.017 | 32.99 | 0.920 | 0.017 |
| M = 4 | 34.97 | 0.946 | 0.021 | 35.04 | 0.947 | 0.021 | 35.00 | 0.947 | 0.021 | 34.94 | 0.946 | 0.021 |
| M = 16 | 35.69 | 0.954 | 0.028 | 35.73 | 0.954 | 0.028 | 35.69 | 0.954 | 0.028 | 35.62 | 0.953 | 0.028 |
| M = 100 | 35.91 | 0.956 | 0.030 | 35.94 | 0.956 | 0.030 | 35.90 | 0.956 | 0.030 | 35.83 | 0.955 | 0.031 |
CelebA Gaussian deblurring, $\tau = 0.15$.
This work was supported in part by the National Science Foundation under CAREER award CCF-2625643 and award CCF-2622128.
@article{shoushtari2026snap,
title = {SNaP: One-Step Posterior Sampling for Noisy Inverse Problems},
author = {Shoushtari, Shirin and Chandler, Edward P. and Shi, Xiao
and Kamilov, Ulugbek S.},
journal = {arXiv preprint arXiv:2609.34071},
year = {2026}
}