Imaging & vision
Restoration: degradation, denoise, super-resolution
A restoration pipeline is a model of how the image got damaged, then an inversion of it.
Restoration is the part of imaging that pays for itself in products. But it is very easy to do it by piling up models. The disciplined way is to write down the degradation first, then invert it.
y = D( H( x ) ) + n
# x clean image
# H blur / resampling / compression already baked into the file
# D downsample or quantise (super-resolution: estimate from this lossy observation)
# n noise (denoise: this is what you undo)
# y what you actually have
# Every restoration task is: given y, estimate x
# — and the art is in modelling H, D, n and the prior or extra measurements.That one line separates the three tasks that are constantly confused:
Denoising — `H` and `D` are identity; you are estimating `x` despite `n`.
Deblurring — `n` is small; you are undoing `H` (and it is ill-posed: many `x` produce the same `y`).
Single-image super-resolution — downsampling `D` removes or mixes detail. The inverse is not unique, so an estimate needs a prior (often learned); extra captures can also provide measurements. One low-resolution image alone cannot uniquely determine arbitrary missing detail.
Which brings up the metric problem, and it is the thing most teams get wrong.
· PSNR / MSE measure pixelwise distortion. When several sharp details are plausible from the same degraded input, minimising average pixel error can favor a smooth compromise that scores well but looks soft. This is a perception–distortion tradeoff, not a rule that every high-PSNR result must be blurry.
· SSIM compares local structure and can complement pixel error, but it is not a universal measure of human preference.
· Perceptual metrics such as LPIPS compare learned feature representations. They capture different aspects of similarity; no single metric is a complete quality oracle.
What does super-resolution actually rely on?
Why is PSNR a poor target for restoration?
Degradation model
y = D(H(x)) + n — write it down before choosing a method.
Ill-posed
Many inputs produce the same output, so the inverse is not unique.
PSNR / SSIM / LPIPS
Pixel error / local structure / learned features; each measures a different aspect.
Review cards
Degradation model
y = D(H(x)) + n — write it down before choosing a method.
Ill-posed
Many inputs produce the same output, so the inverse is not unique.
PSNR / SSIM / LPIPS
Pixel error / local structure / learned features; each measures a different aspect.
Sources for this lesson
Below are the references, editions and original links for further reading and checking.
BookComputer Vision: Algorithms and Applicationsfree
Richard Szeliski
2nd edition, Springer 2022(初稿 2020 起公开征求勘误)
一本书覆盖图像处理到三维重建。第 2 版把深度学习独立成章(第 5 章)。作者官网提供免费 PDF,2026 年秋季多所高校课程仍在用它。
Rafael C. Gonzalez, Richard E. Woods
4th edition, Pearson(40 周年纪念版)
传统图像处理的圣经。第 4 版扩写了深度学习、CNN、SIFT、MSER、图割、超像素、活动轮廓等,并重组了图像变换一章。
CourseSuper-resolution and Image Priorsfree
Computational Photography, MIT CSAIL
Online textbook, section 8.2
Frames single-image super-resolution as an ill-posed inverse problem and explains priors, multi-frame measurements, and limits on recovering detail.
PaperThe Perception-Distortion Tradeofffree
Yochai Blau, Tomer Michaeli
CVPR 2018, pp. 6228–6237
形式化图像恢复中的失真与感知质量权衡;说明单一 PSNR/SSIM 分数不等于完整的视觉质量判断。
Where this is heading
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
ECCV 2020 (oral), arXiv:2003.08934
神经渲染的开端。注意不要与 NeRF++(arXiv:2010.07492)混淆,后者是续作。
3D Gaussian Splatting for Real-Time Radiance Field Rendering
SIGGRAPH 2023, ACM TOG 42(4), arXiv:2308.04079
实时辐射场渲染的当前主流方向:用 3D 高斯代替 MLP,1080p 下可达 30fps 以上。
Lights up these nodes in the hub:i-03 · i-04 · i-05