ALIREZA MOMENZADEH

PhD Graduate

PhD program:: XXXVIII


advisor: prof. Enzo Baccarelli

Thesis title: Enhancing Single Image Super-Resolution via Deep Learning: Stability, Perceptual Quality, and Multi-Scale Modeling

The goal of Single Image Super-Resolution (SISR) is to reconstruct a high-resolution (HR) image from a single low-resolution (LR) observation. This task is ill-posed: the formation of LR image removes high-frequency information, so multiple distinct HR images can map to the same LR input. In medical imaging, this ambiguity becomes severe because hallucinated structures or reconstruction artifacts can mislead diagnosis. Generative models such as Generative adversarial Networks (GANs) and Diffusion Models (DMs) can produce visually sharp super-resolved images, however, they typically introduce non-deterministic textures, artifacts, or hallucinations. In addition, GAN training is very unstable and sensitive to hyperparameters, and DMs require many iterative denoising steps that makes the inference computationally expensive. Motivated by these constraints, this thesis investigates how to bridge the perceptual-quality gap between regression-based Super-Resolution (SR) and generative methods, while preserving the stability, determinism, and artifact-awareness that are needed in medical applications. The main idea is that regression-based models can be improved by (i) architectural designs that enhance representation learning, (ii) a constructed set of perceptual and structural losses that target texture, edges, and frequency content, and (iii) stabilization of convolutional layers that are inspired by weight scaling techniques. The proposed contributions are threefold: First, we develop Twinned Residual Auto-Encoder architecture (TRAE) for denoising SR, including a multi-resolution extension that produces consistent reconstructions over multiple upscaling factors: Multi-Resolution Twinned Residual Auto-Encoder architecture (MR-TRAE). Second, we introduce a perceptual SR model based on Convolutional Neural Network (CNN) backbones that is trained without GANs' adversarial losses but equipped with a set of losses. We combine a robust Charbonnier content loss with feature-based perceptual loss, gradient/edge preservation, frequency-domain alignment using masked Fourier magnitudes, and Gram-matrix style/texture matching. Third, to mimic realistic clinical degradation, we use a stochastic LR synthesis model that includes probabilistic blur (Gaussian kernels with varying variance and kernel size), diverse resampling operators, and additive Gaussian noise, to reduce sensitivity to idealized bicubic assumptions and improving robustness. Our experiments on medical Computed Tomography (CT) imagery (including COVIDx CT-2A) show that the proposed regression-based models improve perceptual fidelity and structural consistency while maintaining stable training. Quantitative comparisons with representative baselines such as EDSR and SRGAN models confirm competitive reconstruction quality, with strong structural similarity behavior that is aligned with the avoidance of artifact required by medical imaging.

Research products

Connessione ad iris non disponibile

© Università degli Studi di Roma "La Sapienza" - Piazzale Aldo Moro 5, 00185 Roma