A single-image GAN framework with adversarially learned inference
Neurocomputing, cilt.700, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 700
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.neucom.2026.134464
- Dergi Adı: Neurocomputing
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Applied Science & Technology Source, Compendex, EMBASE, INSPEC, zbMATH, Academic Search Ultimate (EBSCO), Engineering Source (EBSCO)
- Anahtar Kelimeler: Adversarially learned inference, Bidirectional adversarial learning, Image manipulation, Image-latent inference, Internal learning, Multi-scale generative modeling, Single-image GANs
- Erzincan Binali Yıldırım Üniversitesi Adresli: Evet
Özet
Single-image generative adversarial networks aim to learn a generative prior from a single natural image by exploiting internal patch statistics, spatial self-similarity, and cross-scale regularities. Although substantial progress has been achieved in single-image synthesis and manipulation, existing single-image GANs remain vulnerable to training instability, memorization, limited sample diversity, and the absence of an explicit mechanism for inferring latent representations from observed images. We present a multi-scale single-image generative framework that integrates Adversarially Learned Inference (ALI) into a coarse-to-fine adversarial architecture. The proposed model jointly learns image generation and latent inference from a single training image by aligning generated image-latent pairs with real image-inferred latent pairs at each scale of the pyramid. Each scale consists of a conditioning encoder, generator, inversion encoder, and discriminator, enabling feature-conditioned residual generation and latent inference within a unified adversarial formulation. A hybrid progressive-concurrent training strategy is employed to improve inter-scale consistency while preserving global structural organization and fine-scale visual detail. Experiments on Places, LSUN, and ImageNet, together with quantitative evaluation, visual comparisons, ablation analysis, and human perceptual studies, indicate that the proposed framework provides a favorable trade-off among image fidelity, sample diversity, model compactness, and training efficiency under the reported evaluation protocol. The inferred latent representations are further used to support image manipulation tasks, including editing, harmonization, paint-to-image translation, and animation. Overall, the results suggest that incorporating ALI into multi-scale single-image generative modeling is a promising direction for supporting latent-image consistency, controllability, and efficient synthesis in extremely data-constrained settings.