NCAP: Scene Text Image Super-Resolution with Non-CAtegorical Prior

Dongwoo Park; Suk Pil Ko

2025 WACV WACV 2025

NCAP: Scene Text Image Super-Resolution with Non-CAtegorical Prior

Abstract

Scene text image super-resolution (STISR) enhances the resolution and quality of low-resolution images. Unlike previous studies that treated scene text images as natural images recent methods using a text prior (TP) extracted from a pre-trained text recognizer have shown strong performance. However two major issues emerge: (1) Explicit categorical priors like TP can negatively impact STISR if incorrect. We reveal that these explicit priors are unstable and propose replacing them with Non-CAtegorical Prior (NCAP) using penultimate layer representations. (2) Pre-trained recognizers used to generate TP struggle with low-resolution images. To address this most studies jointly train the recognizer with the STISR network to bridge the domain gap between low- and high-resolution images but this can cause an overconfidence phenomenon in the prior modality. We highlight this issue and propose a method to mitigate it by mixing hard and soft labels. Experiments on the TextZoom dataset demonstrate an improvement by 3.5% while our method significantly enhances generalization performance by 14.8% across four text recognition datasets. Our method generalizes to all TP-guided STISR networks.

🌉 Interdisciplinary Bridge — Computer Vision and Machine Learning

🧭 Keyword Pioneer — text prior

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Dongwoo Park , Suk Pil Ko

Topics

Machine Learning > Core Methods > Representation Learning Machine Learning > Learning Types > Self-Supervised Learning Computer Vision > Processing > Image Restoration

Keywords

domain generalization feature representation image super-resolution prior knowledge scene text recognition text recognition scene text text prior

Download PDF

Related papers

Neural Graph Map: Dense Mapping with Efficient Loop Closure Integration 2025

ELMGS: Enhancing Memory and Computation Scalability through Compression for 3D Gaussian Splatting 2025

Feature Fusion Transferability Aware Transformer for Unsupervised Domain Adaptation 2025

Uncertainty-Aware Online Extrinsic Calibration: A Conformal Prediction Approach 2025

Disentangling Spatio-Temporal Knowledge for Weakly Supervised Object Detection and Segmentation in Surgical Video 2025