Traditional image compression methods aim to reconstruct images for human perception, prioritizing visual fidelity over task relevance. In contrast, Coding for Machines focuses on preserving information essential for automated understanding. Building on this principle, we present an end-to-end compression framework that retains text-specific features for Optical Character Recognition (OCR). The encoder operates at roughly half the computational cost of the OCR module, making it suitable for resource-limited devices. When on-device OCR is infeasible, images can be efficiently compressed and later decoded to recover textual content. Experiments show significant improvements in text extraction accuracy at low bitrates, even outperforming OCR on uncompressed images. Building on these insights, we further explore general-purpose encoders under extreme compression, investigating whether compact, visually degraded representations can still retain recoverable semantic information. Results demonstrate that semantic content can persist despite severe compression, bridging task-oriented text compression and broader machine-centered image coding.

End-to-end semantic preservation in text-aware image compression systems

Della Fiore, Stefano;Gnutti, Alessandro;Dalai, Marco;Migliorati, Pierangelo;Leonardi, Riccardo
2026-01-01

Abstract

Traditional image compression methods aim to reconstruct images for human perception, prioritizing visual fidelity over task relevance. In contrast, Coding for Machines focuses on preserving information essential for automated understanding. Building on this principle, we present an end-to-end compression framework that retains text-specific features for Optical Character Recognition (OCR). The encoder operates at roughly half the computational cost of the OCR module, making it suitable for resource-limited devices. When on-device OCR is infeasible, images can be efficiently compressed and later decoded to recover textual content. Experiments show significant improvements in text extraction accuracy at low bitrates, even outperforming OCR on uncompressed images. Building on these insights, we further explore general-purpose encoders under extreme compression, investigating whether compact, visually degraded representations can still retain recoverable semantic information. Results demonstrate that semantic content can persist despite severe compression, bridging task-oriented text compression and broader machine-centered image coding.
File in questo prodotto:
File Dimensione Formato  
End-to-end semantic preservation in text-aware image compression systems.pdf

accesso aperto

Tipologia: Full Text
Licenza: Copyright dell'editore
Dimensione 4.3 MB
Formato Adobe PDF
4.3 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11379/653705
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact