Abstract:
To address bandwidth constraints and severe multipath interference in underwater acoustic channels, where existing continuous-feature semantic transmission schemes suffer from catastrophic distortion, a robust underwater image transmission method based on discrete semantic tokens is porposed in this paper. The proposed method utilizes a Vision Transformer (ViT) to extract global 1-D latent features, which are then mapped into high-dimensional discrete token sequences via a shared semantic codebook and vector quantization mechanism. This discretization paradigm effectively strips spatial redundancy and background noise while enhancing signal robustness through the error-tolerance of discrete spaces. Combined with codebook-based error correction at the receiver, high-fidelity semantic reconstruction is achieved under extreme channel conditions. Experiments on the URPC2020 dataset demonstrate that under severe multipath distortion, the proposed method significantly outperforms state-of-the-art semantic communication schemes in perceptual quality metrics such as FID and LPIPS.