PACodec: A Low-bitrate Neural Speech Codec with Parallel Additive Vector Quantization

Abstract

This paper proposes PACodec, a novel low-bitrate neural speech codec based on parallel additive vector quantization (PAVQ). Unlike the mainstream residual vector quantization (RVQ) used in most neural speech codecs, where vector quantizers (VQs) are sequentially dependent, the PAVQ strategy adopted in PACodec aggregates parallel quantization results to optimize bitrate usage. Specifically, the PAVQ adopts a ''global-local-global'' (GLG) design: the global encoded features are quantized in parallel by multiple independent VQs, each attending to a local component of the representation, and their outputs are aggregated through addition to yield the final global quantization result for decoding. Experimental results show that PACodec, as each VQ focuses only on local information, supports smaller codebooks and reduces bitrate by 30% compared with baselines at the same decoding quality, with only minor model complexity. Further analysis shows that, owing to the GLG framework of PAVQ, the proposed PACodec is disentanglement-friendly, and each independent VQ captures different aspects of speech, e.g., content, timbre, and acoustic details, suggesting potential for application to downstream tasks such as voice conversion.


I. LibriTTS, 16kHz


Sample 1


Raw Speech PACodec@1.4kbps Encodec@1.5kbps Encodec@2kbps AudioDec@1.5kbps AudioDec@2kbps DAC@1.5kbps
DAC@2kbps APCodec@1.5kbps APCodec@2kbps MDCTCodec@1.5kbps MDCTCodec@2kbps HiFi-Codec@2kbps SQCodec@1.5kbps

Sample 2


Raw Speech PACodec@1.4kbps Encodec@1.5kbps Encodec@2kbps AudioDec@1.5kbps AudioDec@2kbps DAC@1.5kbps
DAC@2kbps APCodec@1.5kbps APCodec@2kbps MDCTCodec@1.5kbps MDCTCodec@2kbps HiFi-Codec@2kbps SQCodec@1.5kbps

Sample 3


Raw Speech PACodec@1.4kbps Encodec@1.5kbps Encodec@2kbps AudioDec@1.5kbps AudioDec@2kbps DAC@1.5kbps
DAC@2kbps APCodec@1.5kbps APCodec@2kbps MDCTCodec@1.5kbps MDCTCodec@2kbps HiFi-Codec@2kbps SQCodec@1.5kbps

II. VCTK, 48kHz


Sample 1


Raw Speech PACodec@4.2kbps Encodec@4.5kbps Encodec@6kbps AudioDec@4.5kbps AudioDec@6kbps DAC@4.5kbps
DAC@6kbps APCodec@4.5kbps APCodec@6kbps MDCTCodec@4.5kbps MDCTCodec@6kbps HiFi-Codec@6kbps

Sample 2


Raw Speech PACodec@4.2kbps Encodec@4.5kbps Encodec@6kbps AudioDec@4.5kbps AudioDec@6kbps DAC@4.5kbps
DAC@6kbps APCodec@4.5kbps APCodec@6kbps MDCTCodec@4.5kbps MDCTCodec@6kbps HiFi-Codec@6kbps

Sample 3


Raw Speech PACodec@4.2kbps Encodec@4.5kbps Encodec@6kbps AudioDec@4.5kbps AudioDec@6kbps DAC@4.5kbps
DAC@6kbps APCodec@4.5kbps APCodec@6kbps MDCTCodec@4.5kbps MDCTCodec@6kbps HiFi-Codec@6kbps

III. Disentanglement Potential Analysis (VCTK, 48kHz)


Sample 1


Raw Speech PACodec PACodec-VQ1 PACodec-VQ2 PACodec-VQ3 PACodec-VQ4

Sample 2


Raw Speech PACodec PACodec-VQ1 PACodec-VQ2 PACodec-VQ3 PACodec-VQ4

Sample 3


Raw Speech PACodec PACodec-VQ1 PACodec-VQ2 PACodec-VQ3 PACodec-VQ4