flashjolt-73f0f1ba·1 events·first seen Aliases: FlashJoLT
Researchers introduce JoLT, a KV cache compression method that treats the cache as a third-order tensor and applies a partial Tucker decomposition on the token and feature axes, then recovers truncation error with a Johnson-Lindenstrauss rotated low-bit residual. A Lagrangian dual jointly allocates Tucker ranks and residual bit-widths per layer group under a single byte budget. The method achieves 2-3x near-lossless compression on Mistral-7B-v0.3 and LLaMA-2-13B, with Frobenius reconstruction error roughly an order of magnitude below cross-layer SVD and 4-bit quantization. A randomized-SVD variant, FlashJoLT, delivers 5-13x compression-time speedup at matched quality.