Discussion about this post

User's avatar
IMSA's avatar

Would be nice if you tested split KV caching (i.e. quant just V-cache, keep K at full precision) since key is much more sensitive to quantization than value. Thanks so much for doing this work!

No posts

Ready for more?