The original post text is unavailable.
Samsung just dropped a method that squeezes a 13B LLM under 1 GB. LittleBit takes Llama2-13B down to about 0.9 GB at 0.1 bits per weight by factorizing the weights, binarizing the factors, and swapping heavy float math for XOR and sign flips. So it was a Large Language Model. Now it is basically just a Language Model that fits on a phone. The name is almost too memeable. LittleBit.