logic65/Qwen3.8-Whittle-16B
Text Generation • 16B • Updated • 5.92k • 2
A 27B whittled on two 8GB GPUs: measured cuts, heals, and restorations. 48L = best untrained, Whittle-16B = best overall.
Note Best of the family at 32/39. Sparse MoE carved from the dense 27B: 5120 shared + 64 experts of 192, top-16, so 17.8B of 26.9B runs per token. English only, 0/3 multilingual.