September 22, 2026
Qwen3.5-4B handles long prompts faster: graft8 delivered 3.7x
On September 22, the graft8 variant processed a 128-thousand-token prompt 3.7 times faster than the original Qwen3.5-4B, but with a loss in accuracy. The author released two graft versions of the model.

On September 22, the graft8 variant processed a 128-thousand-token prompt 3.7 times faster than the original Qwen3.5-4B, but with a loss in accuracy. The author released two graft versions of the model.
The second version, graft16, doubled prompt-processing speed with minimal loss of accuracy.
