Models
Only models released in the last eight months, from Qwen, Gemma, DeepSeek, GLM, Step, Kimi and MiniMax, with thinking on.
Prism ML's ternary Bonsai 2, built on Qwen3.6-27B: a 27B model packed small enough for 16 GB, with vision, tools and thinking.
Google's Gemma 4 12B from May 2026: a compact multimodal model with thinking, good with images and documents.
Google's Gemma 4 mixture-of-experts model: 26B total, 4B active, so it runs fast while seeing images and thinking.
Alibaba's 9B model from February 2026: small enough for an 8 GB card, with a 64K–128K window. It thinks, calls tools and stays quick on modest hardware.
Alibaba's 27B dense model from August 2026, and the strongest model that fits one consumer card. It thinks before it answers, calls tools reliably and reads images, which makes it the default for coding agents and long projects.
Qwen3.8's large mixture-of-experts model: far more knowledge than the 27B, with few active parameters, so it stays fast.