LOCAL AIregistry

Models

Only models released in the last eight months, from Qwen, Gemma, DeepSeek, GLM, Step, Kimi and MiniMax, with thinking on.

Bonsai 2 27B
released Apr 21, 2026visionthinks

Prism ML's ternary Bonsai 2, built on Qwen3.6-27B: a 27B model packed small enough for 16 GB, with vision, tools and thinking.

Runs on 1 GPU: RTX 2000 Ada
Gemma 4 12B
released May 23, 2026visionthinks

Google's Gemma 4 12B from May 2026: a compact multimodal model with thinking, good with images and documents.

Gemma 4 26B A4B
released Mar 12, 2026visionthinks

Google's Gemma 4 mixture-of-experts model: 26B total, 4B active, so it runs fast while seeing images and thinking.

Qwen3.5-9B
released Feb 26, 2026visionthinks

Alibaba's 9B model from February 2026: small enough for an 8 GB card, with a 64K–128K window. It thinks, calls tools and stays quick on modest hardware.

Qwen3.8-27B
released Aug 5, 2026visionthinks

Alibaba's 27B dense model from August 2026, and the strongest model that fits one consumer card. It thinks before it answers, calls tools reliably and reads images, which makes it the default for coding agents and long projects.

Qwen3.8-Flash-Next
released Aug 24, 2026visionthinks

Qwen3.8's large mixture-of-experts model: far more knowledge than the 27B, with few active parameters, so it stays fast.

Runs on 1 GPU: RTX PRO 6000 Blackwell