LOCAL AIregistry
‹ all GPUs

RTX 4090

NVIDIA · 24 GB · 1008 GB/s
Top 3 recipes
Recommended
Qwen3.8-27B
EXL3 3 bpwSGLang200K context

Best for Coding agents, long documents, everyday assistant work.

174 tok/sdecode
1,318 tok/sprefill
200Kcontext
✓ Thinks✓ Tools✓ Long context✓ Vision
Tested on this card · Sep 25, 2026
Longest context
Qwen3.5-9B
EXL3 8 bpwTabbyAPI256K context

Best for Fast answers and light agent work on 8–12 GB cards.

135 tok/sdecode
1,614 tok/sprefill
256Kcontext
✓ Thinks✓ Tools✓ Long context
Tested on this card · Sep 26, 2026
Another family
Gemma 4 26B A4B
EXL3 5.1 bpwTabbyAPI64K context

Best for Fast multimodal work on 16–24 GB cards.

141 tok/sdecode
4,173 tok/sprefill
64Kcontext
✓ Thinks✓ Tools✓ Long context✓ Vision
Tested on this card · Sep 26, 2026