LOCAL AIregistry
‹ all GPUs

RTX 5090

NVIDIA · 32 GB · 1792 GB/s
Top 3 recipes
Recommended
Qwen3.8-27B
EXL3 5 bpwTabbyAPI256K context

Best for Coding agents, long documents, everyday assistant work.

155 tok/sdecode
1,407 tok/sprefill
256Kcontext
✓ Thinks✓ Tools✓ Long context✓ Vision
Tested on this card · Sep 25, 2026
Fastest
Qwen3.5-9B
EXL3 8 bpwTabbyAPI256K context

Best for Fast answers and light agent work on 8–12 GB cards.

206 tok/sdecode
3,164 tok/sprefill
256Kcontext
✓ Thinks✓ Tools✓ Long context
Tested on this card · Sep 26, 2026
Another family
Gemma 4 26B A4B
EXL3 6.1 bpwTabbyAPI128K context

Best for Fast multimodal work on 16–24 GB cards.

173 tok/sdecode
4,483 tok/sprefill
128Kcontext
✓ Thinks✓ Tools✓ Long context✓ Vision
Tested on this card · Sep 26, 2026