LOCAL AIregistry
‹ all GPUs

RTX 4080

NVIDIA · 16 GB · 717 GB/s
Top 3 recipes
Recommended
Qwen3.8-27B
EXL3 3 bpwTabbyAPI128K context

Best for Coding agents, long documents, everyday assistant work.

70 tok/sdecode
612 tok/sprefill
128Kcontext
✓ Thinks✓ Tools✓ Long context✓ Vision
Tested on this card · Sep 25, 2026
Alternative
Qwen3.5-9B
EXL3 6 bpwTabbyAPI128K context

Best for Fast answers and light agent work on 8–12 GB cards.

125 tok/sdecode
2,382 tok/sprefill
128Kcontext
✓ Thinks✓ Tools✓ Long context
Tested on this card · Sep 26, 2026
Fastest
Gemma 4 26B A4B
EXL3 3.1 bpwTabbyAPI32K context

Best for Fast multimodal work on 16–24 GB cards.

136 tok/sdecode
2,573 tok/sprefill
32Kcontext
✓ Thinks✓ Tools✓ Long context✓ Vision
Tested on this card · Sep 26, 2026