Qwen3.5-9B
EXL3 6 bpw · tabbyapi · 128K context · RTX 5080 16 GB
Alibaba's 9B model from February 2026: small enough for an 8 GB card, with a 64K–128K window. It thinks, calls tools and stays quick on modest hardware.
173 tok/sdecode
2,781 tok/sprefill
128Kcontext window
Tested on this cardSep 26, 2026
With Omarchy Local AI it is one button. By hand, it is three steps: the weights, the config, the container. The server then answers on http://localhost:8000/v1.
- 1 Download the weights
hf download TheMelonGod/Qwen3.5-9B-exl3 \ --revision 7fbc2095db50b0d56f7f07ac84903f98f883d3b9 \ --local-dir ~/models/Qwen3.5-9B-exl3-7fbc2095
- 2 Write the server config
cat > qwen3.5-9b.tabbyapi.128k.yml <<'EOF' network: {host: 0.0.0.0, port: 5000, disable_auth: true, disable_fetch_requests: true, send_tracebacks: false, api_servers: [OAI], sse_ping_interval: 15} logging: {log_prompt: false, log_generation_params: false, log_requests: false, log_chat_completion_requests: false} model: model_dir: /workspace/models model_name: Qwen3.5-9B-exl3-7fbc2095 backend: exllamav3 inline_model_loading: false max_seq_len: 131072 cache_size: 264192 cache_mode: Q4 tensor_parallel: false gpu_split_auto: true autosplit_reserve: [256] chunk_size: 2048 output_chunking: true max_batch_size: 2 vision: false reasoning: true start_in_reasoning: auto tool_format: qwen3_coder template_vars_default: {} draft_model: {draft_mode: mtp} memory: {sysmem_recurrent_cache: 4096, sysmem_kv_cache: 0, cuda_malloc_async: true} EOF - 3 Start the server
docker run --rm \ --gpus all \ -p 8000:5000 \ --shm-size 8g \ -v ~/models/Qwen3.5-9B-exl3-7fbc2095:/workspace/models/Qwen3.5-9B-exl3-7fbc2095:ro \ -v $PWD/qwen3.5-9b.tabbyapi.128k.yml:/app/config.yml:ro \ --entrypoint /opt/venv/bin/python3 \ ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1 \ main.py \ --config /app/config.yml
✓ loadLoads and serves within an hour
✓ chatAnswers a plain question and stops on its own
✓ reasoningThinks separately and gets 17 × 23 right
✓ toolsCalls a tool with the right arguments and uses the result
✓ contextRecalls a code buried in a prompt that fills 85% of the window
✓ speedDecodes at 15 tok/s or more
Passed all six checks on a real RTX 5080 on Sep 26, 2026.
- Weights
- TheMelonGod/Qwen3.5-9B-exl3 @ 7fbc2095db ›
- Image
- ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1
- Engine profile
- tabbyapi-exl3 ›
- Recipe file
- registry/recipes/nvidia/rtx-5080-16gb/qwen3.5-9b.tabbyapi.128k.json ›
- Model released
- Feb 26, 2026
- Tested
- vast, RTX 5080