Qwen3.5-9B
EXL3 4 bpw · tabbyapi · 128K context · RTX 4070 Ti 12 GB
Alibaba's 9B model from February 2026: small enough for an 8 GB card, with a 64K–128K window. It thinks, calls tools and stays quick on modest hardware.
78 tok/sdecode
1,632 tok/sprefill
128Kcontext window
Tested on this cardSep 25, 2026
With Omarchy Local AI it is one button. By hand, it is three steps: the weights, the config, the container. The server then answers on http://localhost:8000/v1.
- 1 Download the weights
hf download TheMelonGod/Qwen3.5-9B-exl3 \ --revision e98453aa143a24fab9eb14b718daed02e2fa6eef \ --local-dir ~/models/Qwen3.5-9B-exl3-e98453aa
- 2 Write the server config
cat > qwen3.5-9b.tabbyapi.128k.yml <<'EOF' network: host: 0.0.0.0 port: 5000 disable_auth: true disable_fetch_requests: true send_tracebacks: false api_servers: ["OAI"] sse_ping_interval: 15 logging: log_prompt: false log_generation_params: false log_requests: false log_chat_completion_requests: false model: model_dir: /workspace/models inline_model_loading: false use_dummy_models: false model_name: Qwen3.5-9B-EXL3-6hb-4bpw backend: exllamav3 max_seq_len: 131072 cache_size: 131072 cache_mode: Q4 tensor_parallel: false gpu_split_auto: true autosplit_reserve: [192] chunk_size: 2048 output_chunking: true max_batch_size: 2 # Qwen3.5/3.8 think before answering; split that into reasoning_content instead of leaking it as text reasoning: true start_in_reasoning: auto draft_model: draft_mode: mtp sampling: override_preset: memory: sysmem_recurrent_cache: 4096 sysmem_kv_cache: 0 cuda_malloc_async: true EOF
- 3 Start the server
docker run --rm \ --gpus all \ -p 8000:5000 \ --shm-size 8g \ -e NVIDIA_VISIBLE_DEVICES=all \ -v ~/models/Qwen3.5-9B-exl3-e98453aa:/workspace/models/Qwen3.5-9B-EXL3-6hb-4bpw:ro \ -v $PWD/qwen3.5-9b.tabbyapi.128k.yml:/app/config.yml:ro \ --entrypoint /opt/venv/bin/python3 \ ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1 \ main.py \ --config /app/config.yml
✓ loadLoads and serves within an hour
✓ chatAnswers a plain question and stops on its own
✓ reasoningThinks separately and gets 17 × 23 right
✓ toolsCalls a tool with the right arguments and uses the result
✓ contextRecalls a code buried in a prompt that fills 85% of the window
✓ speedDecodes at 15 tok/s or more
Passed all six checks on a real RTX 4070 Ti on Sep 25, 2026.
- Weights
- TheMelonGod/Qwen3.5-9B-exl3 @ e98453aa14 ›
- Image
- ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1
- Engine profile
- tabbyapi-qwen3.5-9b-exl3-4bpw-128k ›
- Recipe file
- registry/recipes/nvidia/rtx-4070-ti-12gb/qwen3.5-9b.tabbyapi.128k.json ›
- Model released
- Feb 26, 2026
- Tested
- vast, RTX 4070 Ti