LOCAL AIregistry
‹ RTX 4070

Qwen3.5-9B

EXL3 4 bpw · tabbyapi · 128K context · RTX 4070 12 GB

Alibaba's 9B model from February 2026: small enough for an 8 GB card, with a 64K–128K window. It thinks, calls tools and stays quick on modest hardware.

104 tok/sdecode
1,449 tok/sprefill
128Kcontext window
Tested on this cardSep 25, 2026
Run it

With Omarchy Local AI it is one button. By hand, it is three steps: the weights, the config, the container. The server then answers on http://localhost:8000/v1.

  1. 1 Download the weights
    hf download TheMelonGod/Qwen3.5-9B-exl3 \
      --revision e98453aa143a24fab9eb14b718daed02e2fa6eef \
      --local-dir ~/models/Qwen3.5-9B-exl3-e98453aa
  2. 2 Write the server config
    cat > qwen3.5-9b.tabbyapi.128k.yml <<'EOF'
    network:
      host: 0.0.0.0
      port: 5000
      disable_auth: true
      disable_fetch_requests: true
      send_tracebacks: false
      api_servers: ["OAI"]
      sse_ping_interval: 15
    
    logging:
      log_prompt: false
      log_generation_params: false
      log_requests: false
      log_chat_completion_requests: false
    
    model:
      model_dir: /workspace/models
      inline_model_loading: false
      use_dummy_models: false
      model_name: Qwen3.5-9B-EXL3-6hb-4bpw
      backend: exllamav3
      max_seq_len: 131072
      cache_size: 131072
      cache_mode: Q4
      tensor_parallel: false
      gpu_split_auto: true
      autosplit_reserve: [192]
      chunk_size: 2048
      output_chunking: true
      max_batch_size: 2
      # Qwen3.5/3.8 think before answering; split that into reasoning_content instead of leaking it as text
      reasoning: true
      start_in_reasoning: auto
    
    draft_model:
      draft_mode: mtp
    
    sampling:
      override_preset:
    
    memory:
      sysmem_recurrent_cache: 4096
      sysmem_kv_cache: 0
      cuda_malloc_async: true
    EOF
  3. 3 Start the server
    docker run --rm \
      --gpus all \
      -p 8000:5000 \
      --shm-size 8g \
      -e NVIDIA_VISIBLE_DEVICES=all \
      -v ~/models/Qwen3.5-9B-exl3-e98453aa:/workspace/models/Qwen3.5-9B-EXL3-6hb-4bpw:ro \
      -v $PWD/qwen3.5-9b.tabbyapi.128k.yml:/app/config.yml:ro \
      --entrypoint /opt/venv/bin/python3 \
      ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1 \
      main.py \
      --config /app/config.yml
What it passed
✓ loadLoads and serves within an hour
✓ chatAnswers a plain question and stops on its own
✓ reasoningThinks separately and gets 17 × 23 right
✓ toolsCalls a tool with the right arguments and uses the result
✓ contextRecalls a code buried in a prompt that fills 85% of the window
✓ speedDecodes at 15 tok/s or more

Passed all six checks on a real RTX 4070 on Sep 25, 2026.

Details
Weights
TheMelonGod/Qwen3.5-9B-exl3 @ e98453aa14 ›
Image
ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1
Engine profile
tabbyapi-qwen3.5-9b-exl3-4bpw-128k ›
Recipe file
registry/recipes/nvidia/rtx-4070-12gb/qwen3.5-9b.tabbyapi.128k.json ›
Model released
Feb 26, 2026
Tested
vast, RTX 4070