LOCAL AIregistry
‹ RTX PRO 4500 Blackwell

Qwen3.8-27B

EXL3 5 bpw · tabbyapi · 256K context · RTX PRO 4500 Blackwell 32 GB

Alibaba's 27B dense model from August 2026, and the strongest model that fits one consumer card. It thinks before it answers, calls tools reliably and reads images, which makes it the default for coding agents and long projects.

65 tok/sdecode
– tok/sprefill
256Kcontext window
Earlier checkSep 22, 2026
Run it

With Omarchy Local AI it is one button. By hand, it is three steps: the weights, the config, the container. The server then answers on http://localhost:8000/v1.

  1. 1 Download the weights
    hf download turboderp/Qwen3.8-27B-exl3 \
      --revision f33f26d929e2b20ef21361145d582f5239e3831f \
      --local-dir ~/models/Qwen3.8-27B-exl3-f33f26d9
  2. 2 Write the server config
    cat > qwen3.8-27b.tabbyapi.256k.yml <<'EOF'
    network:
      host: 0.0.0.0
      port: 5000
      disable_auth: true
      disable_fetch_requests: true
      send_tracebacks: false
      api_servers: ["OAI"]
      sse_ping_interval: 15
    
    logging:
      log_prompt: false
      log_generation_params: false
      log_requests: false
      log_chat_completion_requests: false
    
    model:
      model_dir: /workspace/models
      inline_model_loading: false
      use_dummy_models: false
      model_name: Qwen3.8-27B-EXL3-SC5bpw-H6-V6
      backend: exllamav3
      max_seq_len: 262144
      cache_size: 263168
      cache_mode: Q4
      tensor_parallel: false
      gpu_split_auto: true
      autosplit_reserve: [512]
      chunk_size: 2048
      output_chunking: true
      max_batch_size: 4
      vision: true
      vision_offload: false
      reasoning: true
      start_in_reasoning: always
      reasoning_start_token: " thinking"
      reasoning_end_token: "</think>"
      template_vars_default:
        enable_thinking: true
      tool_format: qwen3_coder
    
    draft_model:
      draft_mode: mtp
    
    sampling:
      override_preset:
    
    memory:
      sysmem_recurrent_cache: 4096
      sysmem_kv_cache: 0
      cuda_malloc_async: true
    EOF
  3. 3 Start the server
    docker run --rm \
      --gpus all \
      -p 8000:5000 \
      --shm-size 8g \
      -e NVIDIA_VISIBLE_DEVICES=all \
      -v ~/models/Qwen3.8-27B-exl3-f33f26d9:/workspace/models/Qwen3.8-27B-EXL3-SC5bpw-H6-V6:ro \
      -v $PWD/qwen3.8-27b.tabbyapi.256k.yml:/app/config.yml:ro \
      --entrypoint /opt/venv/bin/python3 \
      ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1 \
      main.py \
      --config /app/config.yml
What it passed
✓ loadLoads and serves within an hour
✓ chatAnswers a plain question and stops on its own
– reasoningThinks separately and gets 17 × 23 right
– toolsCalls a tool with the right arguments and uses the result
– contextRecalls a code buried in a prompt that fills 85% of the window
– speedDecodes at 15 tok/s or more

Passed the older check (loads and chats); a full six-check run is pending.

Details
Weights
turboderp/Qwen3.8-27B-exl3 @ f33f26d929 ›
Image
ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1
Engine profile
tabbyapi-qwen3.8-27b-exl3-5bpw-256k ›
Recipe file
registry/recipes/nvidia/rtx-pro-4500-blackwell-32gb/qwen3.8-27b.tabbyapi.256k.json ›
Model released
Aug 5, 2026
Tested
earlier acceptance