LOCAL AIregistry
‹ RTX 3060

Qwen3.8-27B

EXL3 2.2 bpw · tabbyapi · 32K context · RTX 3060 12 GB

Alibaba's 27B dense model from August 2026, and the strongest model that fits one consumer card. It thinks before it answers, calls tools reliably and reads images, which makes it the default for coding agents and long projects.

30 tok/sdecode
336 tok/sprefill
32Kcontext window
Tested on this cardSep 26, 2026
Run it

With Omarchy Local AI it is one button. By hand, it is three steps: the weights, the config, the container. The server then answers on http://localhost:8000/v1.

  1. 1 Download the weights
    hf download turboderp/Qwen3.8-27B-exl3 \
      --revision 25019f1663e0e0bcfc363609f1dbabd5604b04e4 \
      --local-dir ~/models/Qwen3.8-27B-exl3-25019f16
  2. 2 Write the server config
    cat > qwen3.8-27b.tabbyapi.32k.yml <<'EOF'
    network: {host: 0.0.0.0, port: 5000, disable_auth: true, disable_fetch_requests: true, send_tracebacks: false, api_servers: [OAI], sse_ping_interval: 15}
    logging: {log_prompt: false, log_generation_params: false, log_requests: false, log_chat_completion_requests: false}
    model:
      model_dir: /workspace/models
      model_name: Qwen3.8-27B-exl3-25019f16
      backend: exllamav3
      inline_model_loading: false
      max_seq_len: 32768
      cache_size: 33792
      cache_mode: Q4
      tensor_parallel: false
      gpu_split_auto: true
      autosplit_reserve: [256]
      chunk_size: 2048
      output_chunking: true
      max_batch_size: 1
      vision: true
      reasoning: true
      start_in_reasoning: auto
      tool_format: qwen3_coder
      template_vars_default: {}
    draft_model: {draft_mode: mtp}
    memory: {sysmem_recurrent_cache: 4096, sysmem_kv_cache: 0, cuda_malloc_async: true}
    EOF
  3. 3 Start the server
    docker run --rm \
      --gpus all \
      -p 8000:5000 \
      --shm-size 8g \
      -v ~/models/Qwen3.8-27B-exl3-25019f16:/workspace/models/Qwen3.8-27B-exl3-25019f16:ro \
      -v $PWD/qwen3.8-27b.tabbyapi.32k.yml:/app/config.yml:ro \
      --entrypoint /opt/venv/bin/python3 \
      ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1 \
      main.py \
      --config /app/config.yml
What it passed
✓ loadLoads and serves within an hour
✓ chatAnswers a plain question and stops on its own
✓ reasoningThinks separately and gets 17 × 23 right
✓ toolsCalls a tool with the right arguments and uses the result
✓ contextRecalls a code buried in a prompt that fills 85% of the window
✓ speedDecodes at 15 tok/s or more

Passed all six checks on a real RTX 3060 on Sep 26, 2026.

Details
Weights
turboderp/Qwen3.8-27B-exl3 @ 25019f1663 ›
Image
ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1
Engine profile
tabbyapi-exl3 ›
Recipe file
registry/recipes/nvidia/rtx-3060-12gb/qwen3.8-27b.tabbyapi.32k.json ›
Model released
Aug 5, 2026
Tested
vast, RTX 3060