Qwen3.8-27B
EXL3 6 bpw · tabbyapi · 256K context · RTX PRO 6000 Blackwell 96 GB
Alibaba's 27B dense model from August 2026, and the strongest model that fits one consumer card. It thinks before it answers, calls tools reliably and reads images, which makes it the default for coding agents and long projects.
52 tok/sdecode
568 tok/sprefill
256Kcontext window
Tested on this cardSep 25, 2026
With Omarchy Local AI it is one button. By hand, it is three steps: the weights, the config, the container. The server then answers on http://localhost:8000/v1.
- 1 Download the weights
hf download turboderp/Qwen3.8-27B-exl3 \ --revision 60d005a257b39ecb25e4ba23c1dd29df877d5c69 \ --local-dir ~/models/Qwen3.8-27B-exl3-60d005a2
- 2 Write the server config
cat > qwen3.8-27b.tabbyapi.256k.yml <<'EOF' network: {host: 0.0.0.0, port: 5000, disable_auth: true, disable_fetch_requests: true, send_tracebacks: false, api_servers: [OAI], sse_ping_interval: 15} logging: {log_prompt: false, log_generation_params: false, log_requests: false, log_chat_completion_requests: false} model: model_dir: /workspace/models model_name: Qwen3.8-27B-exl3-60d005a2 backend: exllamav3 inline_model_loading: false max_seq_len: 262144 cache_size: 1052672 cache_mode: Q8 tensor_parallel: false gpu_split_auto: true autosplit_reserve: [256] chunk_size: 2048 output_chunking: true max_batch_size: 4 vision: true reasoning: true start_in_reasoning: always tool_format: qwen3_coder template_vars_default: {} draft_model: {draft_mode: mtp} memory: {sysmem_recurrent_cache: 4096, sysmem_kv_cache: 0, cuda_malloc_async: true} EOF - 3 Start the server
docker run --rm \ --gpus all \ -p 8000:5000 \ --shm-size 8g \ -v ~/models/Qwen3.8-27B-exl3-60d005a2:/workspace/models/Qwen3.8-27B-exl3-60d005a2:ro \ -v $PWD/qwen3.8-27b.tabbyapi.256k.yml:/app/config.yml:ro \ --entrypoint /opt/venv/bin/python3 \ ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1 \ main.py \ --config /app/config.yml
✓ loadLoads and serves within an hour
✓ chatAnswers a plain question and stops on its own
✓ reasoningThinks separately and gets 17 × 23 right
✓ toolsCalls a tool with the right arguments and uses the result
✓ contextRecalls a code buried in a prompt that fills 85% of the window
✓ speedDecodes at 15 tok/s or more
Passed all six checks on a real RTX PRO 6000 WS on Sep 25, 2026.
- Weights
- turboderp/Qwen3.8-27B-exl3 @ 60d005a257 ›
- Image
- ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1
- Engine profile
- tabbyapi-exl3 ›
- Recipe file
- registry/recipes/nvidia/rtx-pro-6000-blackwell-96gb/qwen3.8-27b.tabbyapi.256k.json ›
- Model released
- Aug 5, 2026
- Tested
- vast, RTX PRO 6000 WS