Qwen3.8-27B
EXL3 2.2 bpw · tabbyapi · 32K context · RTX 3060 12 GB
Alibaba's 27B dense model from August 2026, and the strongest model that fits one consumer card. It thinks before it answers, calls tools reliably and reads images, which makes it the default for coding agents and long projects.
30 tok/sdecode
336 tok/sprefill
32Kcontext window
Tested on this cardSep 26, 2026
With Omarchy Local AI it is one button. By hand, it is three steps: the weights, the config, the container. The server then answers on http://localhost:8000/v1.
- 1 Download the weights
hf download turboderp/Qwen3.8-27B-exl3 \ --revision 25019f1663e0e0bcfc363609f1dbabd5604b04e4 \ --local-dir ~/models/Qwen3.8-27B-exl3-25019f16
- 2 Write the server config
cat > qwen3.8-27b.tabbyapi.32k.yml <<'EOF' network: {host: 0.0.0.0, port: 5000, disable_auth: true, disable_fetch_requests: true, send_tracebacks: false, api_servers: [OAI], sse_ping_interval: 15} logging: {log_prompt: false, log_generation_params: false, log_requests: false, log_chat_completion_requests: false} model: model_dir: /workspace/models model_name: Qwen3.8-27B-exl3-25019f16 backend: exllamav3 inline_model_loading: false max_seq_len: 32768 cache_size: 33792 cache_mode: Q4 tensor_parallel: false gpu_split_auto: true autosplit_reserve: [256] chunk_size: 2048 output_chunking: true max_batch_size: 1 vision: true reasoning: true start_in_reasoning: auto tool_format: qwen3_coder template_vars_default: {} draft_model: {draft_mode: mtp} memory: {sysmem_recurrent_cache: 4096, sysmem_kv_cache: 0, cuda_malloc_async: true} EOF - 3 Start the server
docker run --rm \ --gpus all \ -p 8000:5000 \ --shm-size 8g \ -v ~/models/Qwen3.8-27B-exl3-25019f16:/workspace/models/Qwen3.8-27B-exl3-25019f16:ro \ -v $PWD/qwen3.8-27b.tabbyapi.32k.yml:/app/config.yml:ro \ --entrypoint /opt/venv/bin/python3 \ ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1 \ main.py \ --config /app/config.yml
✓ loadLoads and serves within an hour
✓ chatAnswers a plain question and stops on its own
✓ reasoningThinks separately and gets 17 × 23 right
✓ toolsCalls a tool with the right arguments and uses the result
✓ contextRecalls a code buried in a prompt that fills 85% of the window
✓ speedDecodes at 15 tok/s or more
Passed all six checks on a real RTX 3060 on Sep 26, 2026.
- Weights
- turboderp/Qwen3.8-27B-exl3 @ 25019f1663 ›
- Image
- ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1
- Engine profile
- tabbyapi-exl3 ›
- Recipe file
- registry/recipes/nvidia/rtx-3060-12gb/qwen3.8-27b.tabbyapi.32k.json ›
- Model released
- Aug 5, 2026
- Tested
- vast, RTX 3060