LOCAL AIregistry
‹ RTX A6000

Gemma 4 26B A4B

EXL3 6.1 bpw · tabbyapi · 128K context · RTX A6000 48 GB

Google's Gemma 4 mixture-of-experts model: 26B total, 4B active, so it runs fast while seeing images and thinking.

87 tok/sdecode
1,894 tok/sprefill
128Kcontext window
Tested on this cardSep 26, 2026
Run it

With Omarchy Local AI it is one button. By hand, it is three steps: the weights, the config, the container. The server then answers on http://localhost:8000/v1.

  1. 1 Download the weights
    hf download turboderp/gemma-4-26B-A4B-it-exl3 \
      --revision 75510f9febfdea0fa22546a2725581cba6984fbb \
      --local-dir ~/models/gemma-4-26B-A4B-it-exl3-75510f9f
  2. 2 Write the server config
    cat > gemma-4-26b-a4b.tabbyapi.128k.yml <<'EOF'
    network: {host: 0.0.0.0, port: 5000, disable_auth: true, disable_fetch_requests: true, send_tracebacks: false, api_servers: [OAI], sse_ping_interval: 15}
    logging: {log_prompt: false, log_generation_params: false, log_requests: false, log_chat_completion_requests: false}
    model:
      model_dir: /workspace/models
      model_name: gemma-4-26B-A4B-it-exl3-75510f9f
      backend: exllamav3
      inline_model_loading: false
      max_seq_len: 131072
      cache_size: 264192
      cache_mode: Q4
      tensor_parallel: false
      gpu_split_auto: true
      autosplit_reserve: [256]
      chunk_size: 2048
      output_chunking: true
      max_batch_size: 2
      vision: true
      reasoning: true
      start_in_reasoning: auto
      tool_format: null
      template_vars_default: {enable_thinking: true}
    draft_model: {}
    memory: {sysmem_recurrent_cache: 4096, sysmem_kv_cache: 0, cuda_malloc_async: true}
    EOF
  3. 3 Start the server
    docker run --rm \
      --gpus all \
      -p 8000:5000 \
      --shm-size 8g \
      -v ~/models/gemma-4-26B-A4B-it-exl3-75510f9f:/workspace/models/gemma-4-26B-A4B-it-exl3-75510f9f:ro \
      -v $PWD/gemma-4-26b-a4b.tabbyapi.128k.yml:/app/config.yml:ro \
      --entrypoint /opt/venv/bin/python3 \
      ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1 \
      main.py \
      --config /app/config.yml
What it passed
✓ loadLoads and serves within an hour
✓ chatAnswers a plain question and stops on its own
✓ reasoningThinks separately and gets 17 × 23 right
✓ toolsCalls a tool with the right arguments and uses the result
✓ contextRecalls a code buried in a prompt that fills 85% of the window
✓ speedDecodes at 15 tok/s or more

Passed all six checks on a real RTX A6000 on Sep 26, 2026.

Details
Weights
turboderp/gemma-4-26B-A4B-it-exl3 @ 75510f9feb ›
Image
ghcr.io/0xsero/tabbyapi-exl3@sha256:0f83e6198dc3be2561652d8df2525a7d1a69733e12fc8e25cbc5bc319a8a3ad1
Engine profile
tabbyapi-exl3 ›
Recipe file
registry/recipes/nvidia/rtx-a6000-48gb/gemma-4-26b-a4b.tabbyapi.128k.json ›
Model released
Mar 12, 2026
Tested
vast, RTX A6000