Skip to content

Bloglocal-llm

Evaluating local models on one 12 GB GPU

Five local models tested on one RTX 4070 (12 GB) with LM Studio and llama.cpp: Qwen3-VL-8B, Gemma 4 12B QAT, Gemma 4 26B A4B, Qwen3.8-27B and Qwen3.6-35B-A3B. Accuracy against 40 hand-labelled documents, time to first token, tokens per second, MoE expert offload tuning (1.1 to 52.8 tok/s), prompt-size limits and vision, thinking and tool-calling checks, as a method to repeat on your own card.

3hship8 min read

Contents

A new local model shows up every week, and the forum numbers rarely match what it does on your card with your work. This is a repeatable way to find out: one real job, a small hand-labelled set, a few speed readings and some quick capability checks.

The worked example is one round on a single RTX 4070 with 12 GB of VRAM, run through LM Studio on llama.cpp. It covered five models:

ModelTypeRan on
Qwen3-VL-8B (Q4_K_M)DenseGPU only
Gemma 4 12B QATDenseGPU only
Gemma 4 26B A4BMixture of expertsGPU and RAM
Qwen3.8-27BDenseGPU and RAM
Qwen3.6-35B-A3BMixture of expertsGPU and RAM

Each was tested for:

  • Accuracy against 40 hand-labelled documents, scored field by field, plus valid JSON
  • Speed: time to the first token and tokens per second
  • Load placement: how far tuning the expert offload moves a mixture-of-experts model (here, 1.1 to 52.8 tok/s)
  • Prompt size: first-token time and recall on prompts of a few thousand tokens
  • Capabilities: text extraction, vision, thinking on and off, and tool calling
  • Memory: peak VRAM against the card’s baseline, and whether a bigger model fits at all

The method at a glance

  1. Pick one real job and hand-label a small set to score against.
  2. Load each model the way you would actually run it, and write the settings down.
  3. Score accuracy and speed together.
  4. Before calling a mixture-of-experts model slow, move its experts.
  5. Test the prompt sizes you will really send.
  6. Run a few small capability checks.
  7. Read memory numbers against a baseline, and do the fit arithmetic before a big download.

1. One job, a small labelled set

Start from work you already do, not a public benchmark. The job here is structured extraction from short documents: for each one the model returns a three-way label with a supporting quote, a yes/no flag, a number and a category. The score is the average agreement with hand labels on those four fields, plus whether the output was valid JSON.

Forty hand-labelled documents was enough to separate the models. Keep the counts in mind when you read the scores: on 40 items, 0.975 is 39 right and 0.90 is 36, so a gap of 0.025 is one document.

Label a second, separate set before you start tuning. The prompt in this round was tuned on the same 40 documents that score it, so the numbers are in-sample, and that is the first thing to change in the next round.

2. Load each model the way you would run it

The round’s setup:

  • GPU: an RTX 4070 with 12 GB of VRAM
  • Runner: LM Studio 0.4.25 on llama.cpp (CUDA)
  • Context window: 8K tokens, 4K for Qwen3.8-27B
  • Layers on the GPU: all of them for the two smaller models, 30 to 60% for the three bigger ones
  • Test set: 40 hand-labelled documents (8 for Qwen3.8-27B)
  1. Gemma 4 12B QAT57.1 s to load
    Share on the GPU
    All of it
    Context window
    8,192 tokens
  2. Qwen3-VL-8B91.6 s to load
    Share on the GPU
    All of it
    Context window
    8,192 tokens
  3. Gemma 4 26B A4B481.2 s to load
    Share on the GPU
    40%
    Context window
    8,192 tokens
  4. Qwen3.8-27B200.3 s to load
    Share on the GPU
    60%
    Context window
    4,096 tokens
  5. Qwen3.6-35B-A3B98.8 s to load
    Share on the GPU
    30%
    Context window
    8,192 tokens

    Qwen3-VL-8B was still loaded in the GPU.

Three things worth knowing before you start:

  • Loading is slow. It took from about a minute to eight minutes per model. Budget for it.
  • Unload everything else first. During the Qwen3.6-35B-A3B run another model was still on the card, and its numbers paid for it.
  • A failed load is not a failed model. In an earlier round Qwen3.8-27B wouldn’t load on an older llama.cpp runtime. On a newer one it loaded and ran. Update the runtime and retry before you drop a model.

3. Score accuracy and speed together

Every model gave zero wrong “yes” answers. Switch the metric to see how the order changes; the dashed tag marks the model that ran on 8 documents instead of 40.

Extraction suite, one row per model. Pick a metric; bars start at zero.

Mean of four agreements with the hand labels. The Qwen3.8-27B value is from 8 documents and cannot be compared with the other four.

Show every metric as a table
Extraction suite and the 2 to 4k-token prompt. Qwen3.8-27B ran on 8 documents; the others on 40. Every model scored zero wrong "yes" answers.
ModelnMean scoreMain labelCategoryValid JSONTokens per second (tok/s)First token (s)First token, long prompt (s)
Qwen3-VL-8B400.9221.0000.8161.0075.50.4918.96
Gemma 4 12B QAT400.9350.9000.8951.0042.31.573.47
Gemma 4 26B A4B400.9470.9750.8681.0024.92.8119.18
Qwen3.6-35B-A3B400.9090.9000.8950.955.236.75156.81
Qwen3.8-27B80.9691.0000.8751.008.424.7329.41

Qwen3-VL-8B (Q4_K_M) won. It matched the main label on all 40, always returned valid JSON, and was the fastest accurate model at half a second to the first token and 75.5 tokens per second. It did not win on the average: its 0.922 is second lowest of the four models that ran all 40 (0.909 to 0.947), and its category agreement, 0.816, was the lowest of the five.

That is the point of scoring fields separately. Decide which field the job can’t get wrong, then pick on that and on speed. A one-document lead on the average shouldn’t decide anything.

Qwen3.8-27B had the highest average, 0.969, but on only 8 documents and at almost 25 seconds to the first token and 8.4 tokens per second. A small sample like that is a reason to rerun it on the full set, not a result.

4. Before calling a model slow, move its experts

Two mixture-of-experts models, Qwen3.6-35B-A3B and Gemma 4 26B A4B, came in far below what forum posts reported. One poster saw 40 to 50 tok/s for Gemma 4 26B A4B on a laptop with 8 GB of VRAM, up from about 20 before tuning[1]. Another saw 52 to 59 tok/s for Qwen3.6-35B-A3B on a 12 GB laptop GPU with short prompts[2], and a third thread reported 14 to 35 tok/s on 12 GB cards[3]. On this card they first ran at 24.9 and 5.2 tok/s.

The difference was where the weights lived. A mixture-of-experts model only uses a few experts per token, so the expert weights can sit in system RAM while every other layer stays on the GPU. In llama.cpp that is --n-cpu-moe N, which keeps the experts of the first N layers on the CPU. A plain GPU layer ratio, which is what LM Studio’s default load uses, moves whole layers instead.

Terminal window
# a commenter's 12 GB RTX 3060 setup, excerpt (llama.cpp)
llama-server -m model.gguf --n-gpu-layers 99 --n-cpu-moe 31 --ctx-size 80128 --flash-attn on --no-mmap
# the 12 GB laptop thread's command shape (llama.cpp)
llama-cli -c 4096 -fa on -ctk q8_0 -ctv q8_0 --fit on -fitt 768 --temp 0 --single-turn
# the first load in this round (LM Studio): a whole-layer ratio, no expert offload
lms load qwen3.6-35b-a3b -y --gpu 0.3 -c 8192

Here is the same Qwen3.6-35B-A3B on the same card at the same 8K context, with only the placement changing:

tok/sPlacement
1.1LM Studio’s default: 16 of 41 layers on the GPU, everything else on the CPU
4.5Every layer on the GPU, all experts in RAM, while another job shared the card
34.1The same load on an idle card, using 7.7 GB of VRAM
52.817 expert layers on the GPU and 23 in RAM, using 10.8 GB of VRAM

About 48 times faster, with the same model on the same hardware, and in line with the forum’s 52 to 59. Gemma 4 26B A4B moved the same way, from 19 to 57.6 tok/s.

To find the setting on your own card:

  1. Start with every expert in RAM (--n-cpu-moe set to the layer count) on an idle GPU.
  2. Lower the number a few layers at a time, so more experts move onto the GPU.
  3. Stop when the speed stops rising or the readings start to swing. Here 23 held at 50 to 55 tok/s, while 22 read 48 once and 29 twice, and lower than that got slower or failed to load.
  4. Take 3 to 5 readings per setting, and look at the spread as well as the best one.

In LM Studio the same control is the per-model load setting for expert layers on the CPU. It saves with the model and becomes --n-cpu-moe when the model loads. One more catch: a thinking model spends tokens before even a short answer, so also time a finished answer, not just tokens per second.

5. Test the prompt sizes you will send

The best model changed with prompt size. For prompts under 2K tokens Qwen3-VL-8B won. For prompts of a few thousand tokens every model found all three hidden facts in the long-context probe, but the wait for the first token was very different:

ModelFirst token
Gemma 4 12B QAT3.5 s
Qwen3-VL-8B19.0 s
Gemma 4 26B A4B19.2 s
Qwen3.8-27B29.4 s
Qwen3.6-35B-A3B156.8 s

Those prompts ran from about 2K to 4K tokens. Load the context window your longest real prompt needs: with an 8K window here, the 16K-and-up class couldn’t be tested at all.

6. Run a few small capability checks

Each model also ran short checks of 2 to 10 items. They tell you whether a feature works, not which model is best at it.

Capability checks and the extraction-suite mean, per model. Each cell is a score from 0 to 1, shaded by rank so close scores stay apart; a blank cell was not claimed, not zero. Hover, focus or tap a cell for its item count.
ModelText extractionVisionVideoTool callingLong contextThinking onThinking offSuite mean
Qwen3-VL-8B1.000.751.001.001.00not claimednot claimed0.922
Gemma 4 12B QAT1.001.001.001.001.000.751.000.935
Gemma 4 26B A4B0.801.001.001.001.001.001.000.947
Qwen3.6-35B-A3B1.000.751.001.001.001.000.750.909
Qwen3.8-27B1.001.001.001.001.001.001.000.969

The checks are small, with 2 to 10 items each. On a four-item check, one item moves a cell from 1.00 to 0.75.

  • Text extraction. Four models got 10 of 10 fields. Gemma 4 26B A4B got 8 of 10.
  • Vision. Qwen3-VL-8B and Qwen3.6-35B-A3B scored 0.75, each missing the code-reading item (Qwen3-VL-8B read 4821 as 4211). The other three scored 1.00.
  • Thinking. Test it on and off. Gemma 4 12B QAT answered 3 of 4 with thinking on and 4 of 4 with it off, while writing about 2,500 tokens per answer against about 240. Qwen3.6-35B-A3B went the other way: 4 of 4 on, 3 of 4 off.
  • Tool calling. All five scored 1.00.

7. Read memory against a baseline

Peak VRAM for every model read 11.4 to 11.7 GiB, which looks like each one filled the card. It didn’t: before four of the five loads the card was already at 11.3 to 11.6 GiB. Read the card before you load a model and subtract, or measure the model on its own. Measured alone, Qwen3-VL-8B used about 9.4 GB.

What the GPU memory reading shows

Peak VRAM per model against a 12 GiB card (12,288 MiB). The dark part is the baseline read before the load, so the bar is the whole GPU, not the model.

  1. Gemma 4 12B QAT11,615 → 11,916 MiB
  2. Qwen3-VL-8B11,812 → 11,832 MiB
  3. Gemma 4 26B A4B11,831 → 11,983 MiB
  4. Qwen3.8-27B11,866 → 11,939 MiB
  5. Qwen3.6-35B-A3B3,027 → 11,917 MiB

Would Qwen3.8-Flash-Next fit?

The smallest version of the model is 67.56 GiB. Choose where it has to go.

RAM + VRAM: 67.56 GiB of weights against 75.2 GiB leaves 7.64 GiB for the operating system and the model's working memory. Unsloth gives 75 GB of RAM plus VRAM as the minimum for this version. So it was "not attempted", not "does not fit".

Show the numbers as a table
Qwen3.8-Flash-Next, Unsloth's smallest 1-bit version: 67.56 GiB. Differences are derived.
WhereAvailableWeights minus available
The GPU alone12.0 GiB of VRAM55.56 GiB short
RAM + VRAM75.2 GiB (63.2 GiB RAM + 12 GiB VRAM)-7.64 GiB spare

Do the fit arithmetic before a big download, and check the free disk space too. The smallest Qwen3.8-Flash-Next download, Unsloth’s 1-bit quant, is 67.6 GiB. Against 63 GiB of RAM plus 12 GiB of VRAM that leaves about 7.6 GiB for everything else, under Unsloth’s stated 75 GB minimum, so it wasn’t attempted. The explorer above runs the same arithmetic.

What came out of it

  • Short documents: Qwen3-VL-8B with the tuned prompt. At about 81 tok/s, 61,000 documents take roughly 18 hours.
  • Prompts of a few thousand tokens: Gemma 4 12B QAT.
  • Next round: a held-out set, a real long-prompt run, the suite rerun with tuned expert placement, and Qwen3.8-27B on all 40.

Sources

  1. r/LocalLLM - 40 t/s with 8 GB VRAM on Gemma 4 26B A4B MoEThe poster reports 40 to 50 tok/s on average for Gemma 4 26B A4B on a laptop with 8 GB of VRAM and 32 GB of RAM, up from about 20 before tuning, with maxed GPU layers, 20 MoE layers on the CPU, a 65,536-token context, a q8 KV cache and MTP speculative decoding with 2 draft tokens.Checked
  2. r/LocalLLaMA - Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12 GBOn a 12 GB laptop GPU with a 4096 context and short prompts, the poster reports Qwen3.6-35B-A3B Q4_K_M at 52 to 59 tok/s, Qwen3.8 27B dense at 7.5 to 9.1 tok/s for Q3 and about 35 tok/s for Q2, where Q2 missed one of six sanity questions. The poster says it is not a real eval.Checked
  3. r/LocalLLM - Best llama.cpp settings for 12 GB VRAM and 32 GB RAMA question thread. Commenters share llama.cpp configs that keep expert weights in RAM (--n-cpu-moe or --cpu-moe) with flash attention on and a q8_0 KV cache: 35 tok/s for Qwen3.6-35B-A3B on a 12 GB RTX 3060, about 14 tok/s at Q6, and 20 to 25 tok/s for the original poster on a 12 GB AMD card.Checked