Evaluating local models on one 12 GB GPU
Five local models tested on one RTX 4070 (12 GB) with LM Studio and llama.cpp: Qwen3-VL-8B, Gemma 4 12B QAT, Gemma 4 26B A4B, Qwen3.8-27B and Qwen3.6-35B-A3B. Accuracy against 40 hand-labelled documents, time to first token, tokens per second, MoE expert offload tuning (1.1 to 52.8 tok/s), prompt-size limits and vision, thinking and tool-calling checks, as a method to repeat on your own card.
A new local model shows up every week, and the forum numbers rarely match what it does on your card with your work. This is a repeatable way to find out: one real job, a small hand-labelled set, a few speed readings and some quick capability checks.
The worked example is one round on a single RTX 4070 with 12 GB of VRAM, run through LM Studio on llama.cpp. It covered five models:
| Model | Type | Ran on |
|---|---|---|
| Qwen3-VL-8B (Q4_K_M) | Dense | GPU only |
| Gemma 4 12B QAT | Dense | GPU only |
| Gemma 4 26B A4B | Mixture of experts | GPU and RAM |
| Qwen3.8-27B | Dense | GPU and RAM |
| Qwen3.6-35B-A3B | Mixture of experts | GPU and RAM |
Each was tested for:
- Accuracy against 40 hand-labelled documents, scored field by field, plus valid JSON
- Speed: time to the first token and tokens per second
- Load placement: how far tuning the expert offload moves a mixture-of-experts model (here, 1.1 to 52.8 tok/s)
- Prompt size: first-token time and recall on prompts of a few thousand tokens
- Capabilities: text extraction, vision, thinking on and off, and tool calling
- Memory: peak VRAM against the card’s baseline, and whether a bigger model fits at all
The method at a glance
- Pick one real job and hand-label a small set to score against.
- Load each model the way you would actually run it, and write the settings down.
- Score accuracy and speed together.
- Before calling a mixture-of-experts model slow, move its experts.
- Test the prompt sizes you will really send.
- Run a few small capability checks.
- Read memory numbers against a baseline, and do the fit arithmetic before a big download.
1. One job, a small labelled set
Start from work you already do, not a public benchmark. The job here is structured extraction from short documents: for each one the model returns a three-way label with a supporting quote, a yes/no flag, a number and a category. The score is the average agreement with hand labels on those four fields, plus whether the output was valid JSON.
Forty hand-labelled documents was enough to separate the models. Keep the counts in mind when you read the scores: on 40 items, 0.975 is 39 right and 0.90 is 36, so a gap of 0.025 is one document.
Label a second, separate set before you start tuning. The prompt in this round was tuned on the same 40 documents that score it, so the numbers are in-sample, and that is the first thing to change in the next round.
2. Load each model the way you would run it
The round’s setup:
- GPU: an RTX 4070 with 12 GB of VRAM
- Runner: LM Studio 0.4.25 on llama.cpp (CUDA)
- Context window: 8K tokens, 4K for Qwen3.8-27B
- Layers on the GPU: all of them for the two smaller models, 30 to 60% for the three bigger ones
- Test set: 40 hand-labelled documents (8 for Qwen3.8-27B)
- Gemma 4 12B QAT57.1 s to load
- Share on the GPU
- All of it
- Context window
- 8,192 tokens
- Qwen3-VL-8B91.6 s to load
- Share on the GPU
- All of it
- Context window
- 8,192 tokens
- Gemma 4 26B A4B481.2 s to load
- Share on the GPU
- 40%
- Context window
- 8,192 tokens
- Qwen3.8-27B200.3 s to load
- Share on the GPU
- 60%
- Context window
- 4,096 tokens
- Qwen3.6-35B-A3B98.8 s to load
- Share on the GPU
- 30%
- Context window
- 8,192 tokens
Qwen3-VL-8B was still loaded in the GPU.
Three things worth knowing before you start:
- Loading is slow. It took from about a minute to eight minutes per model. Budget for it.
- Unload everything else first. During the Qwen3.6-35B-A3B run another model was still on the card, and its numbers paid for it.
- A failed load is not a failed model. In an earlier round Qwen3.8-27B wouldn’t load on an older llama.cpp runtime. On a newer one it loaded and ran. Update the runtime and retry before you drop a model.
3. Score accuracy and speed together
Every model gave zero wrong “yes” answers. Switch the metric to see how the order changes; the dashed tag marks the model that ran on 8 documents instead of 40.
Mean of four agreements with the hand labels. The Qwen3.8-27B value is from 8 documents and cannot be compared with the other four.
Show every metric as a table
| Model | n | Mean score | Main label | Category | Valid JSON | Tokens per second (tok/s) | First token (s) | First token, long prompt (s) |
|---|---|---|---|---|---|---|---|---|
| Qwen3-VL-8B | 40 | 0.922 | 1.000 | 0.816 | 1.00 | 75.5 | 0.49 | 18.96 |
| Gemma 4 12B QAT | 40 | 0.935 | 0.900 | 0.895 | 1.00 | 42.3 | 1.57 | 3.47 |
| Gemma 4 26B A4B | 40 | 0.947 | 0.975 | 0.868 | 1.00 | 24.9 | 2.81 | 19.18 |
| Qwen3.6-35B-A3B | 40 | 0.909 | 0.900 | 0.895 | 0.95 | 5.2 | 36.75 | 156.81 |
| Qwen3.8-27B | 8 | 0.969 | 1.000 | 0.875 | 1.00 | 8.4 | 24.73 | 29.41 |
Qwen3-VL-8B (Q4_K_M) won. It matched the main label on all 40, always returned valid JSON, and was the fastest accurate model at half a second to the first token and 75.5 tokens per second. It did not win on the average: its 0.922 is second lowest of the four models that ran all 40 (0.909 to 0.947), and its category agreement, 0.816, was the lowest of the five.
That is the point of scoring fields separately. Decide which field the job can’t get wrong, then pick on that and on speed. A one-document lead on the average shouldn’t decide anything.
Qwen3.8-27B had the highest average, 0.969, but on only 8 documents and at almost 25 seconds to the first token and 8.4 tokens per second. A small sample like that is a reason to rerun it on the full set, not a result.
4. Before calling a model slow, move its experts
Two mixture-of-experts models, Qwen3.6-35B-A3B and Gemma 4 26B A4B, came in far below what forum posts reported. One poster saw 40 to 50 tok/s for Gemma 4 26B A4B on a laptop with 8 GB of VRAM, up from about 20 before tuning[1]. Another saw 52 to 59 tok/s for Qwen3.6-35B-A3B on a 12 GB laptop GPU with short prompts[2], and a third thread reported 14 to 35 tok/s on 12 GB cards[3]. On this card they first ran at 24.9 and 5.2 tok/s.
The difference was where the weights lived. A mixture-of-experts model only uses a few experts per
token, so the expert weights can sit in system RAM while every other layer stays on the GPU. In
llama.cpp that is --n-cpu-moe N, which keeps the experts of the first N layers on the CPU. A plain
GPU layer ratio, which is what LM Studio’s default load uses, moves whole layers instead.
# a commenter's 12 GB RTX 3060 setup, excerpt (llama.cpp)llama-server -m model.gguf --n-gpu-layers 99 --n-cpu-moe 31 --ctx-size 80128 --flash-attn on --no-mmap
# the 12 GB laptop thread's command shape (llama.cpp)llama-cli -c 4096 -fa on -ctk q8_0 -ctv q8_0 --fit on -fitt 768 --temp 0 --single-turn
# the first load in this round (LM Studio): a whole-layer ratio, no expert offloadlms load qwen3.6-35b-a3b -y --gpu 0.3 -c 8192Here is the same Qwen3.6-35B-A3B on the same card at the same 8K context, with only the placement changing:
| tok/s | Placement |
|---|---|
| 1.1 | LM Studio’s default: 16 of 41 layers on the GPU, everything else on the CPU |
| 4.5 | Every layer on the GPU, all experts in RAM, while another job shared the card |
| 34.1 | The same load on an idle card, using 7.7 GB of VRAM |
| 52.8 | 17 expert layers on the GPU and 23 in RAM, using 10.8 GB of VRAM |
About 48 times faster, with the same model on the same hardware, and in line with the forum’s 52 to 59. Gemma 4 26B A4B moved the same way, from 19 to 57.6 tok/s.
To find the setting on your own card:
- Start with every expert in RAM (
--n-cpu-moeset to the layer count) on an idle GPU. - Lower the number a few layers at a time, so more experts move onto the GPU.
- Stop when the speed stops rising or the readings start to swing. Here 23 held at 50 to 55 tok/s, while 22 read 48 once and 29 twice, and lower than that got slower or failed to load.
- Take 3 to 5 readings per setting, and look at the spread as well as the best one.
In LM Studio the same control is the per-model load setting for expert layers on the CPU. It saves
with the model and becomes --n-cpu-moe when the model loads. One more catch: a thinking model
spends tokens before even a short answer, so also time a finished answer, not just tokens per
second.
5. Test the prompt sizes you will send
The best model changed with prompt size. For prompts under 2K tokens Qwen3-VL-8B won. For prompts of a few thousand tokens every model found all three hidden facts in the long-context probe, but the wait for the first token was very different:
| Model | First token |
|---|---|
| Gemma 4 12B QAT | 3.5 s |
| Qwen3-VL-8B | 19.0 s |
| Gemma 4 26B A4B | 19.2 s |
| Qwen3.8-27B | 29.4 s |
| Qwen3.6-35B-A3B | 156.8 s |
Those prompts ran from about 2K to 4K tokens. Load the context window your longest real prompt needs: with an 8K window here, the 16K-and-up class couldn’t be tested at all.
6. Run a few small capability checks
Each model also ran short checks of 2 to 10 items. They tell you whether a feature works, not which model is best at it.
| Model | Text extraction | Vision | Video | Tool calling | Long context | Thinking on | Thinking off | Suite mean |
|---|---|---|---|---|---|---|---|---|
| Qwen3-VL-8B | 1.00 | 0.75 | 1.00 | 1.00 | 1.00 | not claimed | not claimed | 0.922 |
| Gemma 4 12B QAT | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.75 | 1.00 | 0.935 |
| Gemma 4 26B A4B | 0.80 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.947 |
| Qwen3.6-35B-A3B | 1.00 | 0.75 | 1.00 | 1.00 | 1.00 | 1.00 | 0.75 | 0.909 |
| Qwen3.8-27B | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.969 |
The checks are small, with 2 to 10 items each. On a four-item check, one item moves a cell from 1.00 to 0.75.
- Text extraction. Four models got 10 of 10 fields. Gemma 4 26B A4B got 8 of 10.
- Vision. Qwen3-VL-8B and Qwen3.6-35B-A3B scored 0.75, each missing the code-reading item (Qwen3-VL-8B read 4821 as 4211). The other three scored 1.00.
- Thinking. Test it on and off. Gemma 4 12B QAT answered 3 of 4 with thinking on and 4 of 4 with it off, while writing about 2,500 tokens per answer against about 240. Qwen3.6-35B-A3B went the other way: 4 of 4 on, 3 of 4 off.
- Tool calling. All five scored 1.00.
7. Read memory against a baseline
Peak VRAM for every model read 11.4 to 11.7 GiB, which looks like each one filled the card. It didn’t: before four of the five loads the card was already at 11.3 to 11.6 GiB. Read the card before you load a model and subtract, or measure the model on its own. Measured alone, Qwen3-VL-8B used about 9.4 GB.
What the GPU memory reading shows
Peak VRAM per model against a 12 GiB card (12,288 MiB). The dark part is the baseline read before the load, so the bar is the whole GPU, not the model.
- Gemma 4 12B QAT11,615 → 11,916 MiB
- Qwen3-VL-8B11,812 → 11,832 MiB
- Gemma 4 26B A4B11,831 → 11,983 MiB
- Qwen3.8-27B11,866 → 11,939 MiB
- Qwen3.6-35B-A3B3,027 → 11,917 MiB
Would Qwen3.8-Flash-Next fit?
The smallest version of the model is 67.56 GiB. Choose where it has to go.
RAM + VRAM: 67.56 GiB of weights against 75.2 GiB leaves 7.64 GiB for the operating system and the model's working memory. Unsloth gives 75 GB of RAM plus VRAM as the minimum for this version. So it was "not attempted", not "does not fit".
Show the numbers as a table
| Where | Available | Weights minus available |
|---|---|---|
| The GPU alone | 12.0 GiB of VRAM | 55.56 GiB short |
| RAM + VRAM | 75.2 GiB (63.2 GiB RAM + 12 GiB VRAM) | -7.64 GiB spare |
Do the fit arithmetic before a big download, and check the free disk space too. The smallest Qwen3.8-Flash-Next download, Unsloth’s 1-bit quant, is 67.6 GiB. Against 63 GiB of RAM plus 12 GiB of VRAM that leaves about 7.6 GiB for everything else, under Unsloth’s stated 75 GB minimum, so it wasn’t attempted. The explorer above runs the same arithmetic.
What came out of it
- Short documents: Qwen3-VL-8B with the tuned prompt. At about 81 tok/s, 61,000 documents take roughly 18 hours.
- Prompts of a few thousand tokens: Gemma 4 12B QAT.
- Next round: a held-out set, a real long-prompt run, the suite rerun with tuned expert placement, and Qwen3.8-27B on all 40.
Sources
- r/LocalLLM - 40 t/s with 8 GB VRAM on Gemma 4 26B A4B MoEThe poster reports 40 to 50 tok/s on average for Gemma 4 26B A4B on a laptop with 8 GB of VRAM and 32 GB of RAM, up from about 20 before tuning, with maxed GPU layers, 20 MoE layers on the CPU, a 65,536-token context, a q8 KV cache and MTP speculative decoding with 2 draft tokens.
- r/LocalLLaMA - Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12 GBOn a 12 GB laptop GPU with a 4096 context and short prompts, the poster reports Qwen3.6-35B-A3B Q4_K_M at 52 to 59 tok/s, Qwen3.8 27B dense at 7.5 to 9.1 tok/s for Q3 and about 35 tok/s for Q2, where Q2 missed one of six sanity questions. The poster says it is not a real eval.
- r/LocalLLM - Best llama.cpp settings for 12 GB VRAM and 32 GB RAMA question thread. Commenters share llama.cpp configs that keep expert weights in RAM (--n-cpu-moe or --cpu-moe) with flash attention on and a q8_0 KV cache: 35 tok/s for Qwen3.6-35B-A3B on a 12 GB RTX 3060, about 14 tok/s at Q6, and 20 to 25 tok/s for the original poster on a 12 GB AMD card.