<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>3hship</title><description>How much can you ship in 3 hours?</description><link>https://3hship.com/</link><language>en</language><item><title>Evaluating local models on one 12 GB GPU</title><link>https://3hship.com/blog/local-models-on-one-12gb-gpu/</link><guid isPermaLink="true">https://3hship.com/blog/local-models-on-one-12gb-gpu/</guid><description>Five local models tested on one RTX 4070 (12 GB) with LM Studio and llama.cpp: Qwen3-VL-8B, Gemma 4 12B QAT, Gemma 4 26B A4B, Qwen3.8-27B and Qwen3.6-35B-A3B. Accuracy against 40 hand-labelled documents, time to first token, tokens per second, MoE expert offload tuning (1.1 to 52.8 tok/s), prompt-size limits and vision, thinking and tool-calling checks, as a method to repeat on your own card.</description><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><category>local-llm</category><category>evals</category><category>lm-studio</category><category>agents</category></item></channel></rss>