Saltar al contenido
Murray's Lab

Lab · Local LLM Arena

¿Qué modelo local va mejor en qué?

Mismas pruebas, mismos prompts, juez local. Cada modelo de Ollama pasa por 7 categorías (escritura, código, frontend, agente, razonamiento, instrucciones, roleplay) y se puntúa de 1 a 10. Resultados reproducibles, sin nube.

18 modelos evaluados · 160 pruebas · 15 categorías

gemma4:31b-it-q8_0 ollama 31.3B Q8_0 31.5 GB 262K ctx juez: crimson-31b:latest Vision Tools Thinking 9.8 9.8
  • Núcleo 9.7
  • Hermes 9.7
  • OpenClaw 9.7
9.4
  • Python 9.3
  • SQL 10.0
10.0 10.0 10.0 10.0 9.3 9.7 9.7 93/101 54K ↓26K ↑28K 6.3 1 h 18 min 5 s 27 may 2026
crimson-31b:latest ollama 31.3B Q4_K_M 18.5 GB 262K ctx juez: gemma4:31b Vision Tools Thinking 9.5 9.7
  • Núcleo 9.7
  • Hermes 9.7
  • OpenClaw 9.5
9.4
  • Python 9.3
  • SQL 10.0
10.0 9.9 10.0 10.0 9.4 9.7 9.8 8.0 8.4 9.9 120/128 69K ↓32K ↑37K 7.1 1 h 41 min 3 s 26 may 2026
gemma4:31b ollama 31.3B Q4_K_M 18.5 GB 262K ctx juez: gemma4:e4b Vision Tools Thinking 9.2 9.6
  • Núcleo 9.5
  • Hermes 9.6
  • OpenClaw 9.5
9.2
  • Python 9.0
  • SQL 10.0
9.7 9.7 10.0 10.0 8.4 9.7 8.7 7.5 7.8 9.7 119/128 ⚠1 68K ↓32K ↑35K 9.3 1 h 16 min 36 s 26 may 2026
gemma4:e4b ollama 8.0B Q4_K_M 8.9 GB 131K ctx juez: gemma4:31b Vision audio Tools Thinking 9.0 9.5
  • Núcleo 9.2
  • Hermes 9.5
  • OpenClaw 9.7
9.4
  • Python 9.3
  • SQL 10.0
9.3 9.1 9.4 9.0 8.8 8.8 9.2 8.6 7.5 9.8 8.9 128/128 82K ↓34K ↑49K 27.5 43 min 41 s 26 may 2026
gemma4:e4b-it-q8_0 ollama 8.0B Q8_0 10.8 GB 131K ctx juez: qwen3.6:latest Vision audio Tools Thinking 8.9 9.7
  • Núcleo 10.0
  • Hermes 9.4
  • OpenClaw 9.7
9.2
  • Python 9.0
  • SQL 10.0
8.9 9.4 9.4 10.0 8.3 8.8 9.6 7.4 7.0 9.6 8.9 128/128 80K ↓34K ↑47K 33.3 28 min 30 s 26 may 2026
qwen3.6:latest ollama 36.0B Q4_K_M 22.3 GB 262K ctx juez: deepseek-r1:32b Vision Tools Thinking 8.9 9.4
  • Núcleo 1.4
  • Hermes 9.1
  • OpenClaw 8.7
8.7
  • Python 7.4
  • SQL 10.0
7.2 9.2 9.8 9.3 8.8 9.6 9.0 8.8 7.4 9.6 111/128 ⚠9 76K ↓30K ↑46K 12.7 2 h 13 min 39 s 25 may 2026
gemma4:31b-it-bf16 ollama 31.3B F16 58.3 GB 262K ctx juez: gemma4:31b-it-q8_0 Vision Tools Thinking 8.8 9.7
  • Núcleo 9.5
  • Hermes 9.7
  • OpenClaw 9.7
9.4
  • Python 9.3
  • SQL 10.0
10.0 9.8 10.0 10.0 9.4 9.7 9.8 8.0 8.3 9.8 0.0 124/136 ⚠4 68K ↓33K ↑35K 3.5 3 h 32 min 56 s 27 may 2026
gemma4:e4b-it-bf16 ollama 8.0B F16 14.9 GB 131K ctx juez: gemma4:31b-it-q8_0 Vision audio Tools Thinking 8.8 9.6
  • Núcleo 9.9
  • Hermes 9.4
  • OpenClaw 9.7
9.2
  • Python 9.0
  • SQL 10.0
9.3 9.4 9.2 10.0 8.3 10.0 9.2 7.4 7.5 9.8 9.0 5.0 136/136 85K ↓35K ↑51K 20.0 44 min 8 s 27 may 2026
qwen3.8:latest ollama 27.3B Q4_K_M 16.5 GB 262K ctx juez: gemma4:31b Vision Tools Thinking 8.8 9.8
  • Núcleo 9.4
  • Hermes 9.9
  • OpenClaw 10.0
8.9
  • Python 9.3
  • TypeScript 9.1
  • Rust 8.2
  • Go 9.0
  • Bash 7.2
  • SQL 9.9
9.2 8.6 10.0 9.0 9.2 9.8 9.1 9.7 8.7 7.2 4.8 10.0 152/160 133K ↓42K ↑91K 21.3 25 min 59 s 19 ago 2026
gemma3:12b ollama 12.2B Q4_K_M 7.6 GB 131K ctx juez: gemma4:31b Vision 8.7 9.4
  • Núcleo 9.7
  • Hermes 9.2
  • OpenClaw 9.6
9.8
  • Python 9.7
  • SQL 10.0
7.5 9.4 9.4 7.8 8.3 9.6 8.9 8.1 7.7 8.3 120/128 70K ↓32K ↑39K 22.6 33 min 12 s 26 may 2026
mistral-small3.2:latest ollama 24.0B Q4_K_M 14.1 GB 131K ctx juez: gemma4:31b Vision Tools 8.7 9.8
  • Núcleo 9.4
  • Hermes 10.0
  • OpenClaw 9.8
9.4
  • Python 9.3
  • SQL 10.0
8.3 8.5 9.0 8.0 8.6 10.0 8.8 7.2 7.2 9.3 118/128 ⚠2 105K ↓74K ↑31K 9.4 1 h 8 min 59 s 26 may 2026
mistral-small3.2:24b-instruct-2506-q8_0 ollama 24.0B Q8_0 24.1 GB 131K ctx juez: gemma4:31b Vision Tools 8.6 9.5
  • Núcleo 9.4
  • Hermes 9.4
  • OpenClaw 7.2
9.1
  • Python 8.9
  • SQL 10.0
8.7 8.5 7.8 9.0 8.4 9.6 8.7 7.5 7.5 9.3 119/128 ⚠1 106K ↓75K ↑31K 7.4 1 h 20 min 6 s 26 may 2026
milkey/Seed-OSS-36B-Instruct:q4_K_M ollama 36.2B Q4_K_M 20.3 GB 524K ctx juez: gemma4:31b Tools Thinking 8.4 9.7
  • Núcleo 9.6
  • Hermes 9.9
  • OpenClaw 9.2
9.5
  • Python 9.4
  • SQL 10.0
8.4 9.1 8.2 10.0 7.9 9.3 8.0 7.2 5.2 111/128 ⚠1 83K ↓40K ↑43K 6.9 2 h 22 min 3 s 26 may 2026
deepseek-r1:32b ollama 32.8B Q4_K_M 18.5 GB 131K ctx juez: gemma4:31b Thinking 8.3 9.4
  • Núcleo 10.0
  • Hermes 9.1
  • OpenClaw 9.2
10.0
  • Python 10.0
  • SQL 10.0
8.0 8.7 9.9 8.0 7.8 8.2 7.2 7.8 6.1 112/128 59K ↓29K ↑30K 10.2 57 min 6 s 26 may 2026
mistral-small3.2:24b-instruct-2506-fp16 ollama 24.0B F16 44.7 GB 131K ctx juez: gemma4:31b-it-q8_0 Vision Tools 8.1 9.8
  • Núcleo 9.6
  • Hermes 9.9
  • OpenClaw 9.7
9.2
  • Python 9.0
  • SQL 10.0
8.8 8.7 7.8 10.0 8.7 9.6 9.0 7.4 7.2 9.6 0.0 127/136 ⚠1 114K ↓80K ↑34K 4.8 2 h 26 min 5 s 27 may 2026
gpt-oss:20b ollama 20.9B MXFP4 12.8 GB 131K ctx juez: gemma4:31b Tools Thinking 7.8 9.6
  • Núcleo 9.0
  • Hermes 9.9
  • OpenClaw 9.8
9.7
  • Python 9.6
  • SQL 10.0
7.9 7.8 10.0 8.5 6.8 7.8 7.8 4.5 5.3 111/128 ⚠1 147K ↓33K ↑114K 39.8 1 h 48 s 26 may 2026
qwen2.5vl:7b ollama 8.3B Q4_K_M 5.6 GB 128K ctx juez: gemma4:31b Vision 7.7 8.9
  • Núcleo 9.2
  • Hermes 8.5
  • OpenClaw 9.2
9.2
  • Python 9.0
  • SQL 10.0
6.8 8.7 9.3 8.8 8.5 6.9 6.5 5.8 6.7 6.7 120/128 78K ↓33K ↑44K 31.3 44 min 16 s 26 may 2026
aya-expanse:8b ollama 8.0B Q4_K_M 4.7 GB 8K ctx juez: gemma4:31b Tools 7.4 9.0
  • Núcleo 9.3
  • Hermes 9.3
  • OpenClaw 8.3
8.8
  • Python 8.7
  • SQL 10.0
6.2 8.6 8.0 4.5 7.6 8.1 7.2 7.9 5.6 112/128 57K ↓33K ↑24K 26.9 19 min 42 s 26 may 2026
Mismo banco para todos los modelos

Dónde se ejecutan los benchmarks

  • CPU AMD Ryzen AI Max+ 395 16 núcleos · 32 hilos · hasta 5,19 GHz · Zen 5
  • iGPU AMD Radeon 8060S RDNA 3.5 · 40 CUs · integrada (Strix Halo)
  • VRAM 96 GB memoria unificada asignada a la iGPU
  • RAM sistema 32 GB LPDDR5X · resto compartido con la iGPU
  • Disco NVMe 2 TB PCIe Gen4 · modelos cacheados localmente
  • Software Ollama 0.23 · Ubuntu 24.04 kernel 6.17 · driver amdgpu · ROCm runtime

Es un GMK (local) fanless con SoC AMD Strix Halo: CPU e iGPU comparten un pool de 128 GB de memoria unificada (96 GB reservados para la iGPU). Eso permite cargar modelos de 30 B+ en cuantización Q4 enteramente en VRAM, sin partir entre RAM y GPU. Todos los modelos del ranking corren en exactamente esta misma máquina.

Catálogo de pruebas

Ver las 160 pruebas en detalle →

Cómo funciona

1 · Mismas pruebas para todos

Hay 7 suites (escritura, código, frontend, agentic, reasoning, instruction, roleplay) con ~30 tests entre todas. Cada prueba lleva su prompt y su rúbrica de evaluación. Cualquier modelo nuevo se enfrenta al mismo set.

2 · Auto-checks + juez local

Lo que se puede verificar a máquina (sintaxis Python, JSON parseable, HTML válido, contiene función X) se ejecuta como chequeo automático. El resto lo puntúa otro modelo local que actúa de juez con la rúbrica delante.

3 · 100% local, reproducible

Modelos en Ollama. Juez en Ollama. Sin APIs cloud. Cada ejecución guarda el prompt exacto, la respuesta del modelo, los tokens/s y el veredicto del juez. Re-ejecutable.

4 · Cuando sale un modelo nuevo

Un solo comando lo enfrenta a toda la suite y publica los resultados aquí. Los modelos antiguos siguen comparables.