Models

Model catalog

The open models people run locally, organised by family, version and size. For each one: the video memory it needs, and measured speeds when it has been benchmarked.

11 families · 45 versions · 112 sizes · 6 measured

Qwen

Alibaba

The most varied family: from 0.6B for a laptop to large MoE models, with coding and reasoning variants.

Qwen3.8Qwen3.6Qwen3.5Qwen3 Coder+5

9 versions · 39 sizes · 2 sizes measured

Meta Llama

Meta

The historic reference for open models: Llama 3.1 8B is the reference model of the TokenBench protocol.

Llama 4Llama 3.3Llama 3.2Llama 3.1

4 versions · 8 sizes · 2 sizes measured

Google Gemma

Google

Google's open models, strong at multilingual tasks; the “E” versions are designed for small machines.

Gemma 4Gemma 3nGemma 3Gemma 2

4 versions · 15 sizes

DeepSeek

DeepSeek

Famous for R1 and its “distilled” versions: R1's reasoning transferred into smaller Qwen and Llama models.

DeepSeek V4 FlashR1 0528 (Qwen3)R1 / V3R1 Distill+1

5 versions · 10 sizes · 1 size measured

Mistral

Mistral AI

The French champion: efficient models, including Devstral for coding and Ministral for small setups.

Devstral Small 2Ministral 3Magistral SmallMistral Small 3.2+3

7 versions · 10 sizes

Microsoft Phi

Microsoft

“Small” models trained on carefully curated data: a lot of reasoning for their size.

Phi-4 reasoningPhi-4Phi-3.5

3 versions · 6 sizes · 1 size measured

gpt-oss

OpenAI

OpenAI's open models, MoE models shipped directly in 4-bit (MXFP4).

gpt-oss

1 version · 2 sizes

GLM

Z.ai

Z.ai (formerly Zhipu) models, known for coding and agents; the “Flash” versions target local use.

GLM-5.3 FlashGLM-4.7 FlashGLM-4.5 AirGLM-4 0414

4 versions · 5 sizes

NVIDIA Nemotron

NVIDIA

NVIDIA's models, often hybrid (Mamba + attention): a much lighter context cache.

Nemotron 3.5 LightningNemotron 3Nemotron Nano v2

3 versions · 5 sizes

IBM Granite

IBM

IBM's models, Apache 2.0 licensed and built for business use; Granite 4 is hybrid.

Granite 4.1Granite 4.0 HGranite 3.3

3 versions · 8 sizes

Petits modèles

Liquid AI · Hugging Face

Sub-10B models built to run anywhere, even without a dedicated graphics card.

LFM2.5SmolLM3

2 versions · 4 sizes

Specs and files come from Hugging Face (official repos and public GGUFs). Speeds only come from real measurements.