AI Stays Local

Modelle

Was jedes Modell leistet, auf der Hardware, auf der es gemessen wurde

Laufen, geprüft sein, nutzbar sein und verfügbar sein sind vier verschiedene Tatsachen. Jede Karte nennt sie getrennt, zusammen mit der Hardware, auf der die Nachweise entstanden, und wer sie erbracht hat. Jedes Modell gehört seinem Herausgeber.

Was die Abzeichen bedeuten

Ein Abzeichen ergibt sich aus den Zuordnungsregeln der Produktdaten selbst, nie aus einer Parameterzahl oder einer Geschwindigkeit. Professionelle Hardware beschreibt, wo eine Messung stattfand. Sie ist kein Urteil über das Modell.

Modellnamen, Messwerte und die eigenen Sätze der Produktdaten werden genau so gezeigt, wie die kanonischen Produktdaten sie erfassen, auf Englisch.

Interaktiv
Von AI Stays Local ausgeführt und validiert, und schnell genug, dass ein Mensch damit arbeiten kann, auf der Hardware, auf der es gemessen wurde.
Läuft
Von AI Stays Local ausgeführt und liefert Ausgaben. Nicht vollständig validiert und nicht als interaktiv nutzbar ausgewiesen.
Forschungsdemonstration
Von AI Stays Local ausgeführt und geprüft, um zu lernen, was physikalisch möglich ist. Zu langsam zum Arbeiten und nicht angeboten.
Upstream-Demonstration
Von einem Upstream-Projekt berichtet. Nicht von AI Stays Local nachvollzogen und nie als unser Ergebnis dargestellt.

Von AI Stays Local ausgeführt

OLMoE 1B/7B

Interaktiv

Validated on the recorded professional hardware profile.

7B total1B active per token
Ausführung
Betriebsbereit
Validierung
Validiert
Nutzbarkeit
Interaktiv
Hardware
Hardwareprofil Professional
Verfügbarkeit
Freigegebenes Profil
Nachweise von
AI Stays Local
Herausgeber
Allen Institute for AI
Lizenz
Apache-2.0
Quelle
allenai/OLMoE-1B-7B-0125-Instruct
Festgelegte Revision
b89a7c4bc24fb9e55ce2543c9458ce0ca5c4650e
Quantisierung
int8 merged experts (upstream conversion tool)
Checkpoint auf dem Datenträger
8 GB
Getesteter Kontext
2048 Tokens

Was geprüft wurdeCited-document workflow validated: retrieval, citation of file and location, and refusal when the indexed documents do not contain the answer.

Grenzen
  • Small active model: one billion parameters participate in a token. Instruction following is inconsistent on difficult prompts.
  • No professional acceptance suite has been run, so no professional workflow claim is made.
  • The local API answers stream=true with buffered transport, so the first token arrives with the whole answer.

HerkunftWeights by Allen Institute for AI under Apache-2.0. Converted to int8 by the upstream Colibri conversion tool. Neither the weights nor the conversion are work of AI Stays Local.

Messung, Hardware und Methode

Qwen3.5 122B-A10B

Interaktiv

Validated on the recorded professional hardware profile.

122B total10B active per token (256 experts, top-k 8, 48 layers)
Ausführung
Betriebsbereit
Validierung
Validiert
Nutzbarkeit
Interaktiv
Hardware
Hardwareprofil Professional
Verfügbarkeit
Nur private Validierung
Nachweise von
AI Stays Local
Herausgeber
Alibaba Qwen
Lizenz
Apache-2.0
Quelle
Qwen/Qwen3.5-122B-A10B-GPTQ-Int4
Festgelegte Revision
30cd92cba9707a9aba09d1e490ed4b66b78e9606
Quantisierung
GPTQ INT4, group size 128, symmetric; attention, shared experts, embeddings, lm_head and the vision tower are not quantised
Checkpoint auf dem Datenträger
79 GB
Getesteter Kontext
8192 Tokens

Was geprüft wurdeThirteen of thirteen functional checks, six of six product API checks and the document gates pass, including citation and refusal.

Grenzen
  • No professional acceptance suite has been run, so no professional workflow claim is made.
  • Requires the machine to itself: peak unified memory 90.2 GiB.
  • No distributable profile exists. Validation was private.

HerkunftWeights and quantisation published by Alibaba Qwen under Apache-2.0. Served by vLLM (Apache-2.0). Neither is work of AI Stays Local; the validation is.

Messung, Hardware und Methode

Qwen3.6 35B-A3B

Läuft

Teilweise validiert. Nachweise erhoben auf: hardwareprofil professional.

35B total3B active per token (256 experts, top-k 8)
Ausführung
Erzeugt Ausgabe
Validierung
Teilweise validiert
Nutzbarkeit
Stapelbetrieb
Hardware
Hardwareprofil Professional
Verfügbarkeit
Nicht verfügbar
Nachweise von
AI Stays Local
Herausgeber
Alibaba Qwen (weights); third-party conversion
Lizenz
Apache-2.0
Quelle
Kreuzzelg/qwen36-35b-a3b-colibri-i4-gs64
Festgelegte Revision
c619aa594ad1e70af82168fb6b4878427896e21c
Quantisierung
int4 packed, group size 64
Checkpoint auf dem Datenträger
24 GB
Getesteter Kontext
2048 Tokens

Was geprüft wurdeFour of four instruction-following checks pass. Document gates were not run against this model.

Grenzen
  • Below the interactive threshold at 3.68 tokens per second.
  • The engine's OpenAI endpoint leaks a stop token into the content.
  • Document-correctness gates were not run.

HerkunftThird-party conversion of Qwen weights for the Colibri engine. Neither the weights nor the conversion are work of AI Stays Local. Revision c619aa594ad1e70af82168fb6b4878427896e21c established on 2026-09-21 by querying the Hugging Face model API for the repository and reading the sha field, which is the immutable commit of the default branch at that time. The licence field of the same response reported apache-2.0. The benchmark was taken before the revision was recorded, so the artifact relationship is asserted rather than proven byte for byte.

Messung, Hardware und Methode

GLM-5.2 744B

Forschungsdemonstration

Validated on the recorded professional hardware profile.

744B total~40B active per token (256 experts, top-k 8 plus one shared, across 75 of 78 layers)
Ausführung
Betriebsbereit
Validierung
Validiert
Nutzbarkeit
Nur Forschung
Hardware
Hardwareprofil Professional
Verfügbarkeit
Nicht verfügbar
Nachweise von
AI Stays Local
Herausgeber
Zhipu AI (weights); third-party conversion by mastouri
Lizenz
MIT
Quelle
mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp
Festgelegte Revision
6bbb01ed3e515a8730b694dfae73aadfd6774581
Quantisierung
int4 grouped scales gs=64, int8 embeddings and lm_head, int8 MTP head
Checkpoint auf dem Datenträger
429 GB
Getesteter Kontext
2048 Tokens

Was geprüft wurdeTen of ten functional checks and eight of eight document checks pass, including resistance to a prompt injection planted in an indexed file.

Grenzen
  • 0.17 tokens per second. A 200-token answer takes about 20 minutes.
  • First token at 1018 seconds on a 311-token document prompt.
  • Execution and validation are proven. Usability is research only.

HerkunftWeights by Zhipu AI under MIT. Container converted by a third party (mastouri) under MIT, declaring zai-org/GLM-5.2 as its base model. Engine is upstream Colibri under Apache-2.0. None of the three is work of AI Stays Local; the validation is.

Messung, Hardware und Methode

Nach Einsatzzweck

Open-Weight-Modelle, nach Klasse gruppiert.

Effiziente lokale Assistenten

Modelle mit wenigen aktiven Parametern für alltägliche Rechner. Bisher nur auf professioneller Hardware gemessen; das Profil mit 16 GB wurde nicht gemessen.

Professionelle offene Modelle

Größere Mixture-of-Experts-Modelle für Rechner mit dem nötigen Speicher und Speicherplatz.

Offene Modelle der Spitzenklasse

Die größten offenen Modelle, die ein Rechner mit viel Speicher nach den erfassten Nachweisen mit nutzbarer Geschwindigkeit ausführen kann.

Demonstrationen im Forschungsmaßstab

Nachweise dessen, was physikalisch möglich ist. Keine Fähigkeit des Produkts und nicht angeboten.

Nicht unsere Nachweise

Upstream berichtet, hier nicht nachvollzogen

Aufgeführt, damit die Grenze sichtbar ist. Nichts in diesem Abschnitt wurde von AI Stays Local ausgeführt.

Trillion-parameter MoE families

Upstream-Demonstration

Produced by the upstream project and not reproduced here.

Above one trillion, depending on familyA small fraction per token; the remainder is streamed
Ausführung
Nicht getestet
Validierung
Nicht geprüft
Nutzbarkeit
Nicht verfügbar
Hardware
Keine Hardwareaussage
Verfügbarkeit
Nicht verfügbar
Nachweise von
Upstream-Projekt
Herausgeber
Various
Lizenz
Apache-2.0 (engine); weights vary by publisher
Quelle
github.com/JustVugg/colibri
Festgelegte Revision
dcd73832f293750086643e1f0ccd2cd6d067259c
Quantisierung
Varies
Grenzen
  • Not reproduced by AI Stays Local. No first-party measurement exists.
  • No hardware claim is made: no reliable source has been verified.

HerkunftExecution paths described by the upstream Colibri project. AI Stays Local has not run these families.

Es gibt keine eigene Messung.

Der professionelle Abnahmevertrag

Ein Abzeichen Professionell verlangt jedes dieser 17 Kriterien. Durchsatz und Parameterzahl sind keine Kriterien und können keines ersetzen. Heute erfüllt kein Modell den Vertrag, daher trägt kein Modell das Abzeichen Professionell.

  • stabiles Laden des Modells
  • begrenzte Startzeit
  • dem Ablauf angemessene Latenz
  • wiederholte Inferenz
  • Abbruch
  • Erholung nach Abbruch
  • Zuverlässigkeit der API
  • Kontexterhalt
  • Befolgen von Anweisungen
  • Sprachqualität
  • strukturierte Ausgabe
  • Quellenangaben aus Dokumenten
  • Ablehnung
  • keine verwaisten Prozesse
  • akzeptables Speicherverhalten
  • dokumentierte Hardware
  • dokumentierte Grenzen

Produktdaten-Release v0.3.1-product-layer, Prüfsumme des Inhalts 11797253abc9097a.

Available through a controlled private beta. Larger model profiles are added as they complete hardware validation.