Qué significan las insignias
Una insignia procede de las reglas de asignación de los propios datos de producto, nunca de un número de parámetros ni de una velocidad. Hardware profesional describe dónde se tomó una medición. No es un veredicto sobre el modelo.
Los nombres de los modelos, las mediciones y las frases de los propios datos de producto se muestran exactamente como los registran los datos canónicos del producto, en inglés.
- Interactivo
- Ejecutado y validado por AI Stays Local, y lo bastante rápido para que una persona trabaje con él en el hardware en que se midió.
- Se ejecuta
- Ejecutado por AI Stays Local y produce resultados. No está validado del todo, y no se afirma que sea utilizable de forma interactiva.
- Demostración de investigación
- Ejecutado y comprobado por AI Stays Local para saber qué es físicamente posible. Demasiado lento para trabajar con él, y no se ofrece.
- Demostración de terceros
- Informado por un proyecto de terceros. No reproducido por AI Stays Local, y nunca presentado como nuestro.
OLMoE 1B/7B
InteractivoValidated on the recorded professional hardware profile.
7B total1B active per token
- Ejecución
- Operativo
- Validación
- Validado
- Usabilidad
- Interactivo
- Hardware
- Perfil de hardware Professional
- Disponibilidad
- Perfil aprobado
- Evidencia de
- AI Stays Local
- Editor
- Allen Institute for AI
- Licencia
- Apache-2.0
- Fuente
allenai/OLMoE-1B-7B-0125-Instruct- Revisión fijada
b89a7c4bc24fb9e55ce2543c9458ce0ca5c4650e- Cuantización
- int8 merged experts (upstream conversion tool)
- Checkpoint en disco
- 8 GB
- Contexto probado
- 2048 tokens
Qué se comprobóCited-document workflow validated: retrieval, citation of file and location, and refusal when the indexed documents do not contain the answer.
Limitaciones- Small active model: one billion parameters participate in a token. Instruction following is inconsistent on difficult prompts.
- No professional acceptance suite has been run, so no professional workflow claim is made.
- The local API answers stream=true with buffered transport, so the first token arrives with the whole answer.
ProcedenciaWeights by Allen Institute for AI under Apache-2.0. Converted to int8 by the upstream Colibri conversion tool. Neither the weights nor the conversion are work of AI Stays Local.
Medición, hardware y método
Qwen3.5 122B-A10B
InteractivoValidated on the recorded professional hardware profile.
122B total10B active per token (256 experts, top-k 8, 48 layers)
- Ejecución
- Operativo
- Validación
- Validado
- Usabilidad
- Interactivo
- Hardware
- Perfil de hardware Professional
- Disponibilidad
- Solo validación privada
- Evidencia de
- AI Stays Local
- Editor
- Alibaba Qwen
- Licencia
- Apache-2.0
- Fuente
Qwen/Qwen3.5-122B-A10B-GPTQ-Int4- Revisión fijada
30cd92cba9707a9aba09d1e490ed4b66b78e9606- Cuantización
- GPTQ INT4, group size 128, symmetric; attention, shared experts, embeddings, lm_head and the vision tower are not quantised
- Checkpoint en disco
- 79 GB
- Contexto probado
- 8192 tokens
Qué se comprobóThirteen of thirteen functional checks, six of six product API checks and the document gates pass, including citation and refusal.
Limitaciones- No professional acceptance suite has been run, so no professional workflow claim is made.
- Requires the machine to itself: peak unified memory 90.2 GiB.
- No distributable profile exists. Validation was private.
ProcedenciaWeights and quantisation published by Alibaba Qwen under Apache-2.0. Served by vLLM (Apache-2.0). Neither is work of AI Stays Local; the validation is.
Medición, hardware y método
Qwen3.6 35B-A3B
Se ejecutaValidado en parte. Evidencia tomada en el perfil de hardware professional.
35B total3B active per token (256 experts, top-k 8)
- Ejecución
- Genera
- Validación
- Validado en parte
- Usabilidad
- Por lotes
- Hardware
- Perfil de hardware Professional
- Disponibilidad
- No disponible
- Evidencia de
- AI Stays Local
- Editor
- Alibaba Qwen (weights); third-party conversion
- Licencia
- Apache-2.0
- Fuente
Kreuzzelg/qwen36-35b-a3b-colibri-i4-gs64- Revisión fijada
c619aa594ad1e70af82168fb6b4878427896e21c- Cuantización
- int4 packed, group size 64
- Checkpoint en disco
- 24 GB
- Contexto probado
- 2048 tokens
Qué se comprobóFour of four instruction-following checks pass. Document gates were not run against this model.
Limitaciones- Below the interactive threshold at 3.68 tokens per second.
- The engine's OpenAI endpoint leaks a stop token into the content.
- Document-correctness gates were not run.
ProcedenciaThird-party conversion of Qwen weights for the Colibri engine. Neither the weights nor the conversion are work of AI Stays Local. Revision c619aa594ad1e70af82168fb6b4878427896e21c established on 2026-09-21 by querying the Hugging Face model API for the repository and reading the sha field, which is the immutable commit of the default branch at that time. The licence field of the same response reported apache-2.0. The benchmark was taken before the revision was recorded, so the artifact relationship is asserted rather than proven byte for byte.
Medición, hardware y método
GLM-5.2 744B
Demostración de investigaciónValidated on the recorded professional hardware profile.
744B total~40B active per token (256 experts, top-k 8 plus one shared, across 75 of 78 layers)
- Ejecución
- Operativo
- Validación
- Validado
- Usabilidad
- Solo investigación
- Hardware
- Perfil de hardware Professional
- Disponibilidad
- No disponible
- Evidencia de
- AI Stays Local
- Editor
- Zhipu AI (weights); third-party conversion by mastouri
- Licencia
- MIT
- Fuente
mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp- Revisión fijada
6bbb01ed3e515a8730b694dfae73aadfd6774581- Cuantización
- int4 grouped scales gs=64, int8 embeddings and lm_head, int8 MTP head
- Checkpoint en disco
- 429 GB
- Contexto probado
- 2048 tokens
Qué se comprobóTen of ten functional checks and eight of eight document checks pass, including resistance to a prompt injection planted in an indexed file.
Limitaciones- 0.17 tokens per second. A 200-token answer takes about 20 minutes.
- First token at 1018 seconds on a 311-token document prompt.
- Execution and validation are proven. Usability is research only.
ProcedenciaWeights by Zhipu AI under MIT. Container converted by a third party (mastouri) under MIT, declaring zai-org/GLM-5.2 as its base model. Engine is upstream Colibri under Apache-2.0. None of the three is work of AI Stays Local; the validation is.
Medición, hardware y método