AI Stays Local

النماذج

ما يفعله كل نموذج، على العتاد الذي قيس عليه

أن يعمل النموذج، وأن يُتحقق منه، وأن يكون قابلًا للاستخدام، وأن يكون متاحًا: أربع حقائق مختلفة. تذكر كل بطاقة كلًا منها على حدة، مع العتاد الذي أُخذت عليه الأدلة ومن قدّمها. كل نموذج ملك لناشره.

ما تعنيه الشارات

تأتي الشارة من قواعد الربط الخاصة ببيانات المنتج نفسها، ولا تأتي أبدًا من عدد المعاملات أو من السرعة. العتاد الاحترافي يصف مكان أخذ القياس، وليس حكمًا على النموذج.

أسماء النماذج والقياسات والجمل الواردة في بيانات المنتج نفسها معروضة كما تسجلها بيانات المنتج المرجعية تمامًا، بالإنجليزية.

تفاعلي
نفّذه AI Stays Local وتحقق منه، وهو سريع بما يكفي ليعمل معه شخص على العتاد الذي قيس عليه.
يعمل
نفّذه AI Stays Local وهو يُنتج مخرجات. لم يُتحقق منه بالكامل، ولا يُدَّعى أنه صالح للاستخدام التفاعلي.
عرض بحثي
نفّذه AI Stays Local وتحقق منه لمعرفة ما هو ممكن فعليًا. بطيء جدًا للعمل به، وغير معروض.
عرض من مشروع مصدري
أبلغ عنه مشروع مصدري. لم يكرّره AI Stays Local، ولا يُعرض أبدًا على أنه من عملنا.

نفّذها AI Stays Local

OLMoE 1B/7B

تفاعلي

Validated on the recorded professional hardware profile.

7B total1B active per token
التشغيل
قيد التشغيل
التحقق
مُتحقق منه
قابلية الاستخدام
تفاعلي
العتاد
فئة العتاد Professional
التوفر
فئة معتمدة
الأدلة من
AI Stays Local
الناشر
Allen Institute for AI
الترخيص
Apache-2.0
المصدر
allenai/OLMoE-1B-7B-0125-Instruct
المراجعة المثبَّتة
b89a7c4bc24fb9e55ce2543c9458ce0ca5c4650e
التكميم
int8 merged experts (upstream conversion tool)
نقطة الحفظ على القرص
8 GB
السياق المختبَر
2048 رمز

ما جرى التحقق منهCited-document workflow validated: retrieval, citation of file and location, and refusal when the indexed documents do not contain the answer.

القيود
  • Small active model: one billion parameters participate in a token. Instruction following is inconsistent on difficult prompts.
  • No professional acceptance suite has been run, so no professional workflow claim is made.
  • The local API answers stream=true with buffered transport, so the first token arrives with the whole answer.

المنشأWeights by Allen Institute for AI under Apache-2.0. Converted to int8 by the upstream Colibri conversion tool. Neither the weights nor the conversion are work of AI Stays Local.

القياس والعتاد والمنهجية

Qwen3.5 122B-A10B

تفاعلي

Validated on the recorded professional hardware profile.

122B total10B active per token (256 experts, top-k 8, 48 layers)
التشغيل
قيد التشغيل
التحقق
مُتحقق منه
قابلية الاستخدام
تفاعلي
العتاد
فئة العتاد Professional
التوفر
تحقق خاص فقط
الأدلة من
AI Stays Local
الناشر
Alibaba Qwen
الترخيص
Apache-2.0
المصدر
Qwen/Qwen3.5-122B-A10B-GPTQ-Int4
المراجعة المثبَّتة
30cd92cba9707a9aba09d1e490ed4b66b78e9606
التكميم
GPTQ INT4, group size 128, symmetric; attention, shared experts, embeddings, lm_head and the vision tower are not quantised
نقطة الحفظ على القرص
79 GB
السياق المختبَر
8192 رمز

ما جرى التحقق منهThirteen of thirteen functional checks, six of six product API checks and the document gates pass, including citation and refusal.

القيود
  • No professional acceptance suite has been run, so no professional workflow claim is made.
  • Requires the machine to itself: peak unified memory 90.2 GiB.
  • No distributable profile exists. Validation was private.

المنشأWeights and quantisation published by Alibaba Qwen under Apache-2.0. Served by vLLM (Apache-2.0). Neither is work of AI Stays Local; the validation is.

القياس والعتاد والمنهجية

Qwen3.6 35B-A3B

يعمل

مُتحقق منه جزئيًا. أُخذت الأدلة على: فئة العتاد professional.

35B total3B active per token (256 experts, top-k 8)
التشغيل
يُنتج
التحقق
مُتحقق منه جزئيًا
قابلية الاستخدام
معالجة دفعية
العتاد
فئة العتاد Professional
التوفر
غير متاح
الأدلة من
AI Stays Local
الناشر
Alibaba Qwen (weights); third-party conversion
الترخيص
Apache-2.0
المصدر
Kreuzzelg/qwen36-35b-a3b-colibri-i4-gs64
المراجعة المثبَّتة
c619aa594ad1e70af82168fb6b4878427896e21c
التكميم
int4 packed, group size 64
نقطة الحفظ على القرص
24 GB
السياق المختبَر
2048 رمز

ما جرى التحقق منهFour of four instruction-following checks pass. Document gates were not run against this model.

القيود
  • Below the interactive threshold at 3.68 tokens per second.
  • The engine's OpenAI endpoint leaks a stop token into the content.
  • Document-correctness gates were not run.

المنشأThird-party conversion of Qwen weights for the Colibri engine. Neither the weights nor the conversion are work of AI Stays Local. Revision c619aa594ad1e70af82168fb6b4878427896e21c established on 2026-09-21 by querying the Hugging Face model API for the repository and reading the sha field, which is the immutable commit of the default branch at that time. The licence field of the same response reported apache-2.0. The benchmark was taken before the revision was recorded, so the artifact relationship is asserted rather than proven byte for byte.

القياس والعتاد والمنهجية

GLM-5.2 744B

عرض بحثي

Validated on the recorded professional hardware profile.

744B total~40B active per token (256 experts, top-k 8 plus one shared, across 75 of 78 layers)
التشغيل
قيد التشغيل
التحقق
مُتحقق منه
قابلية الاستخدام
للبحث فقط
العتاد
فئة العتاد Professional
التوفر
غير متاح
الأدلة من
AI Stays Local
الناشر
Zhipu AI (weights); third-party conversion by mastouri
الترخيص
MIT
المصدر
mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp
المراجعة المثبَّتة
6bbb01ed3e515a8730b694dfae73aadfd6774581
التكميم
int4 grouped scales gs=64, int8 embeddings and lm_head, int8 MTP head
نقطة الحفظ على القرص
429 GB
السياق المختبَر
2048 رمز

ما جرى التحقق منهTen of ten functional checks and eight of eight document checks pass, including resistance to a prompt injection planted in an indexed file.

القيود
  • 0.17 tokens per second. A 200-token answer takes about 20 minutes.
  • First token at 1018 seconds on a 311-token document prompt.
  • Execution and validation are proven. Usability is research only.

المنشأWeights by Zhipu AI under MIT. Container converted by a third party (mastouri) under MIT, declaring zai-org/GLM-5.2 as its base model. Engine is upstream Colibri under Apache-2.0. None of the three is work of AI Stays Local; the validation is.

القياس والعتاد والمنهجية

بحسب الغرض

نماذج مفتوحة الأوزان مجمَّعة بحسب الفئة.

مساعدون محليون فعّالون

نماذج بعدد قليل من المعاملات النشطة، مخصصة للأجهزة اليومية. قيست حتى الآن على عتاد احترافي فقط، ولم تُقَس فئة 16 GB.

نماذج مفتوحة احترافية

نماذج Mixture-of-Experts أكبر، للأجهزة التي تملك الذاكرة والتخزين اللازمين لها.

نماذج مفتوحة من الفئة الطليعية

أكبر النماذج المفتوحة التي يستطيع جهاز واحد بذاكرة كبيرة تشغيلها بسرعة قابلة للاستخدام، وفق الأدلة المسجلة.

عروض على نطاق بحثي

أدلة على ما هو ممكن فعليًا. ليست قدرة من قدرات المنتج، وغير معروضة.

ليست أدلتنا

أبلغت عنها مشاريع مصدرية، ولم نكرّرها هنا

مدرجة ليكون الحد واضحًا. لم يشغّل AI Stays Local أي شيء في هذا القسم.

Trillion-parameter MoE families

عرض من مشروع مصدري

Produced by the upstream project and not reproduced here.

Above one trillion, depending on familyA small fraction per token; the remainder is streamed
التشغيل
لم يُختبر
التحقق
غير مُتحقق منه
قابلية الاستخدام
غير متاح
العتاد
لا ادعاء بشأن العتاد
التوفر
غير متاح
الأدلة من
مشروع مصدري
الناشر
Various
الترخيص
Apache-2.0 (engine); weights vary by publisher
المصدر
github.com/JustVugg/colibri
المراجعة المثبَّتة
dcd73832f293750086643e1f0ccd2cd6d067259c
التكميم
Varies
القيود
  • Not reproduced by AI Stays Local. No first-party measurement exists.
  • No hardware claim is made: no reliable source has been verified.

المنشأExecution paths described by the upstream Colibri project. AI Stays Local has not run these families.

لا يوجد قياس من جهتنا.

معايير القبول الاحترافي

تتطلب شارة احترافي اجتياز كل واحد من هذه المعايير الـ 17. الإنتاجية وعدد المعاملات ليسا معيارين ولا يغنيان عن أي معيار. لا يستوفي أي نموذج هذه المعايير اليوم، لذلك لا يحمل أي نموذج شارة احترافي.

  • تحميل مستقر للنموذج
  • بدء تشغيل في زمن محدود
  • زمن استجابة ملائم لسير العمل
  • استدلال متكرر
  • الإلغاء
  • التعافي بعد الإلغاء
  • موثوقية الـ API
  • الاحتفاظ بالسياق
  • اتباع التعليمات
  • جودة اللغة
  • مخرجات منظمة
  • الاستشهاد بالمستندات
  • الامتناع عن الإجابة
  • عدم بقاء عمليات عالقة
  • سلوك ذاكرة مقبول
  • عتاد موثَّق
  • قيود موثَّقة

إصدار بيانات المنتج v0.3.1-product-layer، بصمة المحتوى 11797253abc9097a.

Available through a controlled private beta. Larger model profiles are added as they complete hardware validation.