AI Stays Local

モデル

各モデルにできること(計測したハードウェア上で)

動くこと、確認されていること、使えること、入手できることは、4つの別々の事実です。各カードはそれらを分けて示し、根拠を得たハードウェアと、その根拠を出したのが誰かを添えています。各モデルの権利は公開元に帰属します。

バッジの意味

バッジは製品データ自体の対応付けルールから決まり、パラメーター数や速度から決まることはありません。プロ向けハードウェアとは計測した場所を示すもので、モデルの評価ではありません。

モデル名、計測値、製品データ自体の文は、正規の製品データに記録されているとおり英語で表示しています。

対話可能
AI Stays Local が実行・検証し、計測したハードウェア上で人が作業できる速さがあります。
動作
AI Stays Local が実行し、出力が得られます。完全には検証しておらず、対話的に使えるとは主張しません。
研究デモ
物理的に何が可能かを知るために、AI Stays Local が実行・確認しました。作業に使うには遅すぎ、提供していません。
外部デモ
外部プロジェクトによる報告です。AI Stays Local は再現しておらず、自分たちの成果として示すことはありません。

AI Stays Local が実行したモデル

OLMoE 1B/7B

対話可能

Validated on the recorded professional hardware profile.

7B total1B active per token
実行
稼働
検証
検証済み
使いやすさ
対話可能
ハードウェア
Professional ハードウェアプロファイル
提供状況
承認済みプロファイル
根拠の提供元
AI Stays Local
公開元
Allen Institute for AI
ライセンス
Apache-2.0
ソース
allenai/OLMoE-1B-7B-0125-Instruct
固定したリビジョン
b89a7c4bc24fb9e55ce2543c9458ce0ca5c4650e
量子化
int8 merged experts (upstream conversion tool)
ディスク上のチェックポイント
8 GB
テストしたコンテキスト長
2048 トークン

確認した内容Cited-document workflow validated: retrieval, citation of file and location, and refusal when the indexed documents do not contain the answer.

制限
  • Small active model: one billion parameters participate in a token. Instruction following is inconsistent on difficult prompts.
  • No professional acceptance suite has been run, so no professional workflow claim is made.
  • The local API answers stream=true with buffered transport, so the first token arrives with the whole answer.

出所Weights by Allen Institute for AI under Apache-2.0. Converted to int8 by the upstream Colibri conversion tool. Neither the weights nor the conversion are work of AI Stays Local.

計測、ハードウェア、方法

Qwen3.5 122B-A10B

対話可能

Validated on the recorded professional hardware profile.

122B total10B active per token (256 experts, top-k 8, 48 layers)
実行
稼働
検証
検証済み
使いやすさ
対話可能
ハードウェア
Professional ハードウェアプロファイル
提供状況
非公開の検証のみ
根拠の提供元
AI Stays Local
公開元
Alibaba Qwen
ライセンス
Apache-2.0
ソース
Qwen/Qwen3.5-122B-A10B-GPTQ-Int4
固定したリビジョン
30cd92cba9707a9aba09d1e490ed4b66b78e9606
量子化
GPTQ INT4, group size 128, symmetric; attention, shared experts, embeddings, lm_head and the vision tower are not quantised
ディスク上のチェックポイント
79 GB
テストしたコンテキスト長
8192 トークン

確認した内容Thirteen of thirteen functional checks, six of six product API checks and the document gates pass, including citation and refusal.

制限
  • No professional acceptance suite has been run, so no professional workflow claim is made.
  • Requires the machine to itself: peak unified memory 90.2 GiB.
  • No distributable profile exists. Validation was private.

出所Weights and quantisation published by Alibaba Qwen under Apache-2.0. Served by vLLM (Apache-2.0). Neither is work of AI Stays Local; the validation is.

計測、ハードウェア、方法

Qwen3.6 35B-A3B

動作

一部検証済み。根拠を得た環境:professional ハードウェアプロファイル。

35B total3B active per token (256 experts, top-k 8)
実行
生成可
検証
一部検証済み
使いやすさ
バッチ
ハードウェア
Professional ハードウェアプロファイル
提供状況
提供なし
根拠の提供元
AI Stays Local
公開元
Alibaba Qwen (weights); third-party conversion
ライセンス
Apache-2.0
ソース
Kreuzzelg/qwen36-35b-a3b-colibri-i4-gs64
固定したリビジョン
c619aa594ad1e70af82168fb6b4878427896e21c
量子化
int4 packed, group size 64
ディスク上のチェックポイント
24 GB
テストしたコンテキスト長
2048 トークン

確認した内容Four of four instruction-following checks pass. Document gates were not run against this model.

制限
  • Below the interactive threshold at 3.68 tokens per second.
  • The engine's OpenAI endpoint leaks a stop token into the content.
  • Document-correctness gates were not run.

出所Third-party conversion of Qwen weights for the Colibri engine. Neither the weights nor the conversion are work of AI Stays Local. Revision c619aa594ad1e70af82168fb6b4878427896e21c established on 2026-09-21 by querying the Hugging Face model API for the repository and reading the sha field, which is the immutable commit of the default branch at that time. The licence field of the same response reported apache-2.0. The benchmark was taken before the revision was recorded, so the artifact relationship is asserted rather than proven byte for byte.

計測、ハードウェア、方法

GLM-5.2 744B

研究デモ

Validated on the recorded professional hardware profile.

744B total~40B active per token (256 experts, top-k 8 plus one shared, across 75 of 78 layers)
実行
稼働
検証
検証済み
使いやすさ
研究用のみ
ハードウェア
Professional ハードウェアプロファイル
提供状況
提供なし
根拠の提供元
AI Stays Local
公開元
Zhipu AI (weights); third-party conversion by mastouri
ライセンス
MIT
ソース
mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp
固定したリビジョン
6bbb01ed3e515a8730b694dfae73aadfd6774581
量子化
int4 grouped scales gs=64, int8 embeddings and lm_head, int8 MTP head
ディスク上のチェックポイント
429 GB
テストしたコンテキスト長
2048 トークン

確認した内容Ten of ten functional checks and eight of eight document checks pass, including resistance to a prompt injection planted in an indexed file.

制限
  • 0.17 tokens per second. A 200-token answer takes about 20 minutes.
  • First token at 1018 seconds on a 311-token document prompt.
  • Execution and validation are proven. Usability is research only.

出所Weights by Zhipu AI under MIT. Container converted by a third party (mastouri) under MIT, declaring zai-org/GLM-5.2 as its base model. Engine is upstream Colibri under Apache-2.0. None of the three is work of AI Stays Local; the validation is.

計測、ハードウェア、方法

用途別

クラス別のオープンウェイトモデル。

効率的なローカルアシスタント

日常的なマシン向けの、動作するパラメーターが少ないモデル。これまでの計測はプロ向けハードウェアのみで、16 GB のプロファイルは計測していません。

プロ向けのオープンモデル

必要なメモリとストレージを備えたマシン向けの、より大きな Mixture-of-Experts モデル。

フロンティア級のオープンモデル

記録された根拠によれば、大容量メモリのマシン1台が実用的な速度で動かせる、最も大きなオープンモデル。

研究規模のデモ

物理的に何が可能かについての根拠です。製品の機能ではなく、提供していません。

私たちの根拠ではありません

外部プロジェクトの報告で、ここでは再現していないもの

境界を見えるようにするために載せています。このセクションのものは、どれも AI Stays Local が実行したものではありません。

Trillion-parameter MoE families

外部デモ

Produced by the upstream project and not reproduced here.

Above one trillion, depending on familyA small fraction per token; the remainder is streamed
実行
未テスト
検証
未検証
使いやすさ
利用不可
ハードウェア
ハードウェアについての主張なし
提供状況
提供なし
根拠の提供元
外部プロジェクト
公開元
Various
ライセンス
Apache-2.0 (engine); weights vary by publisher
ソース
github.com/JustVugg/colibri
固定したリビジョン
dcd73832f293750086643e1f0ccd2cd6d067259c
量子化
Varies
制限
  • Not reproduced by AI Stays Local. No first-party measurement exists.
  • No hardware claim is made: no reliable source has been verified.

出所Execution paths described by the upstream Colibri project. AI Stays Local has not run these families.

自社による計測はありません。

業務利用の受け入れ基準

プロフェッショナルのバッジには、次の 17 項目すべてを満たす必要があります。スループットやパラメーター数は基準ではなく、どの項目の代わりにもなりません。 現在この基準を満たすモデルはないため、プロフェッショナルのバッジが付いたモデルはありません。

  • 安定したモデルの読み込み
  • 一定時間内の起動
  • ワークフローに適した応答時間
  • 繰り返しの推論
  • キャンセル
  • キャンセル後の回復
  • API の信頼性
  • コンテキストの保持
  • 指示への追従
  • 言語の品質
  • 構造化出力
  • 文書の引用
  • 回答の拒否
  • プロセスが残らないこと
  • 許容できるメモリのふるまい
  • ハードウェアの文書化
  • 制限の文書化

製品データのリリース v0.3.1-product-layer、ペイロードのダイジェスト 11797253abc9097a.

Available through a controlled private beta. Larger model profiles are added as they complete hardware validation.