배지의 의미
배지는 제품 데이터 자체의 매핑 규칙에서 나오며, 파라미터 수나 속도에서 나오지 않습니다. 전문가용 하드웨어는 측정한 장소를 뜻할 뿐, 모델에 대한 평가가 아닙니다.
모델 이름, 측정값, 제품 데이터 자체의 문장은 정본 제품 데이터에 기록된 그대로 영어로 표시합니다.
- 대화형
- AI Stays Local이 실행하고 검증했으며, 측정한 하드웨어에서 사람이 작업할 수 있을 만큼 빠릅니다.
- 실행됨
- AI Stays Local이 실행했으며 출력이 나옵니다. 완전히 검증되지 않았으며, 대화형으로 쓸 수 있다고 주장하지 않습니다.
- 연구 시연
- 물리적으로 무엇이 가능한지 알아보기 위해 AI Stays Local이 실행하고 확인했습니다. 작업에 쓰기에는 너무 느리며, 제공하지 않습니다.
- 외부 시연
- 외부 프로젝트가 보고한 것입니다. AI Stays Local이 재현하지 않았으며, 저희 성과로 제시하지 않습니다.
OLMoE 1B/7B
대화형Validated on the recorded professional hardware profile.
7B total1B active per token
- 실행
- 운영 가능
- 검증
- 검증됨
- 사용성
- 대화형
- 하드웨어
- Professional 하드웨어 프로필
- 제공 상태
- 승인된 프로필
- 근거 제공
- AI Stays Local
- 배포자
- Allen Institute for AI
- 라이선스
- Apache-2.0
- 출처
allenai/OLMoE-1B-7B-0125-Instruct- 고정된 리비전
b89a7c4bc24fb9e55ce2543c9458ce0ca5c4650e- 양자화
- int8 merged experts (upstream conversion tool)
- 디스크의 체크포인트
- 8 GB
- 테스트한 컨텍스트
- 2048 토큰
확인한 내용Cited-document workflow validated: retrieval, citation of file and location, and refusal when the indexed documents do not contain the answer.
한계- Small active model: one billion parameters participate in a token. Instruction following is inconsistent on difficult prompts.
- No professional acceptance suite has been run, so no professional workflow claim is made.
- The local API answers stream=true with buffered transport, so the first token arrives with the whole answer.
출처 정보Weights by Allen Institute for AI under Apache-2.0. Converted to int8 by the upstream Colibri conversion tool. Neither the weights nor the conversion are work of AI Stays Local.
측정, 하드웨어, 방법
Qwen3.5 122B-A10B
대화형Validated on the recorded professional hardware profile.
122B total10B active per token (256 experts, top-k 8, 48 layers)
- 실행
- 운영 가능
- 검증
- 검증됨
- 사용성
- 대화형
- 하드웨어
- Professional 하드웨어 프로필
- 제공 상태
- 비공개 검증 전용
- 근거 제공
- AI Stays Local
- 배포자
- Alibaba Qwen
- 라이선스
- Apache-2.0
- 출처
Qwen/Qwen3.5-122B-A10B-GPTQ-Int4- 고정된 리비전
30cd92cba9707a9aba09d1e490ed4b66b78e9606- 양자화
- GPTQ INT4, group size 128, symmetric; attention, shared experts, embeddings, lm_head and the vision tower are not quantised
- 디스크의 체크포인트
- 79 GB
- 테스트한 컨텍스트
- 8192 토큰
확인한 내용Thirteen of thirteen functional checks, six of six product API checks and the document gates pass, including citation and refusal.
한계- No professional acceptance suite has been run, so no professional workflow claim is made.
- Requires the machine to itself: peak unified memory 90.2 GiB.
- No distributable profile exists. Validation was private.
출처 정보Weights and quantisation published by Alibaba Qwen under Apache-2.0. Served by vLLM (Apache-2.0). Neither is work of AI Stays Local; the validation is.
측정, 하드웨어, 방법
Qwen3.6 35B-A3B
실행됨일부 검증됨. 근거를 얻은 환경: professional 하드웨어 프로필.
35B total3B active per token (256 experts, top-k 8)
- 실행
- 생성됨
- 검증
- 일부 검증됨
- 사용성
- 일괄 처리
- 하드웨어
- Professional 하드웨어 프로필
- 제공 상태
- 제공 안 됨
- 근거 제공
- AI Stays Local
- 배포자
- Alibaba Qwen (weights); third-party conversion
- 라이선스
- Apache-2.0
- 출처
Kreuzzelg/qwen36-35b-a3b-colibri-i4-gs64- 고정된 리비전
c619aa594ad1e70af82168fb6b4878427896e21c- 양자화
- int4 packed, group size 64
- 디스크의 체크포인트
- 24 GB
- 테스트한 컨텍스트
- 2048 토큰
확인한 내용Four of four instruction-following checks pass. Document gates were not run against this model.
한계- Below the interactive threshold at 3.68 tokens per second.
- The engine's OpenAI endpoint leaks a stop token into the content.
- Document-correctness gates were not run.
출처 정보Third-party conversion of Qwen weights for the Colibri engine. Neither the weights nor the conversion are work of AI Stays Local. Revision c619aa594ad1e70af82168fb6b4878427896e21c established on 2026-09-21 by querying the Hugging Face model API for the repository and reading the sha field, which is the immutable commit of the default branch at that time. The licence field of the same response reported apache-2.0. The benchmark was taken before the revision was recorded, so the artifact relationship is asserted rather than proven byte for byte.
측정, 하드웨어, 방법
GLM-5.2 744B
연구 시연Validated on the recorded professional hardware profile.
744B total~40B active per token (256 experts, top-k 8 plus one shared, across 75 of 78 layers)
- 실행
- 운영 가능
- 검증
- 검증됨
- 사용성
- 연구 전용
- 하드웨어
- Professional 하드웨어 프로필
- 제공 상태
- 제공 안 됨
- 근거 제공
- AI Stays Local
- 배포자
- Zhipu AI (weights); third-party conversion by mastouri
- 라이선스
- MIT
- 출처
mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp- 고정된 리비전
6bbb01ed3e515a8730b694dfae73aadfd6774581- 양자화
- int4 grouped scales gs=64, int8 embeddings and lm_head, int8 MTP head
- 디스크의 체크포인트
- 429 GB
- 테스트한 컨텍스트
- 2048 토큰
확인한 내용Ten of ten functional checks and eight of eight document checks pass, including resistance to a prompt injection planted in an indexed file.
한계- 0.17 tokens per second. A 200-token answer takes about 20 minutes.
- First token at 1018 seconds on a 311-token document prompt.
- Execution and validation are proven. Usability is research only.
출처 정보Weights by Zhipu AI under MIT. Container converted by a third party (mastouri) under MIT, declaring zai-org/GLM-5.2 as its base model. Engine is upstream Colibri under Apache-2.0. None of the three is work of AI Stays Local; the validation is.
측정, 하드웨어, 방법