AI Stays Local

Evidence

What we measured, on hardware we own

Every performance claim on this site links here. A record exists when a run happened on named hardware with a named model and a named runtime — and where we do not have a figure, the row is absent rather than filled with a placeholder.

A record appears here when a run happened on named hardware with a named model and a named runtime. Where we do not have a figure, the row is absent rather than filled with a placeholder. What we measured is kept apart from what an upstream project reported — different section, different styling, its own qualification.

Measured by AI Stays Local

Qwen3.6-35B-A3B — CPU inference

Measured by AI Stays Local
10.762 sCold start7.215 sTime to first token3.68 tok/sGeneration (steady state)31.7 GiBResident set size
4 / 4Instruction following
Date
2026-09
Hardware
NVIDIA DGX Spark
Operating system
Linux, ARM64
Execution
CPU inference, 16 threads
Model
Qwen3.6-35B-A3B
Parameters
35B total, 3B active per token
Expert handling
Resident expert cache
Runtime
AI Stays Local runtime, Colibrì engine
Endpoint
Loopback only; no deployment retained after the run
Result
A 35B-total Mixture-of-Experts model executed on CPU on a single machine we own, with 3B active per token and the frequently used experts held resident.
Limitations
This throughput suits validation and patient work — batch jobs, long documents, questions you are willing to wait for. It is not an instant consumer chat experience, and we are not going to describe it as one. GPU acceleration for this class of model is in validation.

OLMoE 7B-A1B — local inference baseline

Measured by AI Stays Local
OperationalLocal inferenceOperationalOpenAI-compatible APIOperationalStreaming responses
Date
2026-09
Hardware
NVIDIA DGX Spark
Operating system
Linux, ARM64
Execution
CPU inference
Model
OLMoE
Parameters
7B total, 1B active per token
Runtime
AI Stays Local runtime
Endpoint
Loopback only
Result
The baseline profile. Local inference, the compatible API and streaming all work against a small MoE model on the same machine.
Limitations
Timing figures for this profile are not published here yet; the functional results above are what this record establishes. A small model is good at everyday drafting and document questions and will lose to a large hosted model on hard reasoning.

Local documents — coverage and refusal

Measured by AI Stays Local
PASSCovered question — answered with citationPASSUncovered question — correctly refused
Date
2026-09
Hardware
NVIDIA DGX Spark
Operating system
Linux, ARM64
Retrieval
BM25 lexical, local index
Model
Qwen3.6-35B-A3B
Runtime
AI Stays Local runtime, Colibrì engine
Corpus
Cedar internal document set
Result
Asked something the corpus covers, it answered and cited the source. Asked something the corpus does not cover, it declined rather than inventing an answer. The second result is the one that matters.
Limitations
Retrieval is lexical: BM25 matches words, not meaning, so a question phrased in vocabulary the document does not use can miss. Semantic retrieval is on the roadmap. The local index is not encrypted at rest today.
Not our measurement

Upstream research reference

Demonstrated by the upstream Colibrì project and not reproduced by AI Stays Local. It requires substantial local storage, and reported generation is far slower than interactive cloud chat.

GLM-class model, approaching 744B total — expert streaming

Upstream research reference
Date
As reported upstream
Run by
The upstream Colibrì project
Parameters
~744B total; a small active fraction per token
Expert handling
Streamed from fast local storage
Storage
Substantial — a model of this size is its own storage requirement
Result
A model approaching 744B total parameters was executed locally by streaming experts from fast storage. It is evidence that the approach scales, which is why it is on this page.
Limitations
Not reproduced by AI Stays Local, on any hardware, at any time. Reported generation is far slower than interactive cloud chat. This is upstream research evidence about what is physically possible — not a capability of this product, and not something you can do today with AI Stays Local.

This result belongs to the upstream project that produced it. It appears here as evidence that expert streaming scales, and for no other reason. AI Stays Local has not reproduced it.

The model registry these records belong to

Available through a controlled private beta. Larger model profiles are added as they complete hardware validation.