AI Stays Local

Models

What we have actually run

A model is supported when AI Stays Local has executed it and accepted the result. An adapter existing in an upstream engine is not support; it is an adapter. Every entry states total and active parameters separately, because they are different numbers.

Supported

Executed by AI Stays Local on our own hardware and accepted. Evidence is published on the benchmarks page.

In validation

Being brought up and measured now. It is not a capability of the product, and it may not survive validation.

Planned

Intended for a future release. Nothing has been run yet.

Upstream reference

Demonstrated by an upstream project, not reproduced by AI Stays Local. Evidence that the physics works, not evidence that we can do it.

OLMoE 7B-A1B

Supported
7B total1B active per tokenMixture-of-Experts
Hardware class
lightweight
Quantisation
Recorded with the benchmark
Runtime
AI Stays Local runtime
Licence
Apache-2.0 (model weights, per the publisher)
Provenance
Executed by AI Stays Local on NVIDIA DGX Spark, Linux ARM64, CPU inference.

LimitationsA small model. Strong on everyday drafting, summarising and questions about your own documents; it will lose to a large hosted model on hard reasoning.

Qwen3.5-122B-A10B

In validation
122B total10B active per tokenMixture-of-Experts
Hardware class
professional
Quantisation
To be fixed during validation
Runtime
AI Stays Local runtime, multi-engine
Licence
Published by the model publisher; licence verified against the publisher repository before any result is published
Provenance
A published open-weight model from its publisher. AI Stays Local has not completed its own execution and acceptance, so nothing here is attributed to us.
Evidence
No published record yet

LimitationsUnder validation on GPU. It is not operational, we do not claim it, and no third-party benchmark for it is published here as ours.

GLM-class model, approaching 744B total parameters

Upstream reference
~744B totalA small active fraction per token; other experts streamed from local storageMixture-of-Experts with expert streaming
Hardware class
extreme
Quantisation
As reported upstream
Runtime
Upstream Colibrì project
Licence
Per model publisher
Provenance
Reported by the upstream Colibrì project. Not reproduced by AI Stays Local, on any hardware, at any time.

LimitationsUpstream research evidence that expert streaming scales. It needs a great deal of local storage, reported generation is far slower than interactive cloud chat, and it is not a capability of this product.

The upstream entry is included because it shows the approach is physically possible, not because it is something this product does. It has not been reproduced by AI Stays Local on any hardware.

The measured evidence behind these entries

Available through a controlled private beta. Larger model profiles are added as they complete hardware validation.