model-hoster

…

In flight

ModelWaitingServingLongest wait

Usage by key 24h

KeyRequestsPromptCompletionMeanLast used

Memory

DeviceSizeCommittedIn useFreeLoad Recent

Nodes

NodePlatformEnginesHost RAM Processor Paging Cached Held back

Resident

ModelNodeEngineStateReserved Loaded as Leases

Engine arguments

Only for weights that carry one — this pool's 27B coding entry does, and loads it unused today. Measured there at about 2.4x for roughly 950 MiB of draft context, which grows with the context length. It changes what the model writes: unlike the n-gram speculators, its output is not identical to not speculating, so it is a per-model decision rather than a speed setting. A model with no MTP block simply runs without it.
A list tries each type in turn and each machine takes the first that fits it — a Mac with room keeps f16, a card where f16 does not fit takes q8_0 — with nothing set per machine. q8_0 answered like f16 on this pool's tests and found text 120k tokens back; both quantized types decode at about half speed once the context is long, so f16 stays first where it fits.
Blank sizes the model from its weights, context and KV cache type. A figure here is used as it is for every machine and every cache type, so it has to be blank for a list of cache types to be chosen.

A vision model needs its projector named here, or the engine starts and refuses every image. Use the filename as the repository publishes it; it is added to the files this model fetches if it was not already among them.

…

Residency

never evicted, and loaded without waiting to be asked
only meaningful when kept loaded; one machine going away then means not resident
must load on the same machine as this model, which has to be resident already. Keep that one loaded, or this one has nowhere to go
higher survives pressure longer; the lowest is evicted first
seconds idle before it is unloaded on its own. 0 never unloads it

Names

other names a client may ask for, comma separated. A whole list, so removing one means deleting it here
Weights, repository, and removing this model
one pattern per line — what the inline repository fetches. Trimming makes cached copies incomplete, and the next load pulls only what is missing; * means everything