LLM · embedding · VLMThe planner does not ask directly for “Qwen” or
“Claude”: it asks for a role such as wise or
fast.micro. The instance configuration decides which provider,
model, and endpoint will fulfil that role. This allows a model to change without
rewriting the component that uses it.
The Settings > System > Models page shows both
the configured binding and, when the provider supports it, the name of the model
that the service says it is actually running. For example, a tier may be configured
with provider=llamacpp and model=Qwen….gguf; the
Observed model panel lets you compare that value with the identity
returned by the endpoint.
/admin/virt.The page opens in viewing mode. Reload configuration reads the files and renders the page again; it does not save changes, restart models, or force a fresh provider observation. LLM and VLM are edited separately. Embedding remains view-only and is never rewritten from the UI.
This is administrative configuration for the Metnos instance run by that system account. It is not a personal preference of an individual conversation user.
A component declares the kind of work it needs to perform, not the product
that must perform it. Resolution goes through three public interfaces in
runtime/virt:
| Family | Roles | Interface | Result |
|---|---|---|---|
| language model | fast.micro, fast.procedural, fast.fidelity, middle, wise, creative, frontier |
virt.get_llm(role, level=...) |
a language-model provider resolved by the tier router |
| embedding | text, image |
virt.get_embedder(role) |
a provider that produces vectors for the requested modality |
| vision-language | default |
virt.get_vlm(role) |
the effective VLM specification, not an already loaded model |
Tier names express the kind of work being requested; they do not guarantee
the quality of the concrete model. The administrator decides which provider
and model to assign to each role. A partial LLM file overrides only the sections
it contains: omitted fast, middle, and wise
sections retain their defaults. The three fast levels normally share
its binding, but may be customised in [fast.level.micro],
[fast.level.procedural], and [fast.level.fidelity].
When creative is absent, it uses the same provider and model as
wise while retaining its own generation parameters.
frontier is optional: a component that requests it must handle its
possible unavailability.
from virt import get_embedder, get_llm, get_vlm
text_encoder = get_embedder("text")
planner = get_llm("wise")
translator = get_llm("fast", level="fidelity")
vision_spec = get_vlm("default")
Virtualization therefore separates two decisions: the component selects a role; the instance configuration selects the provider and model. Changing a binding among providers already supported by the runtime does not require changes to executors or the planner. Adding a provider type unknown to the runtime still requires an implementation.
For each family, the page presents:
think,
temperature, and reasoning_budget when present;The page shows the effective generation parameters obtained by combining the configuration file with the defaults. An individual operation selects a tier but does not redefine these parameters. It may still impose its own limits, such as maximum response length, available time, and the schema that the output must follow.
Passwords, tokens, keys, credentials, and sensitive URL parts are not shown. They are also excluded from fields submitted by the form. Secrets already present in the document are preserved during a save, but they are managed through the protected credential flow or the service environment, not from this page.
Configured and observed are not synonyms. The first value is what Metnos will request from the provider; the second is what that provider reports through its API. If an endpoint advertises several identities, the page lists them and marks the result as ambiguous. If it is unreachable, does not support the check, or the observation is old, the status says so explicitly. This observation helps reveal configuration drift, but it does not change routing and is not cryptographic proof of the loaded model.
A separate service task checks the endpoints once a minute by default; after three minutes without a refresh, an observation is marked as stale. Opening the page does not instantiate providers, load models, or make network calls: it builds a bounded, secret-free view of the resolved configuration and the latest observations held in memory.
For LLM and VLM, after you press Edit, Metnos makes editable only simple values whose type it can preserve. Sensitive fields and values that cannot be edited remain read-only. On save, the runtime:
0600 permissions;An invalid configuration does not replace the previous file. After a valid save, no reinstall, recompilation, or server restart is required: subsequent calls resolve the new values. A turn already in progress finishes with the resources it has already acquired.
Restore defaults operates on LLM or VLM and asks for confirmation. It writes the initial values supplied by the installed Metnos version. It does not automatically select an old personal configuration or reconstruct choices made during an earlier installation.
Before restoring, Metnos retains a private copy of the current file when the
file exists. Copies made by saves and restores are stored below
$METNOS_USER_STATE/virt-config-history/<family>/; with default
paths, the root is ~/.local/state/metnos/virt-config-history/. The
page does not expose a historical-version picker: recovering one specific copy
is a separate administration operation.
The page always shows the effective path, which is more reliable than a path remembered from another installation. Unless overridden, the documents belong to the system account that runs Metnos.
| Family | File-resolution order |
|---|---|
| LLM | METNOS_LLM_TIERS_CONFIG; then
$METNOS_USER_CONFIG/llm_tiers.toml when it exists; finally the
legacy <install_root>/workspace/.config/llm_tiers.toml |
| embedding | METNOS_EMBEDDING_TIERS_CONFIG; otherwise
$METNOS_USER_CONFIG/embedding_tiers.toml |
| VLM | METNOS_VLM_TIERS_CONFIG; otherwise
$METNOS_USER_CONFIG/vlm_tiers.toml |
$METNOS_USER_CONFIG is normally
~/.config/metnos. If a file does not exist, the runtime uses the
initial values from the installed version and the page says so. If a document
cannot be read or fails validation, the page reports an invalid state and
distinguishes any fallback values from a valid configuration.
The installer reads the optional
~/.config/metnos/services.toml of the user who starts it. An
absent file provisions services locally; a partial file reuses only the
listed endpoints and provisions every other component locally. The file
does not discover servers: inside a container, localhost
means the container itself.
Supported sections are llm, vlm,
searxng, photon and playwright; BGE-M3
remains local. Endpoint checks precede downloads and startup. An unavailable
selected server stops installation without a local copy or cloud fallback.
The profile keeps frontier disabled unless explicitly enabled. A protocol
check does not replace testing a real chat turn or image.
The six-phase path needs administrative access and uses a separate Metnos
account. It prepares components without starting them, verifies the
distribution and initial catalog, then activates the declared system units.
The profile is part of the signed distribution: external services keep their
own management and receive no local units. After phase 2 the profile remains
unchanged during resumption; changing it on an installed instance requires
an administrative update and a restart window. Completed-phase records do
not replace authority checks. This path does not support --skip.
The service catalog also covers the managed installer entrypoint. On first preparation, the store owner creates missing directories after checking their path; symbolic links are not accepted as store directories.
Updates authenticate the original installation inventory. A new component that never belonged to that inventory does not require retiring a file from the development checkout. Existing retirement evidence, service protections and checks for conflicting processes remain mandatory.
The service registry and Tutor observe the profile's verified endpoints; external services offer no local start or stop commands. In a container, the memory check also accounts for the container's assigned limits. The initial Tutor compilation also respects the available CPU quota. Final HTTP verification requires an operational system, not merely a response from the maintenance page.
Status: six-phase integration is under test and has not been published. Examples and instructions: existing-service configuration manual.
| Family | Operational boundary |
|---|---|
| LLM | The router creates the provider declared by the tier. Supported bindings may target a local endpoint or a remote service; credentials and service availability remain separate requirements. |
| embedding | The defaults use local, in-process providers. Changing the backend is a
dedicated administrative migration, not a page control; read-only executors that declare local
computation use get_local_embedder() instead and do not acquire
network authority because of that setting. |
| VLM | get_vlm() returns the configuration. The service starts only
when needed: ensure_vlm_up() checks /health, may attempt the
configured VLM script once per process, and returns false if the
service does not become available. The caller chooses the fallback. |
Virtualization centralizes backend selection; it does not turn a remote endpoint into a local resource, nor does it grant an executor network access, credentials, or additional capabilities by itself.
Production components normally select a canonically named workload, such as
intent.extract or translation.i18n. The central registry
maps that name to a tier and, for fast, to one of its three levels.
An unknown name is rejected. A component may set limits for the individual
operation, such as max_tokens, a time limit, or an output grammar;
it does not select a provider, model, or endpoint.
| Tier | Current use | Why |
|---|---|---|
fast.micro | Intent recognition, short Tutor decisions, small summaries, and result reranking. | Short, structured answers. |
fast.procedural | Optional NLU extractor, disabled in the default installation. | Grammar-constrained JSON extraction. |
fast.fidelity | No workload in the current registry. | The role remains available and configurable without implying that it is already in use. |
middle | Structured classification and extraction, bills, optional Vaglio judgement, telos-fit assessment, manifest normalisation, and URL reranking. | Procedural transformations with a precise result. |
wise | Translation, Tutor answers, extended descriptions, planning, Synt, and durable image workloads. | Extended context, source fidelity, and comparison among alternatives. |
creative | Telos proposals, editorial commentary, manifest refactoring, and Synt descriptions. | Produces alternatives that can be compared. |
frontier | Explicitly requested external consultation. | Intentionally uses the most capable model configured for this role. |
The NLU extractor is the only current path that selects
fast.procedural directly instead of going through the workload
registry. It is disabled by default and can be enabled in observation mode or
as a replacement for the traditional analysis. The Models page still displays
every configurable role, including roles to which no workload is currently
assigned.
A tier may point to a local or remote service and may be changed by an administrator. Changing it changes subsequent calls in every row that uses it; the person configuring tiers must therefore assess capability, cost, privacy, and availability.
LRE uses explicit contract requirements. Provider selection and service lifecycle belong to model virtualization. Jobs do not choose programs to launch, ports or weight files. The admitted model and policy remain verifiable after service startup.
The LRE/Virt boundary exposes resource claims, model identity and readiness. The local vision adapter reuses the existing shared startup path; other providers retain their existing lifecycles. Automatic replica scaling has not been implemented yet.
Multiple replicas require a stable logical endpoint, request routing and resource reservations shared across jobs. Replicating a model on the same CPU is useful only when measured capacity improves; RAM, shared GPU memory and other processes must inform that decision. Service readiness alone cannot guarantee protection from memory exhaustion.
Model readiness, handoff to LRE and workload completion are distinct facts. The domain explicitly reports whether checks required action; LRE shows Exempt only after clean completion with no changes confirmed by every final stage. An existing blocked job found by nightly maintenance produces a partial outcome. See activity outcomes.
runtime/virt/__init__.py — facades and initial embedding
and VLM values;runtime/llm_router.py — LLM tiers, bindings, aliases, and
inference policy;runtime/llm_workloads.py — declarative registry mapping
each production workload to one logical tier;runtime/virt/configuration.py — effective projection,
provenance, and secret redaction;runtime/model_identity.py — bounded, secret-free
observation of the models reported by providers;runtime/virt/config_editor.py — validated editing,
atomic writes, recovery copies, and restore behavior.For the full web-chat map, see the interface navigation guide. To understand how executable components consume these roles, continue with the executor guide or return to the Architecture guide.