# mesa-anyjev documentation — full corpus Each page below begins with its canonical URL followed by its original Markdown, OKF frontmatter included. ---8<--- https://idss-mesa.github.io/mesa-anyjev/getting-started/configuration/ --- title: "Configuration" description: "Configure mesa-anyjev with a YAML file, MESA_ANYJEV_ environment variables or flags; precedence and the mesa-mcp fallbacks." type: Reference tags: - getting-started - configuration - environment generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Configuration Precedence, highest first: command-line flag, environment variable, YAML file (`--config FILE`), built-in default. Environment variables use the `MESA_ANYJEV_` prefix and `__` to descend into a section: `MESA_ANYJEV_BACKEND__KIND=gateway` sets `backend.kind`. Two mesa-mcp names are honoured as fallbacks so an existing `mesa-mcp/.env` keeps working: `MESA_LLM_BASE_URL` (a trailing `/v1` is stripped) and `MESA_LLM_API_KEY`. Never `source` that file in a shell: it contains a hyphenated key; use `export MESA_LLM_API_KEY=$(grep '^MESA_LLM_API_KEY=' ../mesa-mcp/.env | cut -d= -f2-)`. The sections and every field with its default are in `config.yaml.example` and `.env.example` in the repository; the model is `mesa_anyjev.config.Config`. ## Sections | Section | What it configures | |---|---| | `backend` | The AnyJev backend: `fake`, `gateway` (carc-fast, always `logprobs=20`, `max_choice_k=8`), `hf` (local weights), `composite`. | | `decider` | Level request (`auto`), prior, adaptive shifts. | | `planner` | The reasoning model: `gateway` (carc-tools, default), `claude`, `static` (fallback). | | `policy` | Profile (`prod` needs L1 for auto-writes), thresholds file, hosted-provider switch. | | `artifacts`, `provenance`, `ducklake`, `ols`, `motherduck` | Bundle directory, sidecar DSN, local DuckLake catalog, OLS base URL and fixtures, hosted Jev settings. | | `eval_root` | The neon-avu-eval checkout for silver labels and the bench (no default). | ---8<--- https://idss-mesa.github.io/mesa-anyjev/getting-started/gateway/ --- title: "Gateway" description: "Decide with carc-fast on the CARC LiteLLM gateway at L0: the tunnel, the key, what the doctor checks, and what the gateway cannot do." type: Guide tags: - getting-started - gateway - carc generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Gateway The gateway is LiteLLM in front of vLLM on the CARC hosts (RESEARCH.md). It is reached only over a loopback SSH tunnel: `127.0.0.1:8000` (the `carc-litellm-tunnel` user unit) or `127.0.0.1:18000` (mesa-mcp's `scripts/llm_tunnel.sh`). Never `source` mesa-mcp's `.env`; take the key out of it: ```bash export MESA_LLM_API_KEY=$(grep '^MESA_LLM_API_KEY=' ../mesa-mcp/.env | cut -d= -f2-) uv run mesa-anyjev doctor --backend gateway ``` The doctor checks liveness, that the served tokenizer loads at the pinned revision, that the answer labels are single tokens, that a two-option probe rendered through the chat template comes back with both labels in the top-20, and that `logprobs=26` is refused (the cap is 20). Then (`auto` replays recorded OLS responses and records the pairs a real model searches for the first time): ```bash MESA_ANYJEV_BACKEND__KIND=gateway MESA_ANYJEV_OLS__FIXTURES=auto uv run mesa-anyjev annotate \ --level L0 --card tests/fixtures/cards/DP1.10022.001.bet_sorting.md \ --provenance duckdb:///$PWD/.local/prov.duckdb --out .local/gateway_L0.json ``` What the gateway cannot do (DESIGN D11): it drops `allowed_token_ids`, so every request asks for the top-20 logprobs and a label outside them is floored and counted as missing; a batch with a missing label never auto-writes. It serves no hidden states, so L2 needs local weights (milestone M3). Choices with more than eight options use their yes/no twin. ---8<--- https://idss-mesa.github.io/mesa-anyjev/getting-started/install/ --- title: "Install" description: "Install mesa-anyjev with uv, choose the extras for the gateway, local weights, Claude or Postgres, and run the doctor." type: Guide tags: - getting-started - install generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Install ```bash git clone https://github.com/idss-mesa/mesa-anyjev cd mesa-anyjev uv sync --all-extras --no-extra hf # everything except local CUDA weights uv run mesa-anyjev doctor ``` Python 3.11 or newer. `anyjev`, `mesa-mcp` and `mesa-ducklake` come from pinned git commits (`pyproject.toml`, `[tool.uv.sources]`); none of them is on PyPI at a usable version. ## Extras | Extra | Adds | When | |---|---|---| | `gateway` | `transformers` (tokenizer only) | Deciding through the CARC LiteLLM gateway (L0, L1). | | `hf` | `torch`, `transformers`, `accelerate` | Local hidden states for L2 on a CUDA host ([Local weights](local-weights.md)). | | `claude` | `anthropic` | The Claude planner and structured-output provider (M2+). | | `pg` | `psycopg` | The Postgres provenance sidecar (production). | | `e2e` | `mcp` | The end-to-end test over MCP stdio (M4). | The gateway is reached only over a loopback SSH tunnel; see `RESEARCH.md` in the repository for the hosts, the served models and what the gateway does and does not pass through. ---8<--- https://idss-mesa.github.io/mesa-anyjev/getting-started/local-weights/ --- title: "Local weights" description: "Run Qwen/Qwen3-8B on a CUDA host (the GB10 workstation) for hidden states and L2 heads, the composite backend that pairs it with the gateway, and the torch 2.14 Triton switch." type: Guide tags: - getting-started - hf - l2 - gb10 generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Local weights The gateway reads a top-20 logprob list and has no hidden states, so it stops at L1. L2 (a closed-form head on a hidden state, cheaper than L0 because the forward pass stops at the head's block) needs the weights on the host. The `hf` extra installs torch, transformers and accelerate; on the GB10 workstation that resolves to torch 2.14 with CUDA 13 wheels for aarch64 (RESEARCH.md § Hosts). ```bash uv sync --all-extras # includes hf uv run mesa-anyjev doctor --backend hf ``` The doctor loads `backend.hf_model` (default `Qwen/Qwen3-8B`, bf16, 16.4 GB; about two minutes from the local safetensors cache, longer on the first download), reports `n_layers` and `hidden_size`, checks that the answer labels are single tokens, runs a two-state probe and requires the label tokens to carry the answer mass, and confirms the block loop for early stop is available. Then, with the neon-avu-eval labels ingested (see [Learning](../concepts/learning.md)): ```bash uv run mesa-anyjev bench run --backend hf --levels raw,L0,L1,L2 uv run mesa-anyjev learn fit --backend hf --question term.fits --level L2 uv run mesa-anyjev learn promote --backend hf --version --question term.fits uv run mesa-anyjev annotate --backend hf --level auto --card ``` `annotate` and `bench` load the promoted bundle for the configured model before deciding, so `--level auto` resolves `term.fits` at L2 when a head is promoted and every other question at L0. A bundle fitted on local weights is refused by the gateway backend and the other way round (DESIGN D5): the NVFP4 checkpoint the gateway serves and the bf16 checkpoint here are recorded as different models until the bench shows they agree. ## Composite: gateway logprobs plus local hidden states `--backend composite` builds both backends under one name. Logprobs come from the gateway (carc-fast) and hidden states from the local model, so a single process can serve L0 and L1 from the gateway's artifacts and L2 from a local head. The constructor asserts that the two tokenizers agree on every label token id (A..Z, Yes, No, 1..9) and refuses otherwise; the composite still reads the gateway's top-20 list, so `max_choice_k` stays at 8 for L0 and L1 while L2 heads read the hidden state and ignore the cap. It needs the tunnel and the key like the [gateway](gateway.md) page describes. ## torch 2.14 and Triton torch 2.14 routes some eager aten ops (the rotary-embedding outer product in Qwen3) through Triton kernels that are compiled against `Python.h` on first use. A host whose Python has no development headers (the GB10 with the system CPython 3.12) fails inside the forward pass with `fatal error: Python.h: No such file or directory`. `backend.hf_native_triton` is therefore `false` by default: the factory deregisters torch's Triton DSL overrides before loading, and the stock aten kernels run instead. Set it to `true` on a host with headers if you want the Triton path. ## What the host can do | Backend | Levels | Prompts per decision | Where | |---|---|---|---| | `gateway` | raw, L0, L1 | one prefill per cyclic shift | sparky-2 over the tunnel | | `hf` | raw, L0, L1, L2 | one prefill, stopped at the head's block for L2 | this host, CUDA | | `composite` | raw, L0, L1 (gateway) and L2 (local) | as above | both | | `fake` | all (synthetic hidden states) | none | CI | ---8<--- https://idss-mesa.github.io/mesa-anyjev/getting-started/quickstart/ --- title: "Quickstart" description: "Annotate one dataset card end to end with the fake backend, apply the proposals to a local DuckLake, and read the provenance back." type: Tutorial tags: - getting-started - quickstart generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Quickstart Everything below runs offline on the synthetic backend and a local DuckLake catalog; the recorded OLS responses under `tests/fixtures/ols/` stand in for the EMBL-EBI API. ```bash export MESA_ANYJEV_OLS__FIXTURES=replay MESA_ANYJEV_POLICY__PROFILE=dev uv run mesa-anyjev annotate --card tests/fixtures/cards/DP1.10003.001.brd_countdata.md \ --backend fake --planner static \ --provenance duckdb:///$PWD/.local/prov.duckdb --out .local/run.json ``` `annotate` decides and proposes; it writes nothing to iRODS. The run id, the proposals with their probability and level, and the neon-avu-eval result shape are in `.local/run.json`. The shipped policy is proposed-only, so the curator accepts the run: ```bash RUN=$(jq -r .run_id .local/run.json) uv run mesa-anyjev apply --run-id $RUN --accept proposed \ --provenance duckdb:///$PWD/.local/prov.duckdb \ --local-ducklake duckdb:///$PWD/.local/lake.duckdb \ --irods-path /local/mesa-anyjev/DP1.10003.001/brd_countdata.csv --actor $USER uv run mesa-anyjev explain --provenance duckdb:///$PWD/.local/prov.duckdb \ --path /local/mesa-anyjev/DP1.10003.001/brd_countdata.csv ``` `apply` writes one mesa-ducklake snapshot for the run and path with the source tag `mesa-anyjev:apply:human`; `explain` joins each AVU back to the decision that produced it, its level and its snapshot id. To decide with a real model, see [Gateway](gateway.md). ---8<--- https://idss-mesa.github.io/mesa-anyjev/concepts/datacite/ --- title: "DataCite decisions" description: "The frozen DataCite vocabularies as questions: resource type as a yes/no per member, the others as a choice with a yes/no twin, and the record fragment for mesa_avu_apply_datacite." type: Guide tags: - concepts - datacite - questions generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: datacite resource: "https://schema.datacite.org/" title: "DataCite Metadata Schema" author: "DataCite" status: draft --- # DataCite decisions mesa-mcp writes DataCite records as AVUs (`mesa_avu_apply_datacite`) and validates them against the DataCite 4.x controlled vocabularies. Choosing a value from those vocabularies is a closed choice, so M5 freezes five of them in `registry.py` (a test asserts parity with mesa-mcp's enums) and asks them like any other question: | Vocabulary | Members | Question | Shape | |---|---|---|---| | ResourceTypeGeneral | 28 | `datacite.resource_type_fits` | yes/no per member (28 exceeds the 26-option cap) | | ContributorType | 21 | `datacite.contributor_type` (+ `_fits` twin) | choice under a wide backend, twin under the gateway's cap of 8 | | RelationType | 23 | `datacite.relation_type` (+ twin) | same | | DateType | 10 | `datacite.date_type` (+ twin) | same | | DescriptionType | 6 | `datacite.description_type` | choice | The state is the text being classified (a description, a contributor line, a related identifier, a date), the vocabulary name and, for the twins, the candidate value, with the card header when a card is at hand. Levels, calibration and thresholds work as for every other question; there are no labels yet, so every threshold is `auto: null` and the answers are proposals. ```bash uv run mesa-anyjev datacite --card tests/fixtures/cards/DP1.10022.001.bet_sorting.md uv run mesa-anyjev datacite --vocabulary DescriptionType --text "We sampled beetles with pitfall traps." ``` `datacite --card` answers what a dataset card can answer on its own (the general resource type of the product and the description type of its description) and prints the `record` fragment in the shape `mesa_avu_apply_datacite` takes, with the ranked alternatives beside it. ---8<--- https://idss-mesa.github.io/mesa-anyjev/concepts/decision-graph/ --- title: "Decision graph" description: "The fixed-key questions mesa-anyjev asks from a dataset card to a written AVU, which provider answers each, and what happens below threshold." type: Guide tags: - concepts - decisions - questions generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Decision graph Every question is fixed-key (DESIGN D1): its wording and option list are frozen in `questions.py` and pinned in `questions.lock.json`; the variable content (the column, the site, the candidate term, the sibling AVUs) lives in the *state*. That is what lets one question accumulate a batch prior, an L1 temperature and an L2 head across every column and dataset. Open candidate sets are asked as one yes/no per candidate under one key and ranked by `p_true`, which also removes AnyJev's 26-option cap. | Step | Question | Kind | Decides | |---|---|---|---| | Plan | (the reasoning model) | planner | which ontologies are in play, which columns to annotate, what to search; hints only | | Q1 | `column.annotate` | noul | should this column be annotated at all (identifiers are a rule, never a model call) | | Q2 | `column.aspect` | choice, K=8 | taxon, environment, method, measurement, unit, data_type, location, other | | Q3 | `column.ontology` / `column.ontology_fits` | choice K=12, or its noul twin on the gateway | which registry ontology to search; masked by aspect and plan | | S | candidates | code | OLS search, prefix filter, dedup, cap | | Q4 | `term.fits` | noul per candidate | is this candidate the right term at the right specificity (one key for column, site, unit and taxon scopes) | | Q7 | `avu.value_kind` | choice, K=4 | term label, site code, column name, or top data value (after deterministic pre-rules) | | Q8 | `avu.keep` | noul | keep this AVU given its siblings | | H | human pick | MRTR | the curator's choice is authoritative and becomes a label | | W | write | code | one DuckLake snapshot per run and path | The registry, the aspects and the value kinds are in `mesa_anyjev.registry`; the ontology registry is the ten ontologies allowed by the neon-avu-eval prompt plus TAXRANK and GENEPIO (every prefix with at least five valid evaluation AVUs, DESIGN D7). ---8<--- https://idss-mesa.github.io/mesa-anyjev/concepts/hosted-providers/ --- title: "Hosted providers" description: "Hosted Jev through MotherDuck prompt_jev as a third provider for the same questions, and the data-residency switch that keeps it off by default." type: Guide tags: - concepts - motherduck - hosted generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Hosted providers MotherDuck's `prompt_jev` SQL function answers the same three primitives (choice, score, yes/no) and returns a choice, a full probability distribution and a confidence (DESIGN D16). mesa-anyjev asks it the identical questions: the frozen question text is the instruction, the frozen option list is the `choice := [...]` constant, and the `State:` text is the same rendering AnyJev prefills, so a hosted answer and a local answer to one state share the `state_sha256`. The SQL literal for every question is pinned in `questions.lock.json` (`hosted_sql`) beside its key, without entering the lock sha. ## What a hosted record is `provider='motherduck'`, `method='hosted:prompt_jev'`, `level='none'` (no AnyJev level exists for it), `calibration='typesafe'` with the distribution re-ordered into the frozen option order. Three things never become a probability: a `value` the frozen options do not contain, a distribution that does not sum to one, and a NULL answer (MotherDuck returns NULL for a failed request and the query continues). Those records abstain with `calibration='none'` and the reason in `diagnostics`; a NULL is recorded as `decider_unavailable`. The observed return shape (`STRUCT(choice, probabilities[], confidence)` for choice, a DOUBLE for yes/no, a weighted position for score) is pinned in the parser and in the doctor, so a change in the preview fails loudly. ## Scoring stored states ```bash export MOTHERDUCK_TOKEN=... # read by the DuckDB extension; never stored in config MESA_ANYJEV_POLICY__HOSTED_PROVIDERS=allowlist \ uv run mesa-anyjev hosted score --question term.fits --source labels --dry-run MESA_ANYJEV_POLICY__HOSTED_PROVIDERS=allowlist \ uv run mesa-anyjev hosted score --question term.fits,column.aspect --source labels ``` `--dry-run` prints the rows, requests (32 rows per request) and approximate input tokens without sending anything. A real run records a normal sidecar run with `backend_kind='motherduck'`, `data_left_host=true` and the egress region, one decision row per state, and, when the states carry labels, a results file of its own (`bench/results//motherduck__prompt_jev.motherduck.json`), never mixed with gateway or local rows. `--source run --run-id ` re-scores the decisions of a local run so the two providers can be compared state by state. Re-scoring effective AVUs straight from the DuckLake history waits for states that carry the dataset card. ## Data residency Because the state leaves the host, hosted providers are off by default (`policy.hosted_providers: off`) and the CLI, the service and the doctor refuse them. With `allowlist`, a local source (labels, fixtures, a sidecar run on this host) is allowed when `hosted_allow_local_sources` is true, and a project only when its root is listed in `hosted_allow_project_roots` or carries `mesa.hosted_inference=allow`. The `prod` profile allows `auto` for AnyJev-calibrated records only, so a hosted decision may be proposed but never auto-writes until a bench row in its own results file cites its coverage at 5% risk. `mesa-anyjev doctor` adds the MotherDuck checks (extension, token present by name only, `md:` attach, a two-option fixture, the pinned shape) whenever the policy is not `off`. ---8<--- https://idss-mesa.github.io/mesa-anyjev/concepts/learning/ --- title: "Learning" description: "Where labels come from, how L1 temperatures and L2 heads are fitted with leave-one-card-out, the guards against positives-only data, and how an artifact is promoted." type: Guide tags: - concepts - learning - calibration - labels generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Learning Labels are option indices per `(question_key, state)` in the sidecar's `labels` table, each with a source and a weight; when a state has several, the highest weight wins (DESIGN D6): | Source | Weight | Where it comes from | |---|---|---| | `curator` | 1.0 | an explicit pick or "none of these" in the term picker (M4) | | `curator_implicit` | 0.7 | the other candidates offered next to a pick | | `accepted_avu` | 0.8 | an auto-written AVU later deleted through mesa-mcp | | `consensus_all`, `consensus_majority` | 0.8, 0.6 | a term proposed by every, or by at least two, of the four agentic models in neon-avu-eval | | `consensus_negative` | 0.5 | a term proposed by a single model, or an ontology no model used for a column | | `teacher`, `hosted_jev` | 0.5 | allowed in the schema; not produced in 0.1.0 | | `gold` | 1.0 | hand-verified rows | `mesa-anyjev learn ingest --eval-root ` writes the silver from the evaluation: `term.fits` states for every unique (card, CURIE) pair with the candidate's OLS record in the state, `column.annotate`, `column.aspect`, `column.ontology` and its yes/no twin, `avu.value_kind` and `avu.keep`. The silver is *agreement between models*, not truth, and it is circular with the agentic baselines the bench compares against. ## Fitting `mesa-anyjev learn fit --question term.fits --level L1` fits a temperature on the L0 probabilities. Two guards come from the plan's critique: the per-question `min_weight` in `policy_defaults.yaml` lets the single-model negatives in (at the default 0.6 the term labels are positives-only, which would make temperature scaling degenerate and coverage trivially 100%), and every leave-one-card-out fold needs at least 30 labels per class in training and 5 in the held-out card, else it is skipped and reported. Fitting runs on a Decider without adaptive shifts so the artifact freezes the prior it was fitted with. `--level L2` fits a closed-form head (LDA or ridge, chosen by the same held-out score) on the hidden state of one block, on a backend that exposes hidden states (`hf`, `composite`, or the fake backend in tests); the gateway raises `LevelUnavailable`. Both levels go through one fitting path, `learn/fit.py`, which the bench's leave-one-card-out cells now share, so a bench L1 or L2 cell and a `learn fit` report on the same labels are the same computation, and both pool the held-out decisions before computing ECE and coverage (M2's two paths disagreed on term.fits ECE, 0.072 against 0.231, because the fitter averaged per-fold values of metrics that are not linear in the items). L2 needs at least 40 labels and the same per-class fold guards; the manifest records the head's block (`layer_abs`), method and, for L1, the temperature. At serving time a head answers only its own question key (A1): a fitted `term.fits` head is never routed to another yes/no question. The fit is saved as a new immutable artifact version whose manifest carries the pooled held-out numbers; `learn promote --version N` moves `CURRENT` only when held-out accuracy and ECE do not regress against the current version (D15). Serving processes never fit. ---8<--- https://idss-mesa.github.io/mesa-anyjev/concepts/levels-and-policy/ --- title: "Levels and policy" description: "What an AnyJev level and the calibration field promise, and the rules that turn a decision into auto, proposed, human or abstain." type: Guide tags: - concepts - calibration - policy generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Levels and policy Every stored decision carries two fields (DESIGN D3): * `level`: AnyJev's level (`raw`, `L0` zero-label debiasing, `L1` temperature on labels, `L2` closed-form head on hidden states) or `none` for anything that is not an AnyJev decision: a rule, a planner hint, a Claude structured answer, a hosted Jev answer. * `calibration`: where the probabilities come from: `anyjev`, `typesafe` (hosted Jev), or `none`. `probs` is null exactly when calibration is `none`; nothing is ever one-hot. The write policy (`policy_defaults.yaml`, `mesa_anyjev.policy`) reads `p_true` for yes/no questions (never `confidence`, which is the larger of p and 1-p) and `confidence` for choices. A decision becomes: * `auto` only with an AnyJev level at or above both the question's and the profile's floor (`prod`: L1), a calibration the profile allows, a numeric threshold that cites the leave-one-card-out bench cell it came from, and no label missing from the gateway's top-20 in that batch; * `proposed` at or above the propose threshold (the curator sees a ranked picker); * `abstain` otherwise. The shipped defaults are proposed-only (`auto: null` everywhere) until a bench run exists. ---8<--- https://idss-mesa.github.io/mesa-anyjev/concepts/mcp-tools/ --- title: "MCP tools" description: "The mesa_decide_* tools inside mesa-mcp: annotate, apply with one question per candidate group, explain, feedback and health; how they register, and what the elicitation state carries." type: Guide tags: - concepts - mcp - tools - elicitation generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: mesa-mcp-plugins resource: "https://github.com/idss-mesa/mesa-mcp/pull/6" title: "mesa-mcp: load third-party tools from mesa_mcp.tools entry points" author: "team:idss-mesa" status: draft --- # MCP tools mesa-anyjev is a plugin inside mesa-mcp's registry, not a second server (DESIGN D14). The package declares a `mesa_mcp.tools` entry point (`decide = "mesa_anyjev.mcp_tools"`); once mesa-mcp's loader is merged, installing mesa-anyjev next to mesa-mcp adds five tools with the surface tag `decision` in their `_meta`. Until then, `import mesa_anyjev.mcp_tools` before `MesaServer()` registers them under the surface `core`. | Tool | Phase | What it does | |---|---|---| | `mesa_decide_annotate` | decide | Runs the decision graph on a dataset card (`card_text` or `card_path`), records every decision and proposal in the sidecar, writes nothing. | | `mesa_decide_apply` | write | Writes accepted AVUs to an iRODS path and mirrors them into the DuckLake (one snapshot per run and path). `accept="proposed"` asks the user to pick for each open candidate group first. `dry_run=True` by default. | | `mesa_decide_explain` | read | Every decision of a run with its level, calibration and probability, the AVU links with write status, the open groups; or the decisions behind the AVUs on a path. | | `mesa_decide_feedback` | curate | A pick, a reject or a decline on a candidate group; authoritative, stored as an override and as labels. | | `mesa_decide_health` | ops | Questions lock, provenance store, backend and promoted bundle (local weights are not loaded here). | Handlers receive the iRODS `auth_value` mesa-mcp injects; the authenticated user is the actor, and the write goes through mesa-mcp's own `assert_allowed` and AVU helpers with the session from its pool. Without an authenticated user the tool runs in local mode and writes only to the sidecar. ## One question per round trip `mesa_decide_apply(accept="proposed")` uses mesa-mcp's Multi Round-Trip Requests: for the first open group it raises an elicitation with key `term_choice:`, a form whose options are the group's candidate decision ids and whose names read `label (CURIE) p(fits)=0.83 L1` (an unranked group says so instead of showing a number). The request state carries ids only (`tool`, `run_id`, `asked`), never a path, a label or a probability: on resume every label, probability and level is read back from the sidecar (`candidates_for_group`), and an answer whose id was not offered, or a state that belongs to another run, is refused (DESIGN D8). A decline leaves the group unwritten. When no group is open the call writes what is accepted. ## What a pick does A pick is authoritative (outcome `human`): the chosen candidate's link becomes `accepted` (a candidate that was not the winner gets a fresh link built from the same value rule), the other proposed links are `rejected`, the group's winner is updated, an override row records who chose what from which offered list, and labels are written for the next `learn fit`: the pick as a curator positive (weight 1.0), the other offered candidates as implicit negatives (0.7), and an explicit "none of these" as curator negatives (1.0). Serving never fits (D15). ## The chooser for mesa-mcp's own picker `mesa_avu_apply_term` has an eight-candidate picker of its own. `ElicitationChooser` answers it over the real protocol with the fixed-key question `term.fits.chooser` over the only state the picker carries (ontology, value, candidate), and returns the best candidate when p(fits) clears the `propose` threshold, otherwise a decline. It never reuses `term.fits` artifacts (plan amendment B13). `tests/e2e/test_chooser_harness.py` runs it through mesa-ducklake's llm_e2e broker against a live iRODS zone. ---8<--- https://idss-mesa.github.io/mesa-anyjev/concepts/provenance/ --- title: "Provenance" description: "The mesa-anyjev sidecar: runs, decisions, options, groups, AVU links, human overrides and labels, and how it joins the mesa-ducklake AVU history." type: Reference tags: - concepts - provenance - ducklake generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Provenance Decision provenance is a sidecar owned by mesa-anyjev (DESIGN D4): Postgres schema `mesa_anyjev` (its own migrations, `src/mesa_anyjev/provenance/migrations/`) or a DuckDB file for development. It is never a column on mesa-ducklake's `avu_changes` (the AVU triple is a hard contract) and never a table in schema `mesa`. Tables: `runs`, `decisions` (one row per answered question with state, provider, level, calibration, probabilities, thresholds in force and outcome), `decision_options`, `decision_groups` (a ranking over candidates), `avu_links` (the AVU a decision produced and its write status), `human_overrides`, `labels`. The join into the AVU history is `(project_id, snapshot_id, irods_path, attribute, value, unit)`; `snapshot_id` is null for decisions and links that wrote nothing. Every AVU written by mesa-anyjev also carries a per-row `source` tag of the form `mesa-anyjev::` in mesa-ducklake, so `get_history` shows which rows were written automatically and which by a curator without a join. Decisions are never updated; only a link's write status and snapshot id, and a run's status, change. Corrections are new rows. ## Postgres and exports `postgresql://` DSNs use the packaged migration (`migrations/0001_mesa_anyjev.sql`, schema `mesa_anyjev`, JSONB, foreign keys, `NULLS NOT DISTINCT` uniqueness) through a psycopg store with the same protocol as the DuckDB file; `mesa-anyjev provenance migrate --dsn ...` applies it, and the store may share a database with mesa-ducklake's `mesa` schema (D4). The `requires_postgres` test tier runs the DuckDB round trip against a real server. `mesa-anyjev provenance export --run-id --out /.mesa/anyjev` writes `runs`, `decisions`, `decision_groups` and `avu_links` as Parquet (never under `.mesa/ducklake/`). `provenance reconcile --run-id --local-ducklake ` repairs links whose snapshot id is NULL because the mirror step failed after the iRODS write, by matching the DuckLake history rows whose source starts with `mesa-anyjev:` on the exact AVU triple. ---8<--- https://idss-mesa.github.io/mesa-anyjev/concepts/second-opinions/ --- title: "Second opinions" description: "Claude as a recorded second opinion on proposed terms (never a probability, never auto), the specificity step over child terms, and the planner audit." type: Guide tags: - concepts - claude - specificity - audit generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" status: draft --- # Second opinions Three M5 steps refine a proposal without ever raising its level. ## Claude as a recorded second opinion The Claude Messages API exposes no logits, so Claude can never be an AnyJev backend. What it can do is answer the same rendered `State` / `Question` / `Options` text through structured outputs (`messages.parse` with an answer model whose only field is a literal over the frozen options, adaptive thinking, no forced tool choice). Every such record is `provider='claude'`, `level='none'`, `calibration='none'`, `probs=None`; a refusal, an error or an answer outside the options abstains. With `claude.second_opinion` on (or `annotate --second-opinion`), a `proposed` candidate group gets Claude's yes/no over its top candidates, recorded in the same group with the winner as parent. Claude agreeing is noted on the proposal; Claude saying no to the winner turns the outcome into `escalated`, which stays a proposal for a human to settle. Nothing Claude says can make a group `auto` (the profile requires an AnyJev calibration), and nothing it says changes a probability. ## Specificity When a proposed term has children in its ontology, the pipeline asks `term.fits` over those children (a second group that records which term it refines). A child replaces the parent only when p(child) is at least `policy.specificity_delta` (0.10) above p(parent) and the child clears the proposal threshold on its own. Recorded fixtures may lack the children calls; the parent then stands. On the first end-to-end run at L0 (carc-fast) 23 child groups were opened and none replaced its parent ([Bench results](../bench/results.md)). ## Planner audit Before any column is looked at, `dataset.ontology_applies` is asked once per registry entry over the card header. The answers are recorded (they never write) and compared with the planner's ontology list: `annotate` prints the agreement, and the run's result carries `audit.planner_only` and `audit.model_only`, so a planner that keeps naming ontologies the model rejects, or missing ones it accepts, shows up in the sidecar rather than in a hunch. ---8<--- https://idss-mesa.github.io/mesa-anyjev/bench/results/ --- title: "Bench results" description: "Measured raw, L0, L1 and L2 numbers per question on local Qwen3-8B weights, the CARC gateway and the synthetic backend, every cell naming its results JSON." type: Reference tags: - bench - results - calibration generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: hf-results resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/bench/results/2026-09-25/Qwen__Qwen3-8B.hf.json" title: "bench results 2026-09-25, Qwen/Qwen3-8B bf16 on the GB10 workstation" author: "team:idss-mesa" - id: gateway-results resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/bench/results/2026-09-25/RedHatAI__Qwen3-8B-NVFP4.gateway.json" title: "bench results 2026-09-25, carc-fast through the CARC gateway" author: "team:idss-mesa" - id: fake-results resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/bench/results/2026-09-25/fake.fake.json" title: "bench results 2026-09-25, synthetic backend" author: "team:idss-mesa" status: draft stale_after: "2026-12-31T00:00:00Z" --- # Bench results Labels are the neon-avu-eval silver (agreement between four agentic models, not truth; see [Learning](../concepts/learning.md)). `L1 (LOCO)` cells are pooled held-out decisions over seven leave-one-card-out folds. `cov@5%` is the share of items that can be answered before the error rate on the answered set exceeds 5%; `flip` is the answer change under option reversal (choice) or phrasing swap (yes/no). Every number below names its JSON. ## End to end: the whole graph, twice per card (carc-fast at L0), 2026-09-25 `bench e2e` runs the decision graph on the seven fixture cards, two repetitions per planner, and reports agreement rather than asserting determinism: the Jaccard overlap of proposed AVU triples between repetitions, and the recall of the neon-avu-eval consensus terms (proposed by every agentic model, or by at least two). Source: `bench/results/2026-09-25/RedHatAI__Qwen3-8B-NVFP4.gateway.e2e.json`. | planner | rep agreement | consensus-all recall | consensus-majority recall | fallbacks | |---|---|---|---|---| | static | 0.37 | 0.25 | 0.06 | 0 / 7 | | gateway (carc-tools) | 0.18 | 0.05 | 0.05 | 1 / 7 | At L0 on carc-fast the graph is far from repeatable (per card 0.0 to 0.67) and recovers little of the agentic consensus; proposals per card were 1 to 9. The gateway planner made it worse and cost 96 minutes of wall time at ~4.7 tokens per second. The planner audit (`dataset.ontology_applies`) rejected `gaz`, `pato`, `uo` and `ro` on most cards where the static planner had named them. Specificity opened 23 child groups and replaced no parent. These are the numbers the L2 head and curator labels have to move. ## Qwen/Qwen3-8B bf16 on the GB10 workstation (local weights), 2026-09-25 The same labels through AnyJev's `HFBackend` on this host: the full vocabulary is read (no top-20 cap, so no erased labels), and L2 heads read the hidden state of one block. `L1 (LOCO)` and `L2 (LOCO)` cells are pooled held-out decisions over seven leave-one-card-out folds. The `prompts` column reads 0 because the local backend had no request counter in this run (added afterwards; the next run reports prefills). | task | n | level | acc | ece | brier | cov@5% | cov@10% | flip | n_neg | prompts | ms/dec | | neon_annotate | 98 | L0 | 0.439 | 0.474 | 0.970 | 0.000 | 0.000 | 0.143 | 35 | 0 | 817.1 | | neon_annotate | 98 | L1 (LOCO) | 0.294 | 0.529 | 0.569 | 0.000 | 0.000 | | 5 | 0 | 369.2 | | neon_annotate | 98 | L2 (LOCO) | 0.765 | 0.257 | 0.354 | 0.353 | 0.353 | | 5 | 0 | 196.5 | | neon_annotate | 98 | raw | 0.439 | 0.512 | 1.022 | 0.000 | 0.000 | | 35 | 0 | 398.8 | | neon_aspect | 60 | L0 | 0.533 | 0.456 | 0.922 | 0.000 | 0.000 | 0.150 | 0 | 0 | 1917.5 | | neon_aspect | 60 | L1 (LOCO) | 0.517 | 0.162 | 0.633 | 0.000 | 0.000 | | 0 | 0 | 8175.3 | | neon_aspect | 60 | L2 (LOCO) | 0.567 | 0.175 | 0.552 | 0.383 | 0.417 | | 0 | 0 | 2983.5 | | neon_aspect | 60 | raw | 0.550 | 0.420 | 0.841 | 0.117 | 0.117 | 0.250 | 0 | 0 | 447.3 | | neon_keep_avu | 313 | L0 | 0.949 | 0.031 | 0.085 | 0.994 | 1.000 | 0.022 | 11 | 0 | 845.2 | | neon_keep_avu | 313 | L1 | n/a | | | | | | | | | | neon_keep_avu | 313 | L2 | n/a | | | | | | | | | | neon_keep_avu | 313 | raw | 0.962 | 0.036 | 0.074 | 1.000 | 1.000 | | 11 | 0 | 744.1 | | neon_ontology_fits | 190 | L0 | 0.426 | 0.560 | 1.111 | 0.089 | 0.116 | 0.084 | 114 | 0 | 621.4 | | neon_ontology_fits | 190 | L1 (LOCO) | 0.426 | 0.233 | 0.537 | 0.089 | 0.111 | | 114 | 0 | 5381.1 | | neon_ontology_fits | 190 | L2 (LOCO) | 0.837 | 0.078 | 0.231 | 0.547 | 0.768 | | 114 | 0 | 1906.4 | | neon_ontology_fits | 190 | raw | 0.421 | 0.566 | 1.137 | 0.089 | 0.105 | | 114 | 0 | 216.5 | | neon_ontology_for_column | 59 | L0 | 0.441 | 0.508 | 1.027 | 0.169 | 0.271 | 0.186 | 0 | 0 | 3085.4 | | neon_ontology_for_column | 59 | L1 (LOCO) | 0.441 | 0.141 | 0.732 | 0.102 | 0.102 | | 0 | 0 | 6622.9 | | neon_ontology_for_column | 59 | L2 (LOCO) | 0.712 | 0.154 | 0.391 | 0.458 | 0.678 | | 0 | 0 | 1679.8 | | neon_ontology_for_column | 59 | raw | 0.492 | 0.443 | 0.896 | 0.051 | 0.051 | 0.136 | 0 | 0 | 568.5 | | neon_term_choice26 | 44 | L0 | 0.114 | 0.833 | 1.648 | 0.000 | 0.000 | 0.236 | 28 | 0 | 578.2 | | neon_term_choice26 | 44 | L1 | n/a | | | | | | | | | | neon_term_choice26 | 44 | L2 | n/a | | | | | | | | | | neon_term_choice26 | 44 | raw | 0.114 | 0.789 | 1.621 | 0.000 | 0.000 | | 28 | 0 | 191.4 | | neon_term_fits | 285 | L0 | 0.663 | 0.293 | 0.616 | 0.077 | 0.182 | 0.119 | 199 | 0 | 392.3 | | neon_term_fits | 285 | L1 (LOCO) | 0.656 | 0.079 | 0.422 | 0.049 | 0.077 | | 199 | 0 | 2765.1 | | neon_term_fits | 285 | L2 (LOCO) | 0.765 | 0.058 | 0.327 | 0.088 | 0.396 | | 199 | 0 | 1374.7 | | neon_term_fits | 285 | raw | 0.642 | 0.316 | 0.656 | 0.074 | 0.112 | | 199 | 0 | 199.1 | | neon_value_kind | 278 | L0 | 0.421 | 0.477 | 1.022 | 0.014 | 0.014 | 0.385 | 0 | 0 | 1046.0 | | neon_value_kind | 278 | L1 (LOCO) | 0.428 | 0.088 | 0.669 | 0.007 | 0.007 | | 0 | 0 | 1709.4 | | neon_value_kind | 278 | L2 (LOCO) | 0.576 | 0.069 | 0.565 | 0.029 | 0.040 | | 0 | 0 | 1275.0 | | neon_value_kind | 278 | raw | 0.406 | 0.536 | 1.127 | 0.014 | 0.014 | 0.500 | 0 | 0 | 370.4 | How to read it (`bench/results/2026-09-25/Qwen__Qwen3-8B.hf.json`): * **term.fits** (285 states, 199 negatives): accuracy 0.642 raw, 0.663 L0, 0.656 L1, 0.765 L2; ECE 0.316, 0.293, 0.079, 0.058; coverage at 5% risk 0.074, 0.077, 0.049, 0.088 and at 10% risk 0.112, 0.182, 0.077, 0.396. L2 meets the 0.10 ECE bar with the best accuracy, but only 25 of 285 items can be answered at 5% risk, so `term.fits` `auto` stays null. * **column.ontology_fits** (190, 114 negatives) is the one question where L2 changes the picture: accuracy 0.837 and ECE 0.078 with coverage 0.547 at 5% risk and 0.768 at 10%, against 0.426 / 0.233 / 0.089 at L1. The `policy_suggestions` block cites this cell. Before any threshold is adopted the number needs a second label source: these labels are the agentic models' ontology choices, so the head has learnt their habits, not the truth. * **column.annotate** (98, 5 negatives above the weight cut at L1/L2): L2 accuracy 0.765 against 0.294 at L1, but with five negatives the coverage figure (0.353) rests on almost nothing and every fold barely passes the class guard. * **column.aspect** (60) and **avu.value_kind** (278) gain accuracy at L2 (0.567, 0.576) and calibrate (ECE 0.175, 0.069) but stay near zero coverage at 5% risk. * **neon_term_choice26** at raw/L0 scores 0.114 locally against 0.227 / 0.205 on the gateway: reading the full vocabulary does not rescue a 26-way choice with a new key per item; the noul-per-candidate design is the right one for this model. * **avu.keep**: 11 negatives, no fold may fit, unchanged. * Against the gateway at L0 the local bf16 checkpoint scores term.fits 0.663 vs 0.611 and annotate 0.439 vs 0.398, with flip rates 0.119 vs 0.393 on term.fits: the NVFP4 quant and the bf16 weights are not interchangeable, which is why bundles are keyed per model (D5). ## carc-fast (RedHatAI/Qwen3-8B-NVFP4) through the CARC gateway, 2026-09-25 | task | n | level | acc | ece | brier | cov@5% | cov@10% | flip | n_neg | prompts | ms/dec | |---|---|---|---|---|---|---|---|---|---|---|---| | neon_term_fits | 285 | raw | 0.502 | 0.372 | 0.780 | 0.004 | 0.004 | | 199 | 285 | 100.6 | | neon_term_fits | 285 | L0 | 0.611 | 0.262 | 0.605 | 0.007 | 0.049 | 0.393 | 199 | 570 | 126.5 | | neon_term_fits | 285 | L1 (LOCO) | 0.621 | 0.072 | 0.459 | 0.007 | 0.042 | | 199 | 3990 | 728.9 | | neon_term_choice26 | 44 | raw | 0.227 | 0.689 | 1.443 | 0.000 | 0.000 | | 28 | 44 | 151.6 | | neon_term_choice26 | 44 | L0 | 0.205 | 0.689 | 1.460 | 0.000 | 0.000 | 0.343 | 28 | 152 | 354.5 | | neon_term_choice26 | 44 | L1 | n/a | | | | | | | | | | neon_annotate | 98 | raw | 0.378 | 0.517 | 1.062 | 0.000 | 0.000 | | 35 | 98 | 46.8 | | neon_annotate | 98 | L0 | 0.398 | 0.490 | 1.020 | 0.000 | 0.000 | 0.316 | 35 | 196 | 39.1 | | neon_annotate | 98 | L1 (LOCO) | 0.294 | 0.467 | 0.560 | 0.000 | 0.000 | | 5 | 196 | 40.4 | | neon_aspect | 60 | raw | 0.600 | 0.376 | 0.775 | 0.000 | 0.000 | 0.433 | 0 | 120 | 83.2 | | neon_aspect | 60 | L0 | 0.583 | 0.373 | 0.736 | 0.000 | 0.000 | 0.283 | 0 | 444 | 242.0 | | neon_aspect | 60 | L1 (LOCO) | 0.600 | 0.181 | 0.552 | 0.017 | 0.017 | | 0 | 1680 | 718.0 | | neon_ontology_fits | 190 | raw | 0.411 | 0.575 | 1.140 | 0.047 | 0.063 | | 114 | 190 | 47.4 | | neon_ontology_fits | 190 | L0 | 0.474 | 0.417 | 0.868 | 0.032 | 0.116 | 0.289 | 114 | 380 | 47.6 | | neon_ontology_fits | 190 | L1 (LOCO) | 0.458 | 0.182 | 0.494 | 0.068 | 0.105 | | 114 | 2660 | 310.7 | | neon_keep_avu | 313 | raw | 0.965 | 0.025 | 0.064 | 1.000 | 1.000 | | 11 | 313 | 171.9 | | neon_keep_avu | 313 | L0 | 0.895 | 0.044 | 0.160 | 0.786 | 0.974 | 0.323 | 11 | 626 | 178.4 | | neon_keep_avu | 313 | L1 | n/a | | | | | | | | | | neon_value_kind | 278 | raw | 0.507 | 0.457 | 0.937 | 0.004 | 0.004 | 0.259 | 0 | 556 | 73.0 | | neon_value_kind | 278 | L0 | 0.475 | 0.462 | 0.942 | 0.004 | 0.004 | 0.194 | 0 | 1578 | 124.2 | | neon_value_kind | 278 | L1 (LOCO) | 0.468 | 0.124 | 0.642 | 0.004 | 0.004 | | 0 | 4759 | 377.2 | How to read it (`bench/results/2026-09-25/RedHatAI__Qwen3-8B-NVFP4.gateway.json`): * **term.fits** (285 candidate states, 199 negatives): accuracy 0.502 raw, 0.611 at L0, 0.621 at L1; ECE 0.372, 0.262, 0.072. Calibration meets the pre-registered 0.10 bar at L1, but coverage at 5% risk is 0.007 (two items), so every `auto` threshold stays null. * **neon_term_choice26** is the control: one choice per column over its candidates with a new key per item, so it can only run raw/L0 and never learns (accuracy 0.23 / 0.21). * **column.aspect** (K=8) already lost 19 labels to the gateway's top-20 readout at L0; the wider the choice, the more the gateway erases (DESIGN D11). * **avu.keep** has only 11 negatives, so no leave-one-card-out fold is allowed to fit; its raw/L0 numbers reflect a positives-heavy label set, not a calibrated gate. * In M2 the bench and `learn fit` disagreed on term.fits ECE (0.072 here, 0.231 in the bundle's manifest). The cause was not the fitting: `learn fit` averaged per-fold ECE and coverage, which are not linear in the items, while the bench pooled the held-out decisions. Since M3 both share one fitting path and both pool (DESIGN D18). ## Synthetic backend (planted biases; numbers are not a claim about any model) | task | n | level | acc | ece | brier | cov@5% | cov@10% | flip | n_neg | prompts | ms/dec | |---|---|---|---|---|---|---|---|---|---|---|---| | neon_term_fits | 285 | raw | 0.442 | 0.504 | 1.017 | 0.007 | 0.007 | | 199 | 285 | 0.1 | | neon_term_fits | 285 | L0 | 0.442 | 0.488 | 0.946 | 0.007 | 0.007 | 0.004 | 199 | 570 | 0.2 | | neon_term_fits | 285 | L1 (LOCO) | 0.442 | 0.178 | 0.508 | 0.007 | 0.007 | | 199 | 3990 | 1.2 | | neon_term_choice26 | 44 | raw | 0.250 | 0.249 | 0.811 | 0.000 | 0.000 | | 28 | 44 | 0.2 | | neon_term_choice26 | 44 | L0 | 0.159 | 0.317 | 0.841 | 0.000 | 0.000 | 0.411 | 28 | 221 | 0.7 | | neon_term_choice26 | 44 | L1 | n/a | | | | | | | | | | neon_annotate | 98 | raw | 0.643 | 0.238 | 0.573 | 0.020 | 0.020 | | 35 | 98 | 0.1 | | neon_annotate | 98 | L0 | 0.622 | 0.101 | 0.473 | 0.020 | 0.020 | 0.163 | 35 | 196 | 0.2 | | neon_annotate | 98 | L1 (LOCO) | 0.647 | 0.353 | 0.457 | 0.118 | 0.118 | | 5 | 196 | 0.1 | | neon_aspect | 60 | raw | 0.200 | 0.216 | 0.964 | 0.017 | 0.017 | 0.917 | 0 | 120 | 0.3 | | neon_aspect | 60 | L0 | 0.117 | 0.156 | 0.878 | 0.000 | 0.000 | 0.000 | 0 | 861 | 1.8 | | neon_aspect | 60 | L1 (LOCO) | 0.167 | 0.164 | 0.890 | 0.000 | 0.000 | | 0 | 3074 | 6.3 | | neon_ontology_fits | 190 | raw | 0.400 | 0.480 | 0.938 | 0.000 | 0.000 | | 114 | 190 | 0.1 | | neon_ontology_fits | 190 | L0 | 0.432 | 0.167 | 0.549 | 0.000 | 0.000 | 0.189 | 114 | 380 | 0.2 | | neon_ontology_fits | 190 | L1 (LOCO) | 0.432 | 0.123 | 0.501 | 0.000 | 0.000 | | 114 | 2660 | 1.0 | | neon_keep_avu | 313 | raw | 0.965 | 0.085 | 0.083 | 1.000 | 1.000 | | 11 | 313 | 0.2 | | neon_keep_avu | 313 | L0 | 0.911 | 0.318 | 0.352 | 0.888 | 1.000 | 0.173 | 11 | 626 | 0.3 | | neon_keep_avu | 313 | L1 | n/a | | | | | | | | | | neon_value_kind | 278 | raw | 0.478 | 0.130 | 0.691 | 0.007 | 0.007 | 1.000 | 0 | 556 | 0.3 | | neon_value_kind | 278 | L0 | 0.216 | 0.112 | 0.757 | 0.004 | 0.004 | 0.007 | 0 | 2195 | 0.8 | | neon_value_kind | 278 | L1 (LOCO) | 0.227 | 0.107 | 0.750 | 0.004 | 0.004 | | 0 | 7661 | 2.4 | ---8<--- https://idss-mesa.github.io/mesa-anyjev/develop/architecture/ --- title: "Architecture" description: "The mesa-anyjev modules, the annotate and apply phases, and the boundaries with AnyJev, mesa-mcp and mesa-ducklake." type: Reference tags: - develop - architecture generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Architecture Two phases (DESIGN D10): `annotate` decides and proposes (no writes); `apply` writes the accepted AVUs to iRODS through mesa-mcp's shared helpers and records one mesa-ducklake snapshot per run and path. Decisions are written to the sidecar before any iRODS write. Layers: backends (AnyJev's protocol: fake, gateway, local weights, composite) below providers (`DecisionRecord` with level and calibration) below the pipeline; the planner is a separate role that only proposes. `mesa_mcp.ols` provides the OLS client and the canonical AVU transform; `mesa_ducklake.DuckLakeClient` records snapshots. The code map is in `CLAUDE.md`; the decisions in `DESIGN.md`; the facts in `RESEARCH.md`. ## Service and tools (M4) `service.py` holds the collaborators (provider, planner, OLS layer, policy, store) behind one lock with a bounded wait (`policy.max_wait_s`; a busy decider degrades to `decider_unavailable` instead of queueing), and owns the human-feedback path. The CLI and the `mesa_decide_*` tools (`mcp_tools/`) both build on it; `chooser.py` answers mesa-mcp's own picker. See [MCP tools](../concepts/mcp-tools.md). ---8<--- https://idss-mesa.github.io/mesa-anyjev/develop/testing/ --- title: "Testing" description: "How to run the hermetic mesa-anyjev test suite and the opt-in engine, Postgres, end-to-end and hosted tiers." type: Guide tags: - develop - testing generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # Testing ```bash uv run pytest -q # hermetic: FakeBackend, DuckDB under tmp_path uv run ruff check src tests scripts && uv run mypy --strict src ``` Opt-in tiers are excluded by default and selected with markers and environment variables: `engine` (`MESA_ANYJEV_ENGINE=gateway|hf`), `requires_postgres` (`MESA_ANYJEV_TEST_PG_*`), `e2e` (`MESA_ANYJEV_E2E=1`), `live` (`MESA_ANYJEV_LIVE=1`), `hosted` (`MESA_ANYJEV_MOTHERDUCK=1`). Unit tests never touch the network or a GPU. Every question key is pinned by `tests/test_questions_lock.py`; every numeric write threshold is checked by `tests/test_policy_citations.py`. ## The Postgres and end-to-end tiers ```bash docker run -d --rm --name mesa-anyjev-pg -e POSTGRES_PASSWORD=mesa -e POSTGRES_USER=mesa \ -e POSTGRES_DB=mesa_anyjev_test -p 127.0.0.1:55432:5432 postgres:16 MESA_ANYJEV_TEST_PG_DSN=postgresql://mesa:mesa@127.0.0.1:55432/mesa_anyjev_test \ uv run pytest -q -m requires_postgres tests/test_provenance_postgres.py ``` The `e2e` tier (`tests/e2e/test_chooser_harness.py`) spawns the real `mesa-mcp --transport stdio` with the copied elicitation broker and lets the chooser answer `mesa_avu_apply_term`'s picker. It needs `MESA_E2E_IRODS_ROOT`, iRODS credentials in the environment, the `e2e` extra and live EBI OLS; it is skipped otherwise. ---8<--- https://idss-mesa.github.io/mesa-anyjev/about/ai-agents/ --- title: "AI agents" description: "How AI agents should read this documentation bundle: llms.txt, the markdown mirror, trust markers." type: Guide tags: - about - agents - okf generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # AI agents Start from [`llms.txt`](../llms.txt) (an outline with descriptions) or [`llms-full.txt`](../llms-full.txt) (the whole corpus). Any page URL plus `index.md` returns that page's Markdown source. Pages without a `verified:` key are unverified (OKF v0.2 §5.3); `status: draft` pages need review. For decisions and facts, prefer `DESIGN.md` and `RESEARCH.md` in the repository; for behaviour, run `mesa-anyjev` and read its provenance. ---8<--- https://idss-mesa.github.io/mesa-anyjev/about/license/ --- title: "License" description: "mesa-anyjev is MIT licensed by the Regents of the University of New Mexico; AnyJev is Apache-2.0; NEON data are CC BY 4.0." type: Policy tags: - about - license generated: by: "claude/fable-5.1" at: "2026-09-25T00:00:00Z" sources: - id: design resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/DESIGN.md" title: "mesa-anyjev decisions register (DESIGN.md)" author: "team:idss-mesa" - id: research resource: "https://github.com/idss-mesa/mesa-anyjev/blob/main/RESEARCH.md" title: "mesa-anyjev verified facts (RESEARCH.md)" author: "team:idss-mesa" status: draft --- # License mesa-anyjev is released under the MIT License, Copyright (c) 2026 The Regents of the University of New Mexico. AnyJev is Apache-2.0 software by Nokia and is not affiliated with TypeSafe AI or Jev; its `bench/metrics.py` is vendored byte-identical with attribution in `THIRD_PARTY.md`. NEON data used by the bench are CC BY 4.0. See `THIRD_PARTY.md` for every third-party component. ---8<--- https://idss-mesa.github.io/mesa-anyjev/log/ # Directory Update Log ## 2026-09-25 * **Creation**: [Second opinions](concepts/second-opinions.md) and [DataCite decisions](concepts/datacite.md) for milestone M5. * **Update**: [Hosted providers](concepts/hosted-providers.md) for milestone M4b (the MotherDuck provider, `hosted score`, pinned SQL and shape, the residency gate). * **Creation**: [MCP tools](concepts/mcp-tools.md) for milestone M4; [Provenance](concepts/provenance.md) gains Postgres, export and reconcile; [Testing](develop/testing.md) the Postgres and end-to-end tiers; [Architecture](develop/architecture.md) the service layer. * **Creation**: [Local weights](getting-started/local-weights.md) for milestone M3 (the hf and composite backends, the doctor's local checks, the torch 2.14 Triton switch); [Learning](concepts/learning.md) gains the L2 head section and the unified fitting path. * **Creation**: [Bench results](bench/results.md): the first gateway and synthetic bench tables. * **Creation**: [Learning](concepts/learning.md) for milestone M2 (label sources, LOCO fitting, guards, promotion). * **Creation**: [Quickstart](getting-started/quickstart.md) and [Gateway](getting-started/gateway.md) for milestone M1 (annotate, apply, explain; the doctor's gateway checks). * **Creation**: the documentation bundle for milestone M0: [Getting started](getting-started/index.md) ([Install](getting-started/install.md), [Configuration](getting-started/configuration.md)), [Concepts](concepts/index.md) ([Decision graph](concepts/decision-graph.md), [Levels and policy](concepts/levels-and-policy.md), [Provenance](concepts/provenance.md), [Hosted providers](concepts/hosted-providers.md)), [Develop](develop/index.md) ([Architecture](develop/architecture.md), [Testing](develop/testing.md)) and [About](about/index.md) ([AI agents](about/ai-agents.md), [License](about/license.md)).