Data refresh

What each refresh found, and what it did not publish.

ModelTree's dataset is refreshed by agents against primary sources, reviewed by an independent three-rubric panel, and gated deterministically. Every run is recorded here in full — including the runs that published nothing.

A run's working state is never committed. This page transcribes the durable record: the pull request body and the summary issue, both linked from each entry.

Runs recorded
23
Pages fetched
2,128
Claims proposed
1,352
Edits published
611
Items withheld
439

Showing 1–4 of 4 runsfiltered by Ran, changed nothing

Clear filter
  1. 2026-08-31-b7c2d9

    Long-tail depth tranche 2026-08-31 — EleutherAI, Ai2, TII, NVIDIA, IBM, Cohere

    Scope requested: Six creators already in the catalogue, each holding exactly one family — EleutherAI, Ai2, TII, NVIDIA, IBM and Cohere — under the long-tail unanimous 3-of-3 policy of ADR 0002, since none of the six has a reviewed profile. One additional family per creator was sought, with at least one release under each, restricted to origins already approved at the merge-base. The six were taken from the issue’s own candidate table; a coordinating brief naming stability-ai in place of cohere arrived after the panel had returned its verdicts, too late to change what was scouted, so cohere is in this run and stability-ai is not. That substitution is recorded under notCovered along with a read-only probe of what stability-ai would have run into. No profile catalogue was edited, no origin was added, and no existing record was corrected: every claim was kind add, so a correction found along the way would have been a separate issue rather than folded in.Ran, changed nothing0 edits posted · 30 items withheld

    Researched a second model family for six creators that each hold exactly one, intending to take the exactly-one-family count from 30 down to 24, and published none of them. Thirty claims were proposed, all of kind add, carrying 69 evidence entries; every quote was machine-checked as a contiguous verbatim substring of the exact bytes its contentHash names, with a positive and a negative control on the harness, and all 69 passed. Origin discipline was not the blocker: every page came from an origin the repository already approves, and gate-source-approval exited 0 on all six bundles. Dates were not the blocker either, which is what separates this run from the breadth tranche the day before: each proposed family took its firstReleaseDate from a document distinct from the one dating the release it ships with, so no family date was derived from its own release. The panel was the blocker. Under the unanimous 3-of-3 long-tail bar the provenance rubric rejected all 15 family and release claims, and the reason is nearly uniform: releaseSchema and familySchema both require status, and no model card or announcement read in this run states a lifecycle term that forces one member of the enum. An announcement verb — "We are releasing the following three models on Hugging Face" — is not an availability state, and a value filled from what is usual for the field is a guess wearing a vocabulary term. The same rubric rejected accessType: open-weight wherever it rested only on a licence name, since a licence is not a statement that weights are downloadable; it rejected two Falcon 3 parameter counts read out of the model name; and it rejected the Molmo family and release dates, which rested on a page dateline no quote carried. Two further rejections came from other rubrics and are worth recording because they found real defects rather than duplicating provenance: consistency applied the project’s own validateDataset to a clone with all 30 claims applied and showed that the EleutherAI and Cohere releases assert weight-bearing access types with no licence, which releaseSchema refuses outright; and editorial found that the Cohere weights card states "This model is governed by a CC-BY-NC License with an acceptable use addendum" in prose, so presenting those weights as openly available while omitting the licence understated a non-commercial restriction. The 15 source records were accepted unanimously and dropped anyway, because with every citing record rejected they would have landed as orphans. The dataset is unchanged, the exactly-one-family count stays at 30, and this entry is the only edit.

    1. PreflightRanRe-measured the trunk baseline from the committed data rather than quoting the issue: 40 organizations, 64 families, 92 releases, 217 sources, and 30 organizations holding exactly one family — the number this run set out to reduce. Confirmed all six target creators sit at exactly one family, and that all six publishers already exist, so publishers.json needed no change. Merge-base and HEAD were both 3e255145ed, the same trunk commit the issue measured against.
    2. ScoutRanStored 33 page bodies in total, of which 31 belong to the scout proper and 2 to a read-only stability-ai probe run after the panel reported. They cover that many distinct URLs across eight approved origins (huggingface.co, github.com, allenai.org, docs.cohere.com, cohere.com, www.ibm.com, research.nvidia.com, opensource.org). The count is of stored bodies, not fetch operations: a re-fetch overwrites its cached body rather than accumulating, so the number of operations was higher and is not separately recorded. Two fetches returned HTTP 401 behind a gated form (the Cohere Command R 08-2024 raw README and the Aya Expanse 32B card) and their bodies are stored and hashed as the 401 responses they are. Text extraction turned every HTML tag into a newline so that stripping markup could not concatenate text across element boundaries and invent a phrase that appears on no page.
    3. ReviewRanThree reviewers, one per rubric, launched as parallel independent sub-agents. Each received the claims, their evidence, the dataset and its schema, the cached bytes, and its own rubric — and none received the scout’s reasoning, another reviewer’s verdicts, or the running tally. 90 verdicts cast, each with a written rationale. No reviewer was re-run and no tie was broken: with three reviewers and two outcomes there are no ties, and a claim short of its threshold is simply rejected. All three independently rebuilt the contentHash-to-file mapping by re-hashing the cache rather than trusting filenames, and all three confirmed the 69 quotes verbatim.
    4. GatesRangate-source-approval exited 0 on all six bundles. gate-evidence exited 1 on all six, naming exactly the 15 family and release claims as short of the 3-of-3 long-tail threshold — the gate agreeing with the panel that they must not be applied, which is the gate working rather than failing. gate-dataset, gate-scope and gate-ledger each exited 0 over the unchanged dataset. No gate was skipped, overridden, or re-run with a different threshold.
    5. PublishNot applicableNothing reached the unanimous threshold and survived its dependencies, so there was nothing to apply. The dataset documents were never opened for writing. This entry is the only change the run made.
    6. DeployNot applicableNothing merged, so nothing deployed.

    What was found

    Scouts
    6
    Pages fetched and hashed
    33
    Claims proposed
    30
    Claims per creator bundle, with the review threshold its profile set
    CreatorPolicyThresholdClaims
    eleutherailong-tail3-of-34
    ai2long-tail3-of-34
    tiilong-tail3-of-36
    nvidialong-tail3-of-36
    ibmlong-tail3-of-36
    coherelong-tail3-of-34
    What those claims proposed to do
    KindCountEffect
    Add30Every claim proposed a new record: 15 sources, 6 families and 9 releases. None reached the dataset. The run was scoped to additions only, so no existing record was changed or removed.

    Not covered

    • GPT-J-6B and GPT-NeoX-20B were read and withheld before review: neither card nor repository states a release date in prose, and the only date-bearing metadata is the Hub upload timestamp, which is not a stated release date.
    • The whole Cohere Aya line was read and withheld before review: the documentation states status, modality and context figures but no release date for any Aya model, and the Hugging Face weights cards are behind a gated form that returned HTTP 401.
    • command-r-03-2024 was named as the Command R family’s earliest member but never proposed as a release: Cohere states a deprecation date for it and no release date, and its month rests only on the timestamp in its identifier.
    • This run checked six creators for one additional family each. It is not a sweep of everything those six have published, and it says nothing about the other 24 creators holding exactly one family.
    • stability-ai was named in the coordinating brief in place of cohere, and that brief arrived after the panel had already returned its verdicts, so stability-ai was never scouted and no claim was proposed for it. It was probed read-only rather than left silent: its SDXL and Stable Video Diffusion model cards were fetched and searched for lifecycle, status and release-date wording. The SDXL card contains no line matching any of those terms at all, and the two matching lines on the Stable Video Diffusion card state a usage restriction and a download-statistics note, neither of which is a lifecycle state. On the evidence of that probe stability-ai would have failed on the same required status field as the other six, but a probe is not a scout and this is reported as an indication rather than a result.
    • No context window was recorded for Granite 4.0: the only figure sits in an HTML table whose column-to-model mapping is not legible from the text, so a value would have been read from a column position rather than from a statement.

    What was evaluated

    Reviewers
    3
    Verdicts cast
    90
    Accepted by panel
    15
    Rejected by panel
    15
    Deterministic gates and required checks — 6 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one.
    CheckScopeExitResult
    gate-source-approvalall six claim bundles against the approved origin set at merge-base 3e255145ed0PassExit 0 on every bundle. Each of the eight origins cited — huggingface.co, github.com, allenai.org, docs.cohere.com, cohere.com, www.ibm.com, research.nvidia.com and opensource.org — was already approved, so no origin needed widening and none was widened. This run therefore adds no evidence either way about creators whose own domain is unapproved.
    gate-evidenceall six claim bundles, evidence form and review threshold1FailExit 1 on all six bundles, naming the 15 family and release claims as short of the threshold — variously at 0, 1 or 2 of the 3 required accepts under the long-tail policy, never at 3. The evidence form itself was not the complaint: all 69 entries carry the six required fields with a sha256 hash, an https credential-free url, a real fetch date and a quote at or above the 24-character floor. The refusal is the threshold, and it agrees with the panel. Per ADR 0005 this gate checks form and never fetches, so the verbatim check was done separately against the stored bytes.
    gate-datasetthe committed dataset, read-only coherence check0PassExit 0: all gates passed over 472 records. The dataset is unchanged by this run, so this records that trunk stayed coherent, not that anything this run produced was coherent.
    gate-scopemerge-base 3e255145ed to HEAD, plus the working tree0PassExit 0, reporting one dataset document changed since 3e255145ed and in class: refresh-runs.json, this entry. No other tracked file was touched, and in particular no test, no library and no workflow — a single out-of-class path would have disqualified the whole change. The run’s own artefacts — bundles, verdicts and cached page bodies — live under the git-ignored .modeltree-refresh directory and are invisible to this gate by design.
    gate-ledgermerge-base 3e255145ed to HEAD, this entry against the branch diff0PassExit 0, and the pass is narrower than it looks, which is worth stating rather than leaving to be discovered: the gate classified this entry as a transcription, because the branch changes no dataset document other than the ledger itself, and reported in as many words that the record counts were therefore NOT checked against a diff and should be read as unverified. That is the correct behaviour for a run that published nothing — there is no diff for the counts to be checked against — but it means the numbers here rest on the run’s own artefacts and on the reviewers’ verdict files, not on this gate having confirmed them. Counts were instead derived mechanically from the six bundles and three verdict files rather than typed by hand.
    npm run validateweb/, the full test suite plus Astro and TypeScript diagnostics0Pass107 test files, 2395 tests, 0 failures; 232 files checked with 0 errors and 0 warnings. Run from web/ over this entry, so it covers the entry against refresh-log-schema.ts and not merely the unchanged dataset — which is the only thing this run gives it new to check.
    node .github/scripts/ci-preflight.mjsrepository root; the pull-request checks this branch’s diff selects, anchored at the merge-base0PassSelected and passed two check groups, web-ci and skills-ci, the latter including gate-ledger’s own 221 tests over this entry. It does not run the networked source-link-health sweep, the browser end-to-end check, or the second Python interpreter, and it judges this branch rather than this branch merged into a main that has since moved, so a pass here predicts CI rather than binding it.

    Posted 0 edits

    Nothing reached the dataset. No branch, no commit, no pull request.

    Not posted 30 items

    Rejected by the review panel

    • eleutherai-family-gpt-neo-addfamilies record eleutherai-gpt-neo for eleutherai. [provenance] firstReleaseDate 2021-03-21 (day) is well-sourced by the dated release note '**Update 21/03/2021:** ... two pretrained GPT-Neo models', but status:'current' is unsourced. The two attached quotes state a release event and a family definition ('GPT-Neo refers to the class of models...'); neither states any lifecycle/availability state, so nothing forces 'current' over 'legacy'. A required status filled from what is usual for the field is a guess wearing a vocabulary term (SKILL.md provenance, reject bullet 1). [consistency] The family record itself is well-formed and is NOT a duplicate of the creator's existing eleutherai-pythia family (distinct slug, and GPT-Neo 2021-03-21 predates Pythia 2023-04-03). It is rejected on whole-dataset coherence: its only proposed release, eleutherai-release-gpt-neo-2-7b, is schema-invalid (open-weight with no license) and must be dropped, which leaves this family with zero releases. Running the real validateDataset with the release dropped but the family kept fails with 'family eleutherai-gpt-neo has no releases' (validate.ts:549). As submitted the pair cannot land coherently; fix is to add the missing license to the release and re-submit family+release together.Blocked by rubric:provenancerubric:consistency
    • eleutherai-release-gpt-neo-2-7b-addreleases record eleutherai-gpt-neo-2-7b for eleutherai. [provenance] status:'current' is unsourced: the attached quotes describe only architecture, parameter count and training data, none stating a lifecycle/availability state. Separately, accessType:'open-weight' rests on the bare labeled link '2.7B: https://mystic.the-eye.eu/public/AI/gptneo-release/GPT3_2-7B/'; the page's preceding 'The weights can be freely downloaded' sentence is NOT in the attached quote, so read alone the link does not state that weights are downloadable (mi4 / co5 precedent: a link or licence is not a downloadable-weights statement). [consistency] accessType is 'open-weight' but the record carries no license object. releaseSchema.superRefine (schema.ts:419-426) requires a license whenever accessType is open-weight or both. Running the actual datasetSchema over existing+all-30 fails with 'releases.92.license: is required when a release claims downloadable weights'. The bundle's own incomplete note concedes the license (mit) exists only in YAML front matter and was omitted. As submitted this record makes the dataset fail Zod validation, so it cannot be admitted.Blocked by rubric:provenancerubric:consistency
    • ai2-family-molmo-addfamilies record ai2-molmo for ai2. [provenance] Two provenance failures. (1) firstReleaseDate '2024-09' rests on the page's 'September 25, 2024' dateline, which the bundle itself says is not carried as a quote; a publication timestamp is not a stated release date, and no attached quote states any date ('Today we are releasing 4 samples' gives no date). (2) status:'current' has no attached availability quote, and the family's own released checkpoint is a 'preview', so more than one member could fit.Blocked by rubric:provenance
    • ai2-release-molmo-7b-d-addreleases record ai2-molmo-7b-d for ai2. [provenance] status:'preview' is correctly sourced ('This checkpoint is a **preview** of the Molmo release.'). But accessType:'open-weight' and license.weightsDownloadable:true have no supporting quote: 'This model is licensed under Apache 2.0' is a licence name, not a statement that weights are downloadable (co5 reversal precedent), and 'open vision-language models' does not distinguish open-weight from source-available. releaseDate '2024-09' also rests on the unquoted page dateline rather than a stated release date.Blocked by rubric:provenance
    • tii-family-falcon-3-addfamilies record tii-falcon-3 for tii. [provenance] firstReleaseDate 2024-12 is sourced ('- Model Release Date: December 2024'), and '...a set of pretrained and instruct LLMs ranging from 1B to 10B' defines the family, but status:'current' is unsourced. No attached quote states any lifecycle/availability state; the Falcon3 card carries no status wording at all. 'current' is filled from what is usual, which the rubric forbids.Blocked by rubric:provenance
    • tii-release-falcon-3-10b-instruct-addreleases record tii-falcon-3-10b-instruct for tii. [provenance] Three gaps. status:'current' has no attached lifecycle/availability quote. accessType:'open-weight' rests on '- License: TII Falcon-LLM License 2.0' — a licence name, not a statement that weights are downloadable. parameters.totalBillions:10 is read from the model name; no attached quote states a 10B count ('under 10 billion parameters' describes the family, not this model's count). contextWindow 32000 and the date are fine, but the above sink it.Blocked by rubric:provenance
    • tii-release-falcon-3-7b-base-addreleases record tii-falcon-3-7b-base for tii. [provenance] As with the 10B variant: status:'current' has no attached availability quote; accessType:'open-weight' rests only on the licence name '- License: TII Falcon-LLM License 2.0'; and parameters.totalBillions:7 is taken from the model name — the only count-bearing quote, '...ranging from 1B to 10B', is the family range, not this model's parameter count.Blocked by rubric:provenance
    • nvidia-family-nemotron-nano-2-addfamilies record nvidia-nemotron-nano-2 for nvidia. [provenance] Release date 2025-08-18 is sourced ('### Release Date: 08/18/2025'), but status:'current' is unsourced. 'We are releasing the following three models on Hugging Face...' is an announcement verb, not a lifecycle/availability term (mi4 precedent: an availability/CTA statement does not state a status). No attached quote forces 'current'.Blocked by rubric:provenance
    • nvidia-release-nemotron-nano-9b-v2-addreleases record nvidia-nemotron-nano-9b-v2 for nvidia. [provenance] status:'current' has no attached lifecycle/availability quote. accessType:'open-weight' rests on 'Governing Terms: Use of this model is governed by the NVIDIA Open Model License Agreement' — a licence name (despite the word 'Open'), not a statement that weights are downloadable (co5 reversal). Date, contextWindow 128000 and the licence name are otherwise fine, but status and access type are unsourced.Blocked by rubric:provenance
    • nvidia-release-nemotron-nano-12b-v2-base-addreleases record nvidia-nemotron-nano-12b-v2-base for nvidia. [provenance] status:'current' unsourced; accessType:'open-weight' unsourced; and the record's license.name 'NVIDIA Open Model License Agreement' has NO attached quote in this claim at all (its evidence is only the release date, the three-model announcement, 'Model Developer: NVIDIA Corporation', and the OSI index). Nothing attached states downloadable weights or even names the licence.Blocked by rubric:provenance
    • ibm-family-granite-4-0-addfamilies record ibm-granite-4-0 for ibm. [provenance] firstReleaseDate 2025-10-02 is sourced ('- **Release Date**: October 2nd, 2025'), but status:'current' is unsourced. 'The Granite 4.0 collection comprises multiple model sizes and architecture styles...' defines the collection and states no lifecycle/availability state; no attached quote forces 'current' over 'legacy' (the description itself notes a newer Granite 4.2 exists).Blocked by rubric:provenance
    • ibm-release-granite-4-0-h-small-addreleases record ibm-granite-4-0-h-small for ibm. [provenance] parameters (32B total, 9B active) are well-sourced ('...a 32B parameter long-context instruct model...' and 'a hybrid mixture of experts (MoE) model with 32B total parameters (9B active)'), and osiApproved:true is backed by OSI's Apache page ('Open Source Initiative Approved License'). But status:'current' has no attached lifecycle/availability quote, and accessType:'open-weight' rests only on '- **License:** Apache 2.0' — a licence name, not a statement that weights are downloadable (co5 reversal).Blocked by rubric:provenance
    • ibm-release-granite-4-0-h-tiny-addreleases record ibm-granite-4-0-h-tiny for ibm. [provenance] As with H-Small: parameters (7B total, 1B active) are sourced ('...a 7B parameter long-context instruct model...' and 'a hybrid MoE with 7B total parameters (1B active)'), but status:'current' has no attached availability quote and accessType:'open-weight' rests only on the Apache licence name, not a downloadable-weights statement.Blocked by rubric:provenance
    • cohere-family-command-r-addfamilies record cohere-command-r for cohere. [provenance] status:'current' is scoped, not supported. The only lifecycle term attached is 'Deprecated Sept 15, 2025' for command-r-03-2024 — a different, earliest member — which does not establish the family as 'current'; the 'Live' status of command-r-08-2024 is in a different claim's evidence, not this one. A per-release status is not a family status (co5-family precedent). [consistency] The family record itself is well-formed and NOT a duplicate of the creator's existing cohere-command-a family (Command A, 2025-03) - the description correctly keeps the two lines apart, and Command R (2024-03) is the earlier separate line. Its firstReleaseDate 2024-03 also legitimately precedes its only release (2024-08) without tripping the predates-family rule. It is rejected on whole-dataset coherence: its only proposed release, cohere-release-command-r-08-2024, is schema-invalid (accessType 'both' with no license) and must be dropped, leaving this family with zero releases. The real validateDataset then fails with 'family cohere-command-r has no releases' (validate.ts:549). Fix is to supply the release's license and re-submit the pair together.Blocked by rubric:provenancerubric:consistency
    • cohere-release-command-r-08-2024-addreleases record cohere-command-r-08-2024 for cohere. [provenance] status:'current' (mapped from the quoted 'Live'), accessType:'both' ('a large language model with open weights' plus the API 'Live' listing), contextWindow 128000 ('supports a context length of 128K') and parameters 32B ('32 billion parameter') are all properly sourced. But the record and its statement assert a '4k maximum output' (maximumOutput 4000) with NO attached quote: the models-overview fragment stops at 'delivered in August 2024' and never reaches a max-output figure for this model. The claim states more than the attached sources do (codex precedent: attach the quote, do not map from nothing). [consistency] accessType is 'both' (Cohere API plus open weights), which claims downloadable weights, but the record carries no license object. releaseSchema.superRefine (schema.ts:419-426) requires a license for open-weight or both. Running the real datasetSchema over existing+all-30 fails with 'releases.100.license: is required when a release claims downloadable weights'. The bundle's incomplete note concedes the licence (CC-BY-NC) appears only inside a link element and was not captured. As submitted this record makes the dataset fail Zod validation and cannot be admitted. [editorial] Two editorial defects. (1) Licence/access: the record sets accessType "both" (open weights) yet omits `license`, justifying it as "the weights card names one only inside a link element and states none in prose" — but the cited weights card states in prose "This model is governed by a CC-BY-NC License with an acceptable use addendum". Dropping that CC-BY-NC non-commercial/research restriction while presenting the weights as openly available misrepresents how open the model is. (2) Context basis: the summary asserts "Cohere publishes no integers for either", but the cited command-r documentation page states "a long 128,000-token context length" — an explicit integer — so the prose makes a false negative about its own sources and understates the firmness of the figure.Blocked by rubric:provenancerubric:consistencyrubric:editorial

    Accepted by the panel, then dropped

    • eleutherai-gpt-neo-repository-source-addSource record eleutherai-gpt-neo-repository was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • eleutherai-gpt-neo-2-7b-model-card-source-addSource record eleutherai-gpt-neo-2-7b-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • ai2-molmo-announcement-source-addSource record ai2-molmo-announcement was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • ai2-molmo-7b-d-model-card-source-addSource record ai2-molmo-7b-d-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • tii-falcon-3-announcement-source-addSource record tii-falcon-3-announcement was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • tii-falcon-3-10b-instruct-model-card-source-addSource record tii-falcon-3-10b-instruct-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • tii-falcon-3-7b-base-model-card-source-addSource record tii-falcon-3-7b-base-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • nvidia-nemotron-nano-2-announcement-source-addSource record nvidia-nemotron-nano-2-announcement was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • nvidia-nemotron-nano-9b-v2-model-card-source-addSource record nvidia-nemotron-nano-9b-v2-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • nvidia-nemotron-nano-12b-v2-base-model-card-source-addSource record nvidia-nemotron-nano-12b-v2-base-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • ibm-granite-4-0-announcement-source-addSource record ibm-granite-4-0-announcement was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • ibm-granite-4-0-h-small-model-card-source-addSource record ibm-granite-4-0-h-small-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • ibm-granite-4-0-h-tiny-model-card-source-addSource record ibm-granite-4-0-h-tiny-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • cohere-command-r-model-card-source-addSource record cohere-command-r-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record
    • cohere-command-r-08-2024-weights-card-source-addSource record cohere-command-r-08-2024-weights-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked by check-bundle-pairingno surviving citing record

    What this run does not prove

    • A withheld tranche is a real outcome of the policy, not a failure of the sources. Every page came from an origin this repository already approves and every quote was verified verbatim against stored bytes. The run reports six creators it could not record rather than six it stretched a citation to reach.
    • The blocker is a required field with no sourceable value, and better scouting does not reach it. releaseSchema and familySchema both require status; the schema offers preview, current, legacy, deprecated and research and no unknown member; and a Hugging Face model card states no lifecycle term at all. Where a creator does publish a status — Cohere lists command-r-08-2024 as Live — the mapping to current was accepted without argument. So this is not a claim that status is unsourceable in general; it is that a creator who publishes only model cards has no route to a release record, which is most of the long tail.
    • This entry does not claim the panel was wrong, and that was checked rather than assumed. The provenance rubric carries a section written specifically to stop a dissent on a mapped vocabulary field from blocking long-tail claims, on the ground that status and contextWindow are recorded by mapping a creator’s own wording and no source speaks the dataset’s terms — so a panel that treated mapping as inference would have been misapplying its own rubric, and under a unanimous bar that error would silently block every long-tail creator. The rejections were therefore read against that section before this entry was written. They conform: the panel accepted the mapping wherever a status was actually stated, taking Cohere’s quoted "Live" to current and Ai2’s quoted "preview" to preview in as many words, and rejected status only where no attached quote states any lifecycle state at all — which is the mi4-release-large-3 row the rubric lists as a rejection that stands unchanged. Two of the three rejections that did not come from provenance found defects the scout had missed outright, which is further evidence the panel was reading rather than rubber-stamping.
    • The run deliberately did not re-scout to press a different phrase into status. Where the rubric names a remedy it is to attach a quote the page already carries, not to map from nothing, and several rejections here are exactly that repairable kind — but repairing them would not have saved the records, because the unrepairable status gap sits on the same claims. Re-running a rejecting reviewer to obtain a different verdict is forbidden outright and was not done.
    • The six bundles, all 90 verdicts with their rationales, and the 31 stored page bodies are reproducible artefacts of this run, but they live under a git-ignored directory and go with the machine. This entry is the durable record; once that directory is gone, every citation here can only be rechecked by re-fetching a live page that may have changed.
    • Nothing here says the dataset is current. Six creators were checked for one additional family each. A creator not named in this entry was not looked at, and a family not proposed is not thereby known to be absent.

    Follow-ups — proposed, not fixed

    • The tranche this run set out to publish is blocked on a schema-and-rubric interaction rather than on evidence: status is required on every release and every family, the provenance rubric requires a quote that forces one enum member, and a model card states no lifecycle term. Three routes are open and this run took none of them, because all three are outside what a refresh agent may decide: make status optional, or give it an unknown member, so a release can be recorded without asserting a lifecycle nobody stated; write down that a creator publishing weights with no stated lifecycle is recorded as current, which is an editorial convention a human should set rather than a reviewer infer; or add these families through a change that takes the ordinary path, where a human reviews the reading on the way in.
    • accessType: open-weight was rejected on five releases because it rested on a licence name, a download link, or the word "open" in a licence title, none of which states that weights are downloadable. Several of the cards do carry such a statement in prose — the GPT-Neo README says the weights can be freely downloaded a sentence before the link that was quoted instead — so this particular gap is a scouting defect and not a structural one, and is worth re-attempting on its own once the status question is settled.
    • The research behind the 15 unanimously accepted source records is sound and was dropped only because the records citing it fell. If the status question is answered, these six creators do not need re-fetching from scratch, but only if the stored bodies are captured before the machine goes.
    • Two claims failed on defects worth fixing regardless of the status question: the Cohere Command R 08-2024 release asserted a 4k maximum output with no quote attached, and the NVIDIA Nemotron Nano 12B v2 Base release named a licence its evidence never quoted. Both are attach-the-quote-or-drop-the-field problems in the scout, not disagreements with the panel.
  2. 2026-08-30-605b1a

    Long-tail breadth tranche 2026-08-30 — Aleph Alpha, AI Singapore, Nous Research, Reka AI, Liquid AI, Xiaomi

    Scope requested: Six creators with no reviewed profile in tools/updater/profiles — Aleph Alpha, AI Singapore, Nous Research, Reka AI, Liquid AI and the Xiaomi LLM-Core team — under the long-tail unanimous 3-of-3 policy, restricted to the 42 origins already approved at the merge-base. No profile catalogue was edited and no origin was added: widening the trust boundary is deliberately a human act.Ran, changed nothing0 edits posted · 10 items withheld

    Researched six long-tail creators to take the catalogue from 34 organizations to 40 and published none of them. Every fact was fetched and hashed from origins this repository already approves — huggingface.co, github.com, arxiv.org and opensource.org — never read from a search result, and gate-source-approval passed on all six bundles, so origin discipline was not the blocker. The panel was the blocker, and beneath it a structural one. Under the unanimous 3-of-3 long-tail bar the provenance rubric rejected 46 of 54 claims, most of them because the record asserted fields no attached quote stated: lifecycle status, modality lists, downloadable weights, and above all a release date. No approved-origin page states a release date for any of these six models. No model card states a release date for the model it describes. All six carry dates, and every date found belongs to a category that is not a release date: a training-data cutoff (Aleph Alpha, cutoff date 04/2023); Hub revision tags (AI Singapore, Current Version: `14.04.2025` at line 29, and a revision="18.12.2024" argument at line 123 under the comment "Specify the revision here"); paper years in BibTeX blocks (Nous Research, Liquid AI and Xiaomi, year={2025}); benchmark names (Reka AI, AIME-2024; Xiaomi, AIME 2024 and AIME 2025 across six table rows); values inside worked examples and generated output (Liquid AI, "date": "2023-11-20" in a tool-call sample; Aleph Alpha, the founding of Rome and a 2016 census figure in sample completions); EU legal instrument numbers that only resemble slashed dates (Aleph Alpha, Directive (EU) 2019/790 and Regulation (EU) 2016/679); and a changelog line dating a different model (Xiaomi, [2025.05.30], which dates MiMo-7B-RL-0530 and not the MiMo-7B-RL this run targeted). Each is named with its card and line so the exclusion can be checked. This states what the dates found are and why none is a release date; it is not a claim that no other date exists on these pages. The Hub API states createdAt, which is repository creation and not a stated launch; arXiv states when a paper was submitted, which is a fact about the paper and not the model; and the two GitHub repository READMEs that do carry dated release news date other models — SEA-Guard and a March 2026 embedding family for AI Singapore, and MiMo-7B-RL-0530 rather than MiMo-7B-RL for Xiaomi. releaseSchema requires releaseDate, and the schema is explicit that a field no approved source states means the record is withheld rather than guessed, so no release could be written for any creator. The eight claims the panel did accept were three sources, three publishers and two conflict records; applying them would have left sources cited by nothing, which validate.test.ts forbids, and a creator with no release renders in neither branch of the tree. They were dropped rather than applied. The dataset is unchanged and this entry is the only edit.

    1. PreflightRanRe-measured the trunk baseline from the committed data rather than quoting it: 34 organizations, 55 families, 83 releases, 202 sources, 42 publishers. Enumerated the approved origin set from gate-source-approval itself, by running it over a zero-claim probe bundle, which reports anchors.approvedOrigins whether it passes or fails: 42 origins at merge-base c9e01df.
    2. ScoutRanMade 56 fetches over 48 distinct pages — 39 initial targets across four batches, 1 ad-hoc re-check of the AI Singapore card, 6 systematic date re-checks which deliberately re-pulled bytes already held rather than trusting the stored copy, 6 fetches around the Hermes 3 withholding, being four Hermes-4 records, the Hermes-3 LICENSE whose 404 caused it and one arXiv abstract, and 4 GitHub organization and repository pages — of which 8 re-fetched a page already held — across model cards, Hub API records, config.json files, arXiv abstracts, OSI licence pages and GitHub organization and repository pages. The pagesFetched field below records that 56: it counts fetch operations that left a stored body, not distinct pages, of which there are 48. The 8-page gap is the re-fetches, and the AI Singapore card accounts for three of the stored bodies on its own — the original, the ad-hoc re-check and the systematic one — all three byte-identical at sha256:2d4396287f5d. The count is of page bodies only: date-recheck.json sits in the same folder and is the manifest of that re-check rather than a fetched page, so it is excluded. Every quote in every bundle was machine-checked to be a literal substring of the exact bytes whose sha256 the evidence records; the generator refuses to emit a bundle otherwise and caught six bad quotes on its first pass, one of which had only ever looked verified because it sat behind a function the prototype never called.
    3. ReviewRanThree reviewers, one vote each, launched in parallel and given the claim, its evidence, the dataset slice and one rubric — never the chair reasoning, another reviewer output, or a running tally. 162 verdicts over 54 claims. The editorial reviewer first returned a single 138-character rationale repeated 54 times and was sent back for claim-specific rationales with its votes explicitly untouched; the rewrite changed no vote. Provenance accepted 8, consistency 54, editorial 54.
    4. GatesRangate-evidence exit 1 on all six bundles, deriving the 3-of-3 threshold itself from the reviewed-profile set on disk rather than from any policy field in the bundles. gate-source-approval exit 0 on all six. gate-dataset exit 0 on the unchanged dataset. gate-scope reports this log entry out of the ADR 0003 qualifying class, which is expected rather than a problem to work around.
    5. PublishNot applicableNo claim reached the dataset, so there was nothing to publish. The dock agent opened no pull request and recorded no gate verdict on its own work; this run ends at the review gate and hands off to an independent reviewer.
    6. DeployNot applicableNothing was merged, so nothing deployed.

    What was found

    Scouts
    6
    Pages fetched and hashed
    56
    Claims proposed
    54
    Claims per creator bundle, with the review threshold its profile set
    CreatorPolicyThresholdClaims
    Aleph Alphalong-tail3-of-38
    AI Singaporelong-tail3-of-39
    Nous Researchlong-tail3-of-310
    Reka AIlong-tail3-of-38
    Liquid AIlong-tail3-of-310
    Xiaomi LLM-Core Teamlong-tail3-of-39
    What those claims proposed to do
    KindCountEffect
    Add52Six creators as organization, family, release and publisher records, with 28 supporting source records. None reached the dataset.
    Conflict2Two disagreements between primary sources, recorded rather than smoothed: the LFM2-1.2B context window and the Hermes-4-14B parameter total. Both were accepted unanimously and both describe releases that were not published.

    Not covered

    • No creator-owned domain was consulted for any of the six, because none is an approved origin. What those newsrooms say about release dates, lifecycle status and licensing is therefore unknown to this run rather than absent from the world.
    • The /compare page-weight guards in web/src/lib/comparison.test.ts were never exercised against new data, because no release was added. Whether six more creators would breach the per-release or total budget remains untested and is still an open question for whoever owns that guard.
    • Hermes-3-Llama-3.1-8B was withheld before review rather than judged by the panel, so no verdict exists on it.
    • Only one review round was run. The rubric remedy for an unquoted field is to attach the quote, and for the fields other than releaseDate that remedy was not attempted, because releaseDate alone is sufficient to withhold every release record.

    What was evaluated

    Reviewers
    3
    Verdicts cast
    162
    Accepted by panel
    8
    Rejected by panel
    46
    Deterministic gates and required checks — 4 of 6 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one.
    CheckScopeExitResult
    gate-evidenceAll six claim bundles, annotated with the panel verdicts.1FailRefused 46 claims that did not reach the unanimous long-tail bar. The gate derived the 3-of-3 threshold from the reviewed-profile set on disk rather than from the policy field in the bundles, which is the behaviour ADR 0002 requires, and it reached the same answer as the panel independently.
    gate-source-approvalAll six claim bundles against the 42 origins approved at merge-base c9e01dffbf0d9bc92d0b6ee39f286e22120d7246.0PassEvery citation in every bundle sits on an already-approved origin, and no proposed source rests on a creator-owned domain. This is the one gate that passed cleanly on all six, and it is worth recording that the run did not fail on origins: it failed on what the quotes state.
    gate-datasetThe committed dataset, unchanged by this run.0PassThe dataset is unchanged and remains internally coherent: 202 sources, 42 publishers, 34 organizations, 55 families, 83 releases. No orphaned source, no dangling reference, no release predating its family.
    gate-scopeThe branch diff from its merge-base to its tip.1FailReports web/src/data/refresh-runs.json as outOfClass, which is correct and expected rather than a problem to work around: the ADR 0003 qualifying class is the dataset documents raw.ts composes, and the refresh log is deliberately not one of them, so any run that logs itself leaves the class by construction. The gate was not modified, the file was not added to ALLOWED_PATHS, and the log entry was not dropped to stay in class. The consequence is that this change is outside ADR 0003 and merges the ordinary way, with a human — which is what this run wanted anyway, since it published no data.
    npm run validateweb/ — the full vitest suite and Astro/TypeScript diagnostics.0PassRun from web/ after this entry was written, so it covers the entry itself against refresh-log-schema.ts as well as the unchanged dataset.
    node .github/scripts/ci-preflight.mjsThe pull-request checks this branch diff actually triggers, measured from the merge-base.0PassRun from the repository root. It does not cover the networked link-health sweep or the second Python interpreter, which it prints on every run and which remain unrun here.

    Posted 0 edits

    Nothing reached the dataset. No branch, no commit, no pull request.

    Not posted 10 items

    Rejected by the review panel

    • aleph-alphaPharia-1-LLM-7B-control. The parameter count and the 8,192-token sequence length are quoted from Aleph Alpha own card, but no attached quote states a release date, a lifecycle status, a modality set or that the weights are downloadable, and the OSI licence index heading does not by itself establish that the Open Aleph License is absent from the register. The organization record additionally inferred a German identity from a Heidelberg location and did not establish the relationship between the names "Aleph Alpha GmbH", "Aleph Alpha Research" and "IPAI Aleph Alpha Research GmbH", all three of which appear across the sources.Blocked by provenance
    • ai-singaporeLlama-SEA-LION-v3-8B-IT. The card describes AI Singapore in its own words as a national programme supported by the National Research Foundation and hosted by the National University of Singapore, which is what typed the organization research-lab, but the release record rested on an unquoted date, status, modality set and weights statement, and the 128k-to-131072 reconciliation was not supported by an attached quote.Blocked by provenance
    • nous-researchHermes-4-14B. Licence and parameter assertions were not carried by the attached quotes, the website was not quoted, and the release date, status, modalities and weights were unstated. The config source and the publisher record were accepted; the records that would have made them mean something were not.Blocked by provenance
    • reka-aiReka Flash 3. The card states that the model was trained from scratch, which is why no derivation was recorded despite the configuration declaring a Llama architecture class, but the publisher attribution, the website and the release date, status, modality and weights fields were all unquoted.Blocked by provenance
    • liquid-aiLFM2-1.2B. The organization identity and the LFM Open License name were not carried by the attached quotes, and the release date, modalities and status were unstated. The GitHub organization is Liquid4All rather than a name matching the model prefix, which the record noted but did not resolve from a quote.Blocked by provenance
    • xiaomiMiMo-7B-RL. The author attribution, the governance relationship between the LLM-Core team and Xiaomi, and the web endpoints were unquoted, as were the release date, status, modalities and weights. The GitHub organization XiaomiMiMo publishes neither a website nor a location, which is why the record proposed the GitHub profile itself as the organization website rather than asserting xiaomi.com.Blocked by provenance

    Accepted by the panel, then dropped

    • accepted-sources-and-publishersEight claims were accepted unanimously by all three rubrics and none was applied. Six of them were adds: three source records — the Hermes 4 14B and Reka Flash 3 configuration files and the Aleph Alpha GitHub organization profile — and three publishers, Nous Research, Liquid AI and AI Singapore. The remaining two were the conflict records described below. The accepted set spans five creators and completes none of them: Nous Research alone reached both a publisher and a configuration source; Liquid AI and AI Singapore reached a publisher and no source; Reka AI and Aleph Alpha reached a source and no publisher; Xiaomi reached nothing. With no accepted release or organization to cite them, the sources would have been orphaned, which validate.test.ts forbids outright, and a publisher publishing nothing is a record that renders nowhere. This is the same shape of outcome recorded for Mistral AI and Cohere in the 2026-08-27 run.Blocked by ai-singaporealeph-alphanous-researchreka-ailiquid-ai

    Verification date deliberately held back

    • nous-research-hermes-3-llama-3-1-8bWithheld before review and never put to the panel. This model was the tranche original Nous candidate and a prior session had flagged it as possibly uncitable; the flag was checked on all three of its legs and held. The card carries exactly one licence statement, the bare slug "license: llama3", while the same card declares its base model to be meta-llama/Meta-Llama-3.1-8B and the Hub tags declare base_model:meta-llama/Llama-3.1-8B. The slug names the Llama 3 licence and the base names Llama 3.1, and https://huggingface.co/NousResearch/Hermes-3-Llama-3.1-8B/raw/main/LICENSE returns HTTP 404, so nothing on an approved origin resolves the disagreement. An open-weight release requires a licence object, and guessing which licence the slug meant is exactly what this dataset refuses. Hermes-4-14B was carried instead, where the slug apache-2.0 and the base Qwen/Qwen3-14B agree.Blocked by licenseSchema

    Sources conflict, so no value changed

    • liquid-ai-lfm2-1-2b-context-windowThe LFM2-1.2B card states a context length of 32,768 tokens while the repository config.json sets max_position_embeddings to 128000. The conflict was recorded with the creator own stated figure taken as the reading, and was accepted unanimously. It describes a release that was not published.Blocked by liquid-ai
    • nous-research-hermes-4-14b-parameter-totalThe same Hub object reports safetensors.parameters.BF16 as 14768307200 and safetensors.total as 424960 for a model shipping six shards over 29.5 GB. The BF16 figure was taken as the reading. For four of the other five creators the two figures agree exactly, and in each of those four the total also equals the sum of the whole parameters block: AI Singapore 8030261248, Liquid AI 1170340608, Reka AI 20905482240 and Xiaomi 7833409536. The fifth, Aleph Alpha Pharia-1-LLM-7B-control, carries no safetensors block at all in its Hub API record, so it offers no comparison rather than an agreeing one; counting it among the agreeing records would have been the same absence-read-as-agreement error this entry documents elsewhere. For Hermes-4-14B the parameters block sums to the BF16 figure and not to the total, which is what identifies the total as the anomaly. Accepted unanimously; describes a release that was not published.Blocked by nous-research

    What this run does not prove

    • A withheld tranche is a real outcome of the policy rather than a failure of the sources. Everything above rests on pages fetched from origins this repository already approves, and the run reports six creators it could not source rather than six it stretched a citation to reach.
    • The blocker is structural and no amount of better scouting reaches it. Every one of the 42 approved origins is either a generic host — github.com, huggingface.co, arxiv.org, opensource.org, storage.googleapis.com — or the own domain of a creator already in the catalogue. A creator that is not yet in the catalogue therefore has no approved domain of its own, so the announcement page that would state its release date is precisely the page the trust boundary excludes. Approving a new origin is a human act by design, so a refresh agent cannot close this gap from inside a run. What that blocks is the unattended refresh path, and only that. The gates skill names two remedies for extending the trust boundary — add the origin to a profile catalogue, or cite it in a change that takes the ordinary path — and the second requires no origin to be approved in advance, because a human reviews the citation on the way in. This entry should not be read as saying the catalogue cannot grow. It says a run like this one cannot grow it, and the reference below points at where the ordinary path is being taken instead.
    • The statement about dates on these cards was wrong twice before it was right, and both attempts are recorded here rather than quietly replaced. The first said the cards carry no date-bearing line at all; the scan behind it matched month names and ISO forms only, so a dotted Current Version: `14.04.2025` fell straight through. The second said five of the six carry none; the widened scan behind that one also discarded every line longer than 300 characters, so a 480-character paragraph naming AIME-2024 and a 328-character changelog line beginning [2025.05.30] never reached the pattern at all. Two unrelated mechanisms, one a regex gap and one a length filter, produced the same false sentence, which is the evidence that widening the scan a third time would not have fixed it. What the sentence claimed was the defect, not the tool behind it. A scan supports a statement about what it matched and can never support a statement about what is absent, because the forms it does not encode are invisible to it by construction: abbreviated and non-English month names, quarters and half-years, relative dates, epoch seconds, dates in HTML attributes, JSON-LD or YAML front matter, and dates that appear only inside images. The corrected sentence is therefore carried by the categorisation of the dates actually found, each named with its card and line, and by no scan returning nothing. Both errors were caught by fetching the six cards rather than reading the diff, QA finding the first and review the second, and the AI Singapore card hashed byte-identical to the copy the run already held, sha256:2d4396287f5df1e4993d8a0e22b6abdf7b6262b4323476946eb8519be5f3ef7d, so neither error was a matter of the page having changed. The dates this entry names, and the fragments quoted around them, were each re-read from those stored bytes at the line cited — fourteen checks, all of which hold. One did not at first: the AI Singapore value is written inside backticks on the card and this entry had dropped them, corrected here so that a reader grepping the card for what this entry prints finds it.
    • The issue framed part of this tranche as closing a non-transformer gap. That framing is factually wrong and was not carried into any claim. The LFM2-1.2B card describes the architecture in its own words as a hybrid model with multiplicative gates and short convolutions, ten double-gated short-range convolution blocks and six grouped query attention blocks, and the config lists full_attn_idxs [2,5,8,10,12,14] with 32 attention heads and 8 key-value heads. LFM2 contains attention. Writing the issue framing into the dataset would have published a false claim beside a citation that contradicts it.
    • Nous Research was typed company rather than community. The community clause needs a source showing that contributors outside the entity own appointment chain decide its releases, and nothing consulted states that; the ordered procedure then falls through to the company fallback. This is a classification the panel judged rather than a fact any source states.
    • The panel is three model instances reading the same pages, so 2-of-3 or 3-of-3 buys independence of reasoning and not independence of training. A source that is itself wrong could carry all three. The provenance rubric was also the only one of the three that rejected anything here, which is worth noticing rather than smoothing: consistency and editorial each accepted all 54, so the unanimous bar was carried entirely by one rubric.
    • The editorial verdicts were rewritten once. The first pass returned a single sentence repeated across all 54 claims, which fails the requirement that a rationale state what decided it; the reviewer was asked for claim-specific rationales with its votes explicitly untouched, and it changed none. The rewritten rationales are the ones recorded, and a reader should know they were produced on a second pass. They are also not fully claim-specific: the 54 recorded editorial rationales are 28 distinct strings, nine of them reused - two covering six claims each, four covering four, one covering three and two covering two - where provenance and consistency each recorded 54 distinct ones. The account of the first pass is itself unverifiable from what survives, because the rewrite overwrote it, so both that a single rationale was repeated 54 times and that no vote changed are taken from the run record rather than from an artefact a later reader can inspect. That absence was checked across every directory this run wrote rather than only the one holding the verdict files: the sole artefact predating the rewrite is the packet the reviewers were given, which carries the claims and the dataset slice and records no vote or rationale at all. This is the one statement in this entry that no preserved byte can settle.
    • This entry was written by the run rather than transcribed after it, which is why gate-scope reports the branch out of the ADR 0003 class. The entry records gate results for npm run validate and ci-preflight that were observed after the entry itself was written, so those two lines describe the tree as committed rather than the tree as it stood when the earlier gates ran.
    • The account of what the panel accepted was wrong in two places and is corrected here rather than silently replaced. The withheld entry for the accepted sources and publishers stated three sources and three publishers and then named four of each, filing the Liquid AI record as a source and the Reka AI record as a publisher when the bundles have them the other way round; and the follow-up resting on it named Nous Research, Reka AI and Liquid AI as each having reached both a publisher and a configuration source, when only Nous Research did. The aggregate arithmetic was right throughout — 54 claims, 162 verdicts, 8 accepted and 46 rejected — because the error lay inside the eight, so no check that reconciled totals could see it. The first defect was visible as a disagreement between the summary and the detail; the second contradicted nothing in the entry and was found only by counting the accepted claims out of the six bundles again. Both were settled against the bundles rather than by making one field agree with another, which is the only method that could have settled the second.
    • The evidence behind this entry is verifiable, and an earlier version of this caveat said the opposite. Every quote was machine-checked as a literal substring of the exact bytes whose sha256 the bundle records; the six bundles survive; and the fetched page bodies survive too, 56 of them across the six run directories this run wrote, 50 in their pages folders and 6 more re-fetched into the date re-check folder, covering all 31 of the 31 URLs cited as evidence, with no cited URL left without a stored body. By origin they are 38 from huggingface.co, 11 from github.com, 4 from arxiv.org and 3 from opensource.org; the 38th is the re-checked AI Singapore card, which is stored without a host prefix in its filename and so is easy to drop from a breakdown that reads filenames. The earlier version said bodies survived for 19 of 43 pages covering 15 of 31 cited URLs, that every preserved body came from huggingface.co, and that nothing this entry says about a github.com, arxiv.org or opensource.org page could be checked against bytes held here at all. That was false in every part. It was measured over one run directory while the run had been writing to six, and the 19 huggingface.co bodies it found were the true contents of that directory rather than of the evidence set. Three statements it listed as resting on the run record are in fact held in preserved bytes: SEA-Guard appears on the stored AI Singapore GitHub README at lines 948 and 949, and at line 911 as SEA-GUARD in capitals inside a table link, which is a different string and would not be found by a reader grepping for the spelling this entry uses; the March 2026 SEA-LION-Embedding suite is at line 947 of the same page, and MiMo-7B-RL-0530 appears on the stored Xiaomi GitHub README at lines 792 and 799, line 792 being the changelog line dated [2025.05.30]. So are two facts asserted elsewhere in this entry: the Hermes-4-14B conflict figures are in the stored Hub API body, which records safetensors.total as 424960 and safetensors.parameters.BF16 as 14768307200, and the Hermes-3-Llama-3.1-8B LICENSE 404 is stored as a 15-byte body reading Entry not found, the 404 response being itself the artefact. What remains true is the limit ADR 0005 records: gate-evidence validates the shape of a contentHash and a quote and does not, and cannot as built, verify that the digest is the hash of the page at the url, because the gate fetches nothing. The bodies kept here close that gap for a reader who holds them; they never closed it for the gate, which did not look at them.
    • Eight statements in this entry were wrong when first written, counting the two attempts at the date claim separately, and all eight failed in the same way: a correct aggregate carrying a wrong decomposition. Two were absence claims about dates that the scans behind them could not support. One stated the right totals for what the panel accepted — three sources, three publishers and two conflict records — while filing the Liquid AI record as a source and the Reka AI record as a publisher, which the bundles have the other way round. One called 48 fetches 48 unique pages, when the run had in fact made 56 fetches over 48 pages, so the figure it was corrected to was itself too low. One said a single editorial rationale covered six claims when two of them do. One said the safetensors figures agree for the other five creators, when they agree for four and the fifth carries no safetensors block to compare at all, which is an absence read as an agreement. In every case the aggregate itself held: 54 claims, 162 verdicts and 8 accepted against 46 rejected are all correct as written, and every check that reconciled them confirmed them. That is precisely why the errors survived, because a total ties whichever way its parts are attributed, so a check that reconciles totals is structurally blind to a permutation inside one. Each was found by recounting against the artefacts rather than by reading the entry, and the accepted-claim inversion contradicted nothing in the entry at all, so no amount of internal comparison could have surfaced it. The fourth says the most about what the gates cover. It was present, in three places, in the tree the review gate read at commit 473dac1, and review passed that tree without examining pagesFetched at all. That was not carelessness: the same review verified the accepted-claim decomposition down to per-creator attribution out of the bundles, because that claim had been named for it, and it did not generalise from one wrong count to checking every count. Coverage followed the brief it was given. A fifth was then found by the coordinating session walking the run directories instead of trusting this entry, and it is the most self-implicating of the set. The thesis of this entry is that an absence claim must be a matched set - that saying no approved source states a release date obliges you to name every date found and say what each one is instead. Its own absence claim about its own evidence did the thing it warns against. It counted preserved bodies in the pages folder of one run directory, found all 19 of them to be huggingface.co, and stated universally that no github.com, arxiv.org or opensource.org body had ever been persisted, when 18 such bodies were sitting in four of the five sibling directories the same run had written that evening. The aggregate held again - 19 huggingface.co bodies really were in the folder that was looked at - and the decomposition was again what was wrong. The review gate confirmed that claim rather than catching it, reporting independently that all 19 were huggingface.co, because it inherited the directory scope instead of re-deriving it. That is the sharpest result this run produced about what agreement is worth: independence of reviewers does not give independence of scope, and two parties looking in the same wrong place agree perfectly. The correction was then widened twice more, once from one run directory to six and once again from the pages folders to the date re-check folder, and neither widening began in a disagreement about a number. Each came from asking what else the run had written rather than whether a figure was right, which is a different question and is the one that found every instance. The last of them was caught by the review gate, and that balances the account. Correcting the coverage caveat moved the preserved-body count from 19 to 56, but found.pagesFetched and the sentence explaining it kept the figures computed before the widening, so the entry asserted 48 fetches over 43 distinct pages in one place and 56 stored bodies in another, and named the first as correct as written. A stored body is the product of a fetch, so the two could not both stand. Review found that contradiction; neither this run nor the session coordinating it did. The lesson is narrower and more useful than checking the counts again: when a correction changes the size of a set, every number computed over that set has to be re-derived, including numbers in structured fields the correction did not edit. The prose and the field were two views of one measurement, and only one of them moved. Of the eight defects, two were caught by gates that had been pointed at them, one by the review gate looking where it had not been told to look, three by this run auditing its own artefacts and two by a reader who widened the measurement. The general form is that a document internal agreement is not evidence about the world, and where it is the only evidence available the honest report is unverified rather than consistent.

    Follow-ups — proposed, not fixed

    • A new creator cannot be added by an unattended refresh from the approved origin set alone, because releaseDate is required and only a creator own announcement states one. Three routes stay open and this run took none of them, because all three are outside what a refresh agent may do: approve each new creator origins in a profile catalogue as a reviewed human step before its refresh runs; add the creator through a change that takes the ordinary path, where a human reviews the citation and no origin has to be approved in advance; or change what the schema requires of a release whose date no approved source states.
    • The six bundles and all 162 verdicts are reproducible artefacts of this run but live under a git-ignored directory, so they disappear with the machine. The fetched bytes are kept in full there: 56 body files across six run directories cover all 31 of the URLs cited as evidence, so every citation here can be checked against stored bytes while that directory exists and only by re-fetching a live page that may have changed once it does not. If the origin question above is answered, the research does not need repeating from scratch, but only if it is captured before that happens.
    • Five of the six creators reached at least one unanimously accepted record: Nous Research a publisher and a configuration source, Liquid AI and AI Singapore a publisher only, Reka AI and Aleph Alpha a source only, and Xiaomi none. Nous Research has the least left to fetch and would be the cheapest to revisit first once a date-bearing origin is available. None of them is close to publishable in the meantime, because every release still needs a date that no approved origin states.
  3. 2026-08-28-cff539

    Long-tail refresh 2026-08-28 — Cohere

    Scope requested: One long-tail creator, Cohere, under the unanimous 3-of-3 policy, restricted to the two origins PR #445 approved for it: docs.cohere.com under /docs/ and cohere.com under /blog.Ran, changed nothing0 edits posted · 6 items withheld

    A second attempt at Cohere, run after PR #459 amended the provenance rubric to treat vocabulary mapping as a recording step rather than an inference. The amendment worked and is not why this run publishes nothing: the provenance reviewer stated in terms that "mapping Live to current would be a permitted recording step at the same entity level", conceding the exact objection that sank run 2026-08-27-4f1c9e. Three different blockers stopped it instead. No release claim could be assembled at all — familySchema.firstReleaseDate is a day-precision isoDate and no approved origin states one for any Cohere family, while the one model with a day-precision date, Command A+, is released under two deployment routes and so requires a license object whose osiApproved field no approved origin states. The Command A+ family claim that was proposed was rejected 0-of-3, all three rubrics independently holding that Cohere's own words "the last model in the Command A family" make it a model inside a family rather than a family. The organization and publisher claims were rejected 2-of-3 by provenance on the same ground as last time — no quote states the schema classification "company" — which PR #459 did not address. Two source claims were accepted unanimously and still were not applied: with no record citing them they would be unreferenced sources, which validate.test.ts forbids. The dataset is unchanged; this entry is the only edit.

    1. PreflightRandrydock is not on PATH, so the manual posture applied: implement, commit, report, and stop at the review gate rather than landing anything. Clean tree rebased onto origin/main at 1512da2, then rebased again mid-run onto 73f92f3 after three pull requests landed — #462 correcting the featured criterion, #463 adding the A-Z provider directory and #464 the lineage explorer. The second rebase was clean and releases.json did not conflict. Every gate below was recomputed against the final anchor 73f92f3 rather than carried forward, because a gate result binds to the commit it was computed against.
    2. ScoutRan26 pages fetched over real HTTP from the two approved origins only, all HTTP 200, with the exact bytes and their sha256 recorded in a fetch manifest before any quote was written. Every quote was then extracted programmatically from those hashed bytes by a builder that throws if the quote is absent, ambiguous or under the 24-character minimum; it caught two real faults during construction — a search-offset bug, and an identifier that appears in both a model table and a serving-platform table and so could not be quoted unambiguously. docs.cohere.com/changelog/, /llms.txt and /v2/docs/, and cohere.com/research/, were deliberately not fetched: they fall outside the allowed_paths in the approved origins catalogue.
    3. ReviewRanOne blind panel, three rubrics, run as three parallel sub-agents on three different model families so no rubric shared a training lineage with another. Each reviewer saw only the stripped bundle — statement, values, evidence, dataset slice and its own rubric — and never the scout's reasoning, another reviewer's verdict, or the running tally. 18 verdicts cast, every one carrying a rationale, all published verbatim in the pull request body. No reviewer was re-run and no threshold was adjusted.
    4. GatesRangate-evidence and gate-source-approval ran before anything would have been applied, anchored on the committed dataset at merge-base 73f92f3 computed by the gate rather than supplied. gate-dataset, npm run validate and gate-scope ran afterwards. gate-evidence was first invoked wrongly and exited 2; exit 2 is "the gate could not run", so it was treated as a failure and re-run correctly rather than read as a pass.
    5. PublishRanAn ordinary pull request carrying this log entry and no dataset change. Auto-merge was neither requested nor eligible: refresh-runs.json is outside the gate-scope qualifying class, so ADR 0003 does not cover this change and a human merges it.
    6. DeployNot applicableNo dataset change to deploy. The rendered tree is byte-identical to the one already on main; only the /refresh/ page gains this entry.

    What was found

    Scouts
    1
    Pages fetched and hashed
    26
    Claims proposed
    6
    Claims per creator bundle, with the review threshold its profile set
    CreatorPolicyThresholdClaims
    coherelong-tail3-of-36
    What those claims proposed to do
    KindCountEffect
    Add6Would have added three sources, one publisher, one organization and one family. No release claim was proposed, because no release could be assembled from the approved origins without guessing a required field. None was applied.

    Not covered

    • No release claim for any Cohere model. familySchema.firstReleaseDate is a day-precision isoDate with no datePrecision companion, and no approved origin states a first release date at that precision for the Command, Aya, Embed, Rerank, Parse or Transcribe families. The best available are month-precision prose — "delivered in August 2024", "delivered in December 2024" — and the snapshot dates encoded in identifiers such as command-a-03-2025.
    • command-a-plus-05-2026 has a quotable day-precision release date and was still not claimed: the model card states both "Cohere Deployment" and an Apache 2.0 open-source deployment, which is accessType "both", and releaseSchema then requires a license object whose osiApproved is a non-optional boolean no approved origin states. This is open issue #461.
    • rerank-v3.5 carries the only other day-precision date on an approved origin, but the Rerank table has no Status column, no approved origin states its accessType, and its single "Modalities" column does not separate input from output. Three required fields unstated.
    • huggingface.co/CohereLabs states the Apache 2.0 licence and would close the osiApproved gap. It is explicitly not an approved origin for this creator, and a run may never approve one, so it was not fetched and not cited.
    • Blog publication timestamps on cohere.com/blog were deliberately not used as release dates. The amended provenance rubric states that a page’s publication timestamp is not a statement of a release date, and this run held to that even though it left the run with no usable dates.
    • Bedrock, SageMaker, Azure AI Foundry and Oracle OCI identifiers appear throughout the Cohere docs and were never used to fill a release’s accessType. They state what a serving platform offers, not how the creator releases the model.

    What was evaluated

    Reviewers
    3
    Verdicts cast
    18
    Accepted by panel
    2
    Rejected by panel
    4
    Deterministic gates and required checks — 3 of 5 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one.
    CheckScopeExitResult
    gate-evidence1 bundle, 6 claims1FailFailed on exactly the four claims the panel put below the unanimous 3-of-3 long-tail threshold: co6-source-command-a-plus-card-add, co6-publisher-cohere-add, co6-organization-cohere-add and co6-family-command-a-plus-add. It reported no evidence-format failure of any kind — every quote was present, verbatim against the hashed bytes, and over the 24-character minimum. The failure is the review threshold, not the evidence. Nothing was applied, so nothing shipped past it.
    gate-source-approval1 bundle, 12 citations0PassAnchored on the committed dataset at merge-base 73f92f3 with refs/remotes/origin/main, computed by the gate rather than supplied, so the run could not approve its own source. 24 already-trusted origins; all 12 citations landed on the two origins PR #445 approved for Cohere. No new host was admitted and none was proposed.
    gate-datasetworking tree0PassThe dataset is unchanged by this run and remains internally coherent: 109 sources, 15 publishers, 7 organizations, 26 families, 49 releases. No orphaned source, no dangling reference, no release predating its family.
    npm run validateweb/0PassFull test suite and Astro/TypeScript diagnostics green — 985 tests across 37 files, 0 errors — including the refresh-log schema test this new entry must satisfy, and validate.test.ts, whose unreferenced-source rule is the reason the two accepted sources were not applied.
    gate-scopebranch vs merge-base 73f92f31FailReports web/src/data/refresh-runs.json as outOfClass, which is correct and expected rather than a problem to work around: the ADR 0003 qualifying class is the eleven dataset documents raw.ts composes, and the refresh log is deliberately not one of them, so any run that logs itself leaves the class by construction. The gate was not modified, the file was not added to ALLOWED_PATHS, and the log entry was not dropped to stay in class. The consequence is simply that this change is outside ADR 0003 and merges the ordinary way, with a human.

    Posted 0 edits

    Nothing reached the dataset. No branch, no commit, no pull request.

    Not posted 6 items

    Rejected by the review panel

    • cohere-command-a-plus-familyRejected 0-of-3 — the only claim in either Cohere run to be refused by all three rubrics independently, and the clearest finding here. Cohere’s own documentation, quoted in the claim, calls Command A+ "the last model in the Command A family". Provenance held that the Live status and release date are quoted about the dated command-a-plus-05-2026 release and so add unsupported scope when applied to a family. Consistency held that filing a model as a family breaks lineage coherence and contradicts the claim’s own description. Editorial held that the profile’s naming rule makes the dated identifier the release and the dateless form its alias, so elevating one model to a family nested inside the family Cohere actually names mis-levels a release. This is a modelling defect the run got wrong, not a sourcing gap, and it is recorded rather than patched: correcting it would need a second panel on a claim whose release dependency had already failed.Blocked by co6-family-command-a-plus-add
    • cohere-organization-and-publisherBoth accepted by consistency and editorial and both rejected by provenance, 2-of-3, short of the unanimous long-tail bar. The objection is precisely the one that sank the same claim in run 2026-08-27-4f1c9e: no quote states the schema classification "company", and "Name of the model provider: Cohere Inc." supplies a provider name rather than a classification. PR #459 legitimised selecting a vocabulary member that a quoted term denotes; it did not create a quoted term here, and the rubric’s "the source states nothing on the point" rejection still applies. That is the amendment working as written, not failing.Blocked by co6-organization-cohere-addco6-publisher-cohere-add
    • cohere-command-a-plus-model-card-sourceAccepted by consistency and editorial, rejected by provenance 2-of-3. The page heading "Cohere’s Command A+ Model" is verbatim in the hashed bytes, but provenance held it identifies a page about the model without stating that the page is a structured model card, so the source type "model-card" rested on an unsupported classification. Editorial reached the opposite conclusion on the same page, noting it carries provider, release date, modalities and model size under explicit headings. The disagreement is recorded rather than resolved by the chair, who does not vote.Blocked by co6-source-command-a-plus-card-add

    Accepted by the panel, then dropped

    • cohere-accepted-sourcesThe two documentation sources — the Models Overview page and The Cohere Platform page — were the only claims to clear the unanimous bar, accepted 3-of-3 by all three rubrics. They were still not applied. Every claim that would have cited them was rejected, so applying them would add sources that no record references, and check-bundle-pairing refuses exactly this: "an unreferenced source is dead provenance", a rule validate.test.ts enforces in the suite. Applying two records that reference nothing to report a non-empty diff would be the pipeline working against its own purpose.Blocked by cohere-organization-and-publishercohere-command-a-plus-family

    Verification date deliberately held back

    • cohere-command-a-plus-05-2026The strongest release candidate Cohere publishes, and it was never put to the panel because it cannot be assembled honestly. Its model card states both "Cohere Deployment" and "Open Source Deployment: Command A+ is available under an Apache 2.0 License on Hugging Face", which is accessType "both"; releaseSchema.superRefine then requires a license object, and licenseSchema requires osiApproved as a non-optional boolean. No approved Cohere origin states OSI approval, and schema.ts states in terms that downloadable weights and OSI approval are separate claims where neither implies the other, so it cannot be derived from the Apache 2.0 quote. huggingface.co/CohereLabs would settle it and is explicitly not an approved origin; a run may never approve one. Open issue #461. Withheld rather than filled — a required field is not a licence to guess.Blocked by web/src/data/schema.tstools/updater/profiles/origins/cohere.json
    • cohere-remaining-model-table-rowsCommand A, Command R, Command R7B, Command R+, Command A Translate, Command A Reasoning, Command A Vision, Aya Expanse, Aya Vision, Tiny Aya, Embed, Parse, Cohere Transcribe, rerank-v3.5 and rerank-v4.0 — the full breadth this run was asked to claim. Each needs a family, and familySchema.firstReleaseDate is a day-precision isoDate with no datePrecision companion even though partialDate exists in the same file and is used for evaluationDate, release-event dates and usage windows. Only two day-precision Cohere release dates exist anywhere on an approved origin, and both belong to explicitly non-first family members, so every family here would need a day the source never gave. rerank-v3.5 additionally lacks a stated status and accessType and its "Modalities" column does not separate input from output. Breadth was the goal and the schema, not the sources, is what refused it.Blocked by web/src/data/schema.ts

    What this run does not prove

    • Nothing was applied to the dataset. The Others branch is unchanged and still renders no creators; this entry is the only edit in the pull request.
    • PR #459 did what it set out to do, and this run is evidence for that rather than against it. The provenance reviewer wrote that "mapping Live to current would be a permitted recording step at the same entity level" — the precise objection that ended the previous Cohere attempt, now conceded. It rejected on entity-level scope instead, which is a different and untouched part of the rubric. Reading this run as PR #459 having failed would be the wrong lesson.
    • The two remaining blockers are schema shape, not evidence and not review. A day-precision firstReleaseDate and a required osiApproved are both fields the schema demands and the approved sources do not state, and no amount of better scouting reaches them. Until one of the follow-ups below lands, Cohere is unpublishable from its currently approved origins however good the evidence is.
    • The 0-of-3 family rejection is a fault in this run’s own modelling, and it is reported as one. Command A+ is a model in the Command A family and filing it as a family was wrong; all three rubrics caught it independently, which is the panel doing exactly what it exists to do.
    • Since PR #463, an organization with at least one family but no releases now renders in the /providers A-Z directory rather than nowhere, so the reason the previous run gave for dropping accepted organizations no longer holds in full. It did not change this run’s outcome — the organization and family claims were rejected on their own merits — but the constraint recorded in run 2026-08-27-4f1c9e is now out of date.
    • A withheld creator is a real outcome of the policy, not a failure of the sources. Everything above rests on facts Cohere publishes plainly on origins this repository already approved.

    Follow-ups — proposed, not fixed

    • Make familySchema.firstReleaseDate accept partialDate, the same YYYY(-MM(-DD)) shape already used in this file for evaluationDate, release-event dates and usage windows, and pair it with a datePrecision companion as releases already have. It is the single change that would let this run publish: Cohere states month precision for several families and day precision for almost none, and a day-precision-only field forces a fabricated -01 day or nothing at all.
    • Resolve issue #461. licenseSchema requires osiApproved whenever a license object is present, and releaseSchema requires a license object whenever accessType is "both" or "open-weight", so a well-sourced Apache 2.0 release is unrecordable unless an origin stating OSI approval is approved. Either make osiApproved optional, or approve huggingface.co/CohereLabs for this creator. Today the only honest option is to withhold the release entirely.
    • Update the constraint recorded in run 2026-08-27-4f1c9e and in issue #441: after PR #463, model-tree.ts is no longer the only renderer, and provider-directory.ts places an organization with a family and no releases into the A-Z directory. An organization-and-family record is no longer invisible, so "renders nowhere" is no longer a sufficient reason on its own to drop accepted claims.
    • Decide whether a source’s type field — "model-card" against "official-docs" — is a claim needing its own verbatim support or a cataloguing decision. Provenance and editorial split on exactly this for the Command A+ page, and the same split will recur on every model card the project ever records.
    • Consider whether the organization type "company" needs a quote at all. It is the second Cohere run stopped partly by the absence of a source sentence classifying a company as a company, and no approved origin is ever likely to contain one. If the field is a cataloguing decision rather than a claim about the world, the rubric should say so; if it is a claim, organizations may simply be unrecordable for most creators.
    • docs.cohere.com/changelog/ and docs.cohere.com/v2/docs/ sit outside the allowed_paths in the approved origins catalogue and were not fetched. The changelog is the one place a creator normally states dated releases in prose, so approving that path is likely worth more to this creator than any further scouting of /docs/.
  4. 2026-08-27-4f1c9e

    Long-tail refresh 2026-08-27 — Microsoft, Mistral AI, xAI, Cohere

    Scope requested: Four creators with no reviewed profile — Microsoft, Mistral AI, xAI and Cohere — under the long-tail unanimous 3-of-3 policy, restricted to the eight origins PR #445 approved for them.Ran, changed nothing0 edits posted · 7 items withheld

    Researched four long-tail creators to populate the Others branch of the tree and published none of them. Every fact came from the origins PR #445 approved for these creators, fetched and hashed rather than read from search results. Microsoft and xAI were withheld before review: no approved-origin page states an access type or modality set for MAI-Thinking-1, and both are required fields, while x.ai now self-describes as "SpaceXAI LLC" on both of its approved origins and no approved page names a Grok family. Mistral AI and Cohere reached a review panel across five rounds. A final round re-scouted Cohere against a model reference on an already-approved origin that carries an explicit Status column and a machine-readable specification table, which answered the earlier objections about status, categories and the creator relationship; the resulting organization and release claims were accepted by the consistency and editorial rubrics but not by provenance, so they did not reach the unanimous 3-of-3 long-tail bar. Because model-tree.ts renders a creator only when it has at least one release, the organizations and families that were accepted would have added records that render nowhere, so they were dropped rather than applied. The dataset is unchanged; this entry is the only edit.

    1. PreflightRanClean tree rebased onto main at 843f2d4. main moved three times during this run — f665967 to 3e6f3e3 to 7fd5ea2 to 843f2d4 — and every gate below was recomputed against the final anchor rather than carried forward, because a gate result binds to the commit it was computed against. drydock is not on PATH, so the manual posture applied. ADR 0003 and all four gate scripts were read directly rather than from a summary; gate-scope ALLOWED_PATHS was re-read after #448 grew the qualifying class from 9 files to 11.
    2. ScoutRan28 pages fetched over real HTTP from approved origins only and hashed with sha256; every quote was verified verbatim against those bytes before any claim was written, and a mechanical check rejected three round-4 quotes that had been copied from a stripped-text rendering rather than from the hashed HTML. x.ai/news returned a Cloudflare 403 and was never readable. docs.mistral.ai/getting-started/models/models_overview/ redirected outside that origin’s allowed_paths, so it was fetched but deliberately not cited, leaving Mistral without a context-window figure. The round-5 Cohere re-scout admitted no new host and no new page: it re-read bytes already fetched and already proposed as a source, and quoted the parts of them that state the facts under review.
    3. ReviewRanFive blind panel rounds, three rubrics each, run on three different model families with the rubric-to-family assignment rotated every round so no rubric kept one family. Reviewers never saw the scout’s reasoning, each other’s verdicts, or the running tally. No threshold was lowered and no reviewer was re-run on an unchanged claim; each corrected claim was given a new id and put to a fresh panel, and claims already accepted unanimously kept their original verdicts rather than being re-reviewed. The run stopped after round 5 rather than re-panelling again: the surviving provenance objections are about how much of a dataset record a quote must literally state, which no further evidence can resolve, and re-sampling reviewers until three agree is the failure the panel exists to prevent.
    4. GatesRangate-evidence and gate-source-approval were run before anything was applied, anchored on the committed dataset at merge-base 843f2d4 so the run could not approve its own writes, and re-run at that anchor after the round-5 bundle was written because a gate result binds to the commit it was computed against. gate-dataset, npm run validate and gate-scope were run afterwards. gate-evidence failed exactly as the panel did, naming the sub-threshold claims and nothing else.
    5. PublishRanAn ordinary pull request carrying this log entry and no dataset change. Auto-merge was not requested and not eligible: refresh-runs.json is outside the gate-scope qualifying class, so ADR 0003 does not cover this change and a human merges it.
    6. DeployNot applicableNo dataset change to deploy. The rendered tree is byte-identical to the one already on main; only the /refresh/ page gains this entry.

    What was found

    Scouts
    4
    Pages fetched and hashed
    28
    Claims proposed
    12
    Claims per creator bundle, with the review threshold its profile set
    CreatorPolicyThresholdClaims
    mistral-ailong-tail3-of-35
    coherelong-tail3-of-37
    What those claims proposed to do
    KindCountEffect
    Add12Would have added two organizations, two families, two releases, two publishers and four sources. None was applied.

    Not covered

    • Microsoft produced no bundle that reached final adjudication: MAI-Thinking-1 has no approved-origin statement of accessType or modalities, and both are required by releaseSchema.
    • xAI produced no bundle that reached final adjudication: the SpaceXAI/xAI naming conflict is unresolved across both approved origins and no approved page names a Grok family.
    • azure.microsoft.com, learn.microsoft.com and api-docs.deepseek.com were deliberately not cited. Microsoft as a serving platform and product vendor is a different entity from Microsoft as a model creator.
    • No context-window figure for any Mistral release: the only page stating one redirects outside the approved allowed_paths.
    • No approved Cohere origin states OSI approval for the Apache 2.0 licence, and licenseSchema requires osiApproved whenever a license object is present, so the whole object was omitted rather than guessed. opensource.org is not an approved origin and was not cited to close the gap.
    • No approved Cohere origin states the Command A family’s first release date in prose; the only available date is the publication timestamp of the launch announcement.

    What was evaluated

    Reviewers
    3
    Verdicts cast
    45
    Accepted by panel
    8
    Rejected by panel
    4
    Deterministic gates and required checks — 3 of 5 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one.
    CheckScopeExitResult
    gate-evidence2 bundles, 12 claims1FailFailed on exactly the four claims the panel put below the unanimous 3-of-3 long-tail threshold: mi4-release-large-3-add, co5-organization-add, co5-family-command-a-add and co5-release-command-a-plus-add. Every quote was present, verbatim against the hashed bytes and over the 24-character minimum; the failure is the review threshold, not the evidence format. Nothing was applied, so nothing shipped past it.
    gate-source-approval2 bundles, 57 citations0PassAnchored on the committed dataset at merge-base 843f2d4 with refs/remotes/origin/main, computed by the gate rather than supplied, so the run never approved its own source. 24 already-trusted origins; every citation in both bundles, including the enlarged round-5 Cohere bundle, landed on origins the reviewed profile catalogues merged in PR #445 already stand behind. No new host was admitted. Re-run after #448 changed sources.json, after #451 moved main, and again after the round-5 bundle was written, because a gate result binds to its anchor.
    gate-datasetworking tree0PassThe dataset is unchanged by this run and remains internally coherent. No orphaned source, no dangling reference, no release predating its family.
    npm run validateweb/0PassFull test suite and Astro/TypeScript diagnostics green, including the refresh-log schema test that this new entry must satisfy.
    gate-scopebranch vs merge-base 843f2d41FailReports web/src/data/refresh-runs.json as outOfClass, which is correct and expected: the ADR 0003 qualifying class is the 11 dataset documents raw.ts composes, and the refresh log is deliberately not one of them. The gate was not modified and the file was not added to ALLOWED_PATHS; the consequence is simply that this change is outside ADR 0003 and merges the ordinary way, with a human, rather than auto-merging.

    Posted 0 edits

    Nothing reached the dataset. No branch, no commit, no pull request.

    Not posted 7 items

    Rejected by the review panel

    • mistral-large-3The only Mistral release claim, rejected by the provenance rubric in two consecutive rounds on different grounds and by different model families. Round 3 objected to an intendedUse clause asserting the model is "demanding to self-host", which no quote stated; that clause was deleted. Round 4 then objected that the quotes do not state the status, category or open-weight access type. Consistency and editorial accepted it both times. Under the unanimous 3-of-3 long-tail bar it did not carry.Blocked by mi3-release-large-3-addmi4-release-large-3-add
    • cohere-command-a-plus-05-2026The only Cohere release claim, put to three separate panels on corrected inputs and never unanimous. Round 3 caught a real schema error — contextWindow given as an object where releaseSchema requires a positive integer — which was corrected. Round 4 provenance objected that the quotes state neither osiApproved nor the dataset categories nor the intendedUse guidance. Round 5 rewrote the claim against Cohere’s own specification table, dropped the licence object entirely rather than guess osiApproved, and rebuilt categories, status, context limits and intendedUse from quoted text; consistency and editorial both accepted it, and editorial specifically credited the dated canonical name and the exclusion of Azure Foundry ids and North product metrics. Provenance still rejected, holding that mapping the table’s "Live" to the schema’s "current", normalising "128K" to 128000, and reading a release date off the announcement’s datePublished are each inferences the quotes do not contain.Blocked by co3-release-command-a-plus-addco4-release-command-a-plus-addco5-release-command-a-plus-add
    • cohere-command-a-familyRejected by two rubrics in round 5, and the only claim in this run to draw a substantive objection from editorial. The family record borrowed its description from the Command A release row and carried the modality of a later sibling: it listed multimodal-generalist because Command A+ accepts images, while the same table lists Command A itself as text-only, and it adopted the marketing word "excelling" from release copy into family-level prose. Editorial held that family and release were not kept distinct. That is a real modelling defect rather than a sourcing gap, and it is recorded here rather than patched, because correcting it would need a fourth panel on a claim whose release dependency had already failed.Blocked by co5-family-command-a-add

    Accepted by the panel, then dropped

    • mistral-ai-organization-and-familyThe Mistral AI organization, the Mistral 3 family, the publisher and the announcement source were all accepted unanimously, 3-of-3. They were still not applied: model-tree.ts places a creator in a branch only when it has at least one release, and drops families holding zero releases, so an organization and family without an accepted release would have added records that render nowhere and read as data while showing nothing. This is the same defect as issue #441.Blocked by mistral-large-3
    • cohere-organization-and-sourcesThe Cohere publisher and three sources were accepted unanimously across rounds 3 and 4 and were carried into round 5 with their original verdicts rather than re-reviewed. The organization claim was rewritten in round 5 to rest on quotes that state the creator relationship directly and was accepted by consistency and editorial, but provenance rejected it because no quote classifies Cohere as a "company" and the sentence naming the Command models resolves its subject only from surrounding context. Nothing was applied: with no accepted release the creator renders in neither branch, and applying the sources alone would leave them orphaned, which the validator forbids.Blocked by cohere-command-a-plus-05-2026

    Verification date deliberately held back

    • microsoft-mai-thinking-1No approved-origin page states an accessType or an input/output modality set for MAI-Thinking-1. Both are required fields on releaseSchema. The only availability statement found — "available in public preview on Microsoft Foundry" — describes a serving platform, not the model's access type, and treating it as one would collapse creator and platform into a single entity. Recorded as unknown rather than inferred.Blocked by tools/updater/profiles/origins/microsoft.jsonweb/src/data/schema.ts

    Sources conflict, so no value changed

    • xai-grokx.ai and docs.x.ai both now self-describe as "SpaceXAI LLC" — docs.x.ai is titled "SpaceXAI Docs" — so the naming conflict spans every origin approved for this creator rather than being contradicted by one of them. The profile records it as an unresolved ambiguity and it stays explicit rather than being smoothed over. Separately, no approved page names a Grok family, and x.ai/news returns Cloudflare 403.Blocked by tools/updater/profiles/origins/xai.json

    What this run does not prove

    • Nothing was applied to the dataset. The Others branch is unchanged and still renders no creators; this entry is the only edit in the pull request.
    • The counts here describe the final adjudicated bundles — the round-4 Mistral bundle and the round-5 Cohere bundle. Five review rounds ran in total, and the superseded bundles proposed more claims than the 12 counted above; they are published in the pull request body with every verdict and dissent verbatim, not summarised here.
    • A unanimous 3-of-3 bar means a single rubric decides the outcome, and provenance was that rubric in every rejection but one. Three provenance reviewers on different model families rejected these releases across three rounds on shifting grounds, so this is not one model’s idiosyncrasy — but neither is it evidence that the underlying published facts are wrong.
    • Round 5 is the useful control. It answered round 4’s stated objections with better evidence from the same approved origin — an explicit Status column, a specification table giving licence, size, context length and modalities — and the claims still did not carry, because the surviving objections are about the dataset’s own vocabulary rather than about the sources. If mapping a creator’s "Live" to the schema’s "current" is an unsupported inference, then no release record can ever clear a 3-of-3 provenance bar, whatever the evidence.
    • A withheld creator is a real outcome of the policy, not a failure of the sources. Every fact these releases rest on is published by the creator on an approved origin.

    Follow-ups — proposed, not fixed

    • The three rubrics have no shared rule for whether dataset-vocabulary fields — status, categories, accessType, license.osiApproved — are assertions about the world needing their own verbatim quote, or encodings of facts already quoted. Round 5 isolated this as the single blocking question: the evidence objections were answered and the claims still failed. Settle it before the next long-tail attempt, because at present no release record can reliably clear a 3-of-3 bar.
    • Related and narrower: decide whether normalising a quoted "128K" into the integer 128000, and reading a release date from an announcement’s datePublished, count as inference. The committed dataset already does both — openai-gpt-4-1 takes its firstReleaseDate from its announcement — so the standard applied to new claims is currently stricter than the one the existing data was built to.
    • licenseSchema requires osiApproved whenever a license object is present, which makes an otherwise well-sourced Apache 2.0 licence unrecordable unless an origin stating OSI approval is approved. Either make osiApproved optional or approve such an origin; today the only honest option is to drop the licence entirely, losing a fact the creator publishes plainly.
    • Microsoft cannot be recorded as a model creator from its currently approved origins: microsoft.ai and the bounded www.microsoft.com research-blog path state no access type or modality set for MAI-Thinking-1. Either an additional creator-side origin is approved, or Microsoft stays absent; azure.microsoft.com and learn.microsoft.com are the wrong entity and must not be used to close the gap.
    • x.ai self-describes as "SpaceXAI LLC" across both approved origins while the profile and the wider world still say xAI. The conflict is recorded and unresolved, and until it is resolved the creator has no defensible canonical name.
    • x.ai/news returns a Cloudflare 403 to this fetcher, so xAI release announcements were never readable from an approved origin.
    • No Mistral context window is recorded anywhere in the dataset: the only page stating one redirects outside the approved allowed_paths for docs.mistral.ai.