What each refresh found, and what it did not publish.
ModelTree's dataset is refreshed by agents against primary sources, reviewed by an independent three-rubric panel, and gated deterministically. Every run is recorded here in full — including the runs that published nothing.
A run's working state is never committed. This page transcribes the durable record: the pull request body and the summary issue, both linked from each entry.
- Runs recorded
- 23
- Pages fetched
- 2,128
- Claims proposed
- 1,352
- Edits published
- 611
- Items withheld
- 439
Showing 11–20 of 23 runs
2026-09-01-c3f81a
Long-tail depth 2026-09-01 — Cohere, TII, NVIDIA
Scope requested: Three creators already in the catalogue, each holding exactly one family — Cohere, TII and NVIDIA — targeted at one added family plus at least one release each. Deliberately excluded by the coordinating brief and untouched here: ai2-molmo and the other family ids held by in-flight issue #740, the four held by #751, and Stability AI, which is reserved to #754. No id proposed by this run collides with any of those.Published4 edits posted · 10 items withheldResearched a second model family for three creators that each hold exactly one — Cohere, TII and NVIDIA — and published one of the three. Fourteen claims were proposed across three bundles, all of kind add, carrying 44 evidence entries; every quote was machine-checked as a contiguous verbatim substring of the exact bytes its contentHash names, with positive and negative controls on the harness, and all 44 passed. The panel accepted eleven claims unanimously and blocked three, and the three blocks are what decided the run. The provenance rubric rejected the Falcon 3 release because its inputModalities and outputModalities were mapped to text from no quote at all: the card states no modality anywhere, so there was nothing to transcribe and the remedy of attaching a quote does not exist for that page. The editorial rubric rejected the Cohere model card and the Cohere release on entity attribution, holding that tools/updater/profiles/origins/cohere.json records huggingface.co/CohereLabs in deferred_origins as not approved, that resolving whether that org is Cohere's own voice is a human catalogue decision, and that this run did not take it. Because releaseSchema requires both modality fields and accessType with no unknown member, neither rejection could be repaired by dropping an optional field. Each rejection removed the creator's only release, which would have left an empty family that validateDataset refuses, so Cohere and TII were withheld whole — seven further claims that the panel had accepted unanimously were dropped after acceptance rather than applied and orphaned, per the scout contract's pairing rule. NVIDIA cleared unanimously on all four claims and is the run's entire output: the Nemotron Nano 2 family and its Nemotron Nano 9B v2 release, with two supporting sources. Both records carry status unknown, which under ADR 0008 records that the card states no lifecycle state and is a sourced value rather than a fallback; the family and its only release agree at unknown, so the run adds no family whose status outruns its releases. The release's accessType open-weight rests on the ADR 0009 composite — a Hugging Face Hub API record reporting the repository ungated and enabled with its safetensors shards enumerated — and not on the licence name, which states nothing about whether weights can be downloaded.
- PreflightRanRe-measured the base from the committed data rather than quoting the brief. The worktree started at trunk but trunk moved during planning, so with zero commits of its own and no divergence the branch was fast-forwarded to refs/remotes/origin/main at b435c69056; that fast-forward is recorded as an assumption in the run summary. At that anchor the dataset held 71 families, 100 releases and 239 sources, schema.ts carried unknown in lifecycleStatus, and none of cohere-command-r, tii-falcon-3 or nvidia-nemotron-nano-2 existed. The reviewed-profile set at tools/updater/profiles holds seven creators and none of these three, so gate-evidence derives long-tail for all three bundles and the declared policy matches the derived one.
- ScoutRanStored 18 page bodies and cut every quote out of those exact bytes between two literal markers, so a non-verbatim quote throws at build time rather than being emitted. Fourteen claims across three bundles, 44 evidence entries, all retrieval fetch and no search snippets. Two fields confirmed optional in schema.ts were dropped rather than sourced: maximumOutput on the Cohere release and parameters.totalBillions on both the Falcon 3 and Nemotron releases, where the size appears only in the model name. The Cohere family date came from a Hub createdAt with dateBasis platform-repository-created, because firstReleaseDate is required and Cohere's own docs date only the August 2024 update rather than the line.
- ReviewRanThree reviewers, one per rubric, run in parallel and independently: none saw the scout's reasoning, another reviewer's verdict, or any statement of the threshold or the running tally. 42 verdicts over 14 claims. Each was pointed at web/src/data/schema.ts as the definition of the controlled vocabulary, because .github/skills/modeltree-review/SKILL.md has drifted from ADR 0008 and still states that no field has an unknown member; the drift is recorded as a follow-up rather than edited here, since SKILL.md is outside the qualifying class gate-scope enforces. No reviewer that rejected was re-run. The editorial reviewer was asked once to re-transcribe rationales lost to output truncation, with its votes stated as final and closed to revision; it returned wording only.
- GatesRangate-evidence and gate-source-approval were run before any dataset file was touched. gate-evidence exited 0 on the published NVIDIA bundle and, run over the full Cohere and TII bundles, exited 1 naming exactly the three claims the panel had blocked — an independent confirmation of the withholding rather than the run's own word for it. gate-dataset and gate-scope both exited 0 after the change was applied. gate-ledger is the one gate this branch cannot pass by itself and the reason is structural, not a defect in the data: a published entry must name its pull request, and a dock does not open one.
- PublishNot runDeliberately not run. The dock's boundary ends at a reviewable commit: it does not open a pull request, does not merge, does not rebase and does not gate its own work. The dataset change is committed to the branch and this entry is handed to the publishing step to commit second, once the pull request exists to be named — the same two-commit shape run 2026-09-01-b41087 used. Landing remains GitHub's once CI is green.
What was found
- Scouts
- 3
- Pages fetched and hashed
- 18
- Claims proposed
- 14
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims cohere long-tail 3-of-3 5 tii long-tail 3-of-3 5 nvidia long-tail 3-of-3 4 What those claims proposed to do Kind Count Effect Add 14 Four applied to the dataset; ten withheld, three of them blocked by the panel and seven dropped after unanimous acceptance to avoid orphaning them. Not covered
- No second Falcon 3 or Command R variant was scouted after the first was blocked. Re-scouting the same creator for a page that scores better is vote-rigging however it is framed, and the Falcon 3 block in particular is a property of the card rather than of the variant chosen.
- The build.nvidia.com API Catalog channel the Nemotron card names was not read, so the release records accessType open-weight rather than both. That under-claims rather than over-claims.
- Link health was not swept and the second Python interpreter was not exercised; ci-preflight names both as outside what it covers.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 42
- Accepted by panel
- 11
- Rejected by panel
- 3
Deterministic gates and required checks — 6 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidencethe published NVIDIA bundle at .modeltree-refresh/runs/2026-09-01-c3f81a/published/nvidia.claims.json 0 Pass4 claims admissible under the long-tail policy, all 4 applying to the dataset. The policy was derived from the reviewed-profile set rather than taken from the bundle, and the derived long-tail matched the declared long-tail, so the unanimous 3-of-3 threshold is the one actually applied. gate-evidencenot requiredthe full Cohere and TII bundles, run to confirm the withholding mechanically 1 FailExit 1 on both, naming cohere-source-command-r-08-2024-card-add, cohere-release-command-r-08-2024-add and tii-release-falcon-3-7b-instruct-add as marked add but reaching only 2 of 3 required accepts. This is the expected and wanted result: it is the gate independently identifying the same three claims the panel blocked, and none of them is in the commit. Recorded as not required because these bundles are not what the run published. gate-source-approvalthe published NVIDIA bundle against the approved origin set at merge-base b435c69056 0 Pass13 citations rest on approved sources: 1 inherited from the dataset at the merge base and 2 proposed on already-trusted origins. No origin was widened and none needed to be. gate-datasetweb/src/data after the four claims were applied 0 PassAll gates passed over 545 records. In particular the new family carries a release, so the empty-family refusal that withheld Cohere and TII does not bite here, and the family's status unknown agrees with its only release. gate-scopethe branch against its computed merge-base with refs/remotes/origin/main 0 PassOnly families.json, releases.json and sources.json changed, all three inside the ADR 0003 qualifying class. No test, workflow, skill or document was touched: the SKILL.md drift found during review was left alone precisely because editing it would have taken the change out of class. npm run validateweb/, tests plus Astro and TypeScript diagnostics 0 PassGreen, with a green baseline measured on the unmodified base first so that any failure would have been attributable to this change. No pinned-enumeration test needed editing: the dataset-count assertions are all lower bounds. ci-preflightthe repository root, selecting the pull-request checks this branch's diff triggers 0 PassRun from the repository root because npm run validate reads only web/ and would not have exercised the checks a diff outside it triggers. This diff is dataset-only, so the selection is narrow. Posted 4 edits
4 edits across 3 documents, a net change of 4 records.
Dataset documents this run changed Document Before After What changed sources.json239 241 Two sources added: the Nemotron Nano 9B v2 model card and its Hugging Face Hub API record. families.json71 72 One family added: nvidia-nemotron-nano-2. releases.json100 101 One release added: nvidia-nemotron-nano-9b-v2. Records added
- Nemotron Nano 2in the treefamilies
nvidia-nemotron-nano-2New NVIDIA family, kept distinct from the earlier Nemotron-4 line. firstReleaseDate 2025-08-18 is the card's own stated release date and carries no dateBasis; status unknown. - Nemotron Nano 9B v2passportreleases
nvidia-nemotron-nano-9b-v2First release under that family. status unknown, accessType open-weight from the ADR 0009 platform composite, contextWindow 128000, modalities quoted from the card's Input Type(s) and Output Type(s) lines.
Not posted 10 items
Rejected by the review panel
tii-release-falcon-3-7b-instruct-addreleases record tii-falcon-3-7b-instruct for tii. [provenance] inputModalities ['text'] and outputModalities ['text'] are unsupported: the card contains no modality statement of any kind and no modality quote is attached, so mapping text/text is inference from the word LLM rather than a recording step. Every other field on the record — releaseDate 2024-12, status unknown, contextWindow 32000, the licence block with osiApproved false, and accessType open-weight on the ADR 0009 composite — was found properly sourced, which is what makes the single defect decisive: releaseSchema requires both modality fields and neither is optional, so the record cannot be written without them.Blocked byrubric:provenancecohere-source-command-r-08-2024-card-addsources record cohere-command-r-08-2024-model-card. [editorial] The claim attributes the card to publisherId cohere as Cohere's own voice, but tools/updater/profiles/origins/cohere.json records huggingface.co/CohereLabs in deferred_origins as not approved, reasoning that nothing there establishes it is Cohere's own org rather than a name anyone could take, and that resolving it is an organization mapping for a human. The run's own notes conceded that decision was not taken. The provenance rubric reached the opposite conclusion on the strength of an approved docs.cohere.com page naming the CohereLabs org, so this is a genuine split rather than a weak claim, and under the unanimous long-tail bar the split blocks it.Blocked byrubric:editorialprofile:cohere deferred_originscohere-release-command-r-08-2024-addreleases record cohere-command-r-08-2024 for cohere. [editorial] Its creator-authored facts — accessType both with the open-weight half, the CC-BY-NC licence block, parameters.totalBillions 32 and the modality lines — rest solely on the model card whose attribution the same rubric rejected, so accessType both over-attributes open weights to Cohere when only the docs' Live-on-API row is solidly Cohere's voice. It separately found variant 'Standard' to be a tier no source states. The unsourced maximumOutput 4000 the previous run carried was already dropped rather than sourced.Blocked byrubric:editorialclaim:cohere-source-command-r-08-2024-card-add
Accepted by the panel, then dropped
cohere-family-command-r-addfamilies record cohere-command-r. Accepted 3-of-3 on its own merits: firstReleaseDate 2024-03-11 with dateBasis platform-repository-created, status unknown, and sources that do not include the disputed card. Dropped because its only proposed release was blocked, which would leave a family with zero releases — the condition validateDataset refuses with 'family has no releases'. Publishing the family alone would have been incoherent, so family and release stand or fall together.Blocked byclaim:cohere-release-command-r-08-2024-addcohere-source-command-r-docs-addsources record cohere-command-r-docs, accepted 3-of-3 on an approved docs.cohere.com origin. Dropped because the only records citing it were the withheld Cohere family and release; applying it alone would leave an uncited source in the dataset.Blocked byclaim:cohere-family-command-r-addcohere-source-command-r-v01-hub-record-addsources record hugging-face-cohere-command-r-v01-hub-record, accepted 3-of-3 as a platform-authored ADR 0009 record. Dropped for the same orphaning reason: only the withheld Cohere family cited it.Blocked byclaim:cohere-family-command-r-addtii-family-falcon-3-addfamilies record tii-falcon-3. Accepted 3-of-3, with firstReleaseDate 2024-12 taken from the card's creator-stated Model Release Date: December 2024 and status unknown. Dropped because its only proposed release was blocked on modalities, leaving an empty family. This is the closest miss of the run: the family claim itself has no known defect.Blocked byclaim:tii-release-falcon-3-7b-instruct-addtii-source-falcon3-7b-instruct-card-addsources record tii-falcon3-7b-instruct-model-card, accepted 3-of-3. Dropped because both records citing it were withheld.Blocked byclaim:tii-family-falcon-3-addtii-source-falcon3-announcement-addsources record tii-falcon3-announcement, accepted 3-of-3 as an official-announcement in TII's first-person voice. Dropped for orphaning.Blocked byclaim:tii-family-falcon-3-addtii-source-falcon3-7b-instruct-hub-record-addsources record hugging-face-falcon3-7b-instruct-hub-record, accepted 3-of-3 as a platform-authored ADR 0009 record. Dropped for orphaning.Blocked byclaim:tii-release-falcon-3-7b-instruct-add
What this run does not prove
- One creator of three landed. The issue's exactly-one-family count moves by one, not by three, and this run does not by itself satisfy #689.
- status unknown is honest but it is not informative. Three of the dataset's families now sit at unknown plus this one, and a reader looking for whether Nemotron Nano 2 is current will not find that here, because NVIDIA's card does not say.
- accessType open-weight on the Nemotron release rests on a platform record reporting an ungated repository with safetensors shards, not on a statement by NVIDIA that weights are downloadable. That is the ADR 0009 composite and the panel accepted it unanimously, but it is a weaker basis than a creator sentence saying so, and it can go stale if the repository is later gated.
- The Cohere block is an unresolved human decision, not a finding that Cohere's card is wrong. Two rubrics reached opposite conclusions on whether huggingface.co/CohereLabs is Cohere's voice, and until a human updates that catalogue entry the question stays open and the evidence stays unused.
- The panel is three instances of one model family reading the same pages, so it buys independence of reasoning and not independence of training. A page that is itself wrong could carry all three.
- gate-ledger cannot pass on this branch as the dock leaves it, because this entry is not committed and cannot be until the pull request it must name exists.
Follow-ups — proposed, not fixed
- tools/updater/profiles/origins/cohere.json defers huggingface.co/CohereLabs as not approved on the ground that nothing establishes it as Cohere's own org. This run found first-party evidence bearing on exactly that: docs.cohere.com, an approved Cohere origin, links to a CohereLabs collection and describes it as where Cohere publishes open-weight models on Hugging Face. A human should decide whether that settles the mapping; the Cohere family and release are fully scouted and re-runnable the moment it does.
- .github/skills/modeltree-review/SKILL.md states that none of the controlled-vocabulary fields has an unknown member and lists five members of lifecycleStatus. ADR 0008 added a sixth, unknown, and web/src/data/schema.ts carries it. The rubric text and the schema now disagree on a point that decides claims, and reviewers in this run had to be pointed at the schema to resolve it. Editing SKILL.md was deliberately not done here because it is outside the ADR 0003 qualifying class and would have taken the dataset change out of class.
- The Falcon3 model card states no modality anywhere, which blocked an otherwise clean release. Whether TII states modalities on another Falcon 3 surface is unknown; a future run would need a different page, not a different reading of this one.
2026-09-01-b41087
Long-tail sweep: all 40 creators re-verified
Scope requested: All 40 creators in organizations.json. Seven carry a reviewed profile in tools/updater/profiles and were judged at 2-of-3 — alibaba-cloud, amazon, anthropic, google-deepmind, meta, microsoft, openai — and the other 33 were judged at the unanimous 3-of-3 long-tail bar. Predominantly re-verification: asking whether recorded facts still hold, and advancing the dates that say when they were last checked. Discovery was attempted only where a creator index page made it cheap, which surfaced exactly one unrecorded release. The model-fit guidance layer was deliberately left alone: all seven statements were re-verified on 2026-08-31, so the five records they pin needed nothing this run.Published179 edits posted · 5 items withheldAn agent-run, source-backed refresh under ADR 0003, and the breadth run the previous entry asked for. Run 2026-08-31-ae0342 traded breadth for depth, scouted six pilot creators and recorded the untouched long tail as the largest gap in its own entry; this run scouted every one of the 40 creators in organizations.json against 233 pages fetched and hashed during the run, 7 at the 2-of-3 pilot threshold and 33 at the unanimous 3-of-3 long-tail threshold. 202 claims were proposed and three rubrics voted blind on every one, casting 606 verdicts. 199 met their threshold, 197 were applied, and 5 were withheld with their rationales published. Not one claim was carried over a dissent, which is not a sign the rubrics went easy but a fact about this run being almost entirely re-verification, where there is little to disagree about: all four rejections landed on the three claims that attempted something harder than confirming a date. Run 2026-08-30-c0b6e9 also recorded zero dissents, so this is not a first. 179 field edits landed across three documents — 89 source lastCheckedDate values and 90 record verifiedAt values — and no record was added or removed. The one genuinely new release found this run, Gemini 2.5 Flash-Lite, was refused by the panel and is not in the dataset. No human reviewed any claim, verdict or edit.
- PreflightRanTree clean, gh authenticated, and no open pull request from a previous refresh when the run started. Run id 2026-09-01-b41087, with the run directory under .modeltree-refresh/ confirmed git-ignored before anything was written to it. Both npm and drydock were probed in both shim forms before either was relied on or ruled out: the bare names resolve to PowerShell shims this machine’s execution policy refuses, while npm.cmd reports 11.9.0 and drydock.cmd reports 0.1.0. Both are installed and blocked in bare form rather than absent, and the execution policy was not changed to work around it — that would be a machine-wide security change made to satisfy a probe. Dependencies were installed with npm ci after npm run validate failed on a missing vitest; package-lock.json is unmodified.
- ScoutRan240 URLs were attempted and 233 returned 200 and were hashed; the 7 failures are recorded under notCovered rather than retried until they looked like successes. Every quote was cut as a byte-exact slice of a page body saved during this run and asserted as a substring of those bytes by the builder, which throws rather than emit an unverifiable pair — so no claim here rests on recollection or on a search snippet. Discovery findings were genuinely mixed and both directions are recorded: OpenAI’s and Anthropic’s model catalogues list nothing the dataset does not already carry, which is a real negative result rather than an unchecked one, while Google’s listed Gemini 2.5 Flash-Lite, which the dataset lacks.
- ReviewRanThree rubrics — provenance, consistency, editorial — voted independently on all 202 claims across 9 reviewer invocations, each in its own context and each given only its own rubric text, the claim, the evidence and the relevant dataset slice. No reviewer saw another rubric’s verdicts, the running tally, the threshold, or the scout’s reasoning. 606 verdicts were cast, every one with a rationale, and all are published on the pull request. Four rejections were returned and every one was left to stand: no claim was edited and re-reviewed to chase a verdict, no reviewer was re-run, and the coordinating agent cast no vote and overruled none — including the one rejection whose stated premise is demonstrably wrong, which is recorded in the caveats and filed as a follow-up instead.
- GatesRangate-evidence and gate-source-approval ran across all 40 gated bundles before anything was applied, since gate-source-approval anchors on the committed dataset and running it afterwards would be asking the run’s own writes whether the run’s own sources are trustworthy. gate-dataset, npm run validate, gate-scope, gate-ledger and ci-preflight.mjs ran afterwards. origin/main moved during the review stage, so the branch was re-anchored from d9c5c403 onto 329a719f and both pre-apply gates were re-run against the new anchor rather than carried over. No gate was skipped, forced, or re-run to obtain a different answer, no threshold was lowered, and nothing was pushed to main.
- PublishRanOpened as pull request #752 carrying the full evidence trail — every claim, every quote with the SHA-256 of the bytes it was cut from, and all 606 rationales. The trail is roughly 260 KB against GitHub’s 65 536-character body limit, so the body carries the summary, the gate results, the source-approval block and all five withheld claims in full, and the per-claim detail follows as five comments on the same pull request rather than being summarised away. Merged by GitHub via --auto --squash once web-ci went green; the agent did not merge and did not use --admin. The data and this entry are two commits on the branch and one commit on main, so a bad run is still one revert.
- DeployNot runNot yet run when this entry was written, and recorded that way rather than predicted. The entry necessarily ships in the commit whose merge triggers pages.yml, so it cannot report the outcome of its own deploy without guessing. The run waits for the deploy after the merge, checks it against the merge SHA, and reverts by pull request if it failed; the result is reported in the run’s summary issue, which is where a reader should look to close this line.
What was found
- Scouts
- 40
- Pages fetched and hashed
- 233
- Claims proposed
- 202
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims 01-ai long-tail 3-of-3 6 ai-singapore long-tail 3-of-3 4 ai2 long-tail 3-of-3 5 ai21-labs long-tail 3-of-3 4 aleph-alpha long-tail 3-of-3 4 alibaba-cloud pilot 2-of-3 6 amazon pilot 2-of-3 5 anthropic pilot 2-of-3 6 apple long-tail 3-of-3 4 baidu long-tail 3-of-3 5 bytedance-seed long-tail 3-of-3 4 cohere long-tail 3-of-3 5 databricks long-tail 3-of-3 4 deepseek long-tail 3-of-3 6 eleutherai long-tail 3-of-3 6 google-deepmind pilot 2-of-3 9 hugging-face long-tail 3-of-3 5 ibm long-tail 3-of-3 6 lg-ai-research long-tail 3-of-3 4 liquid-ai long-tail 3-of-3 4 meta pilot 2-of-3 6 microsoft pilot 2-of-3 6 minimax long-tail 3-of-3 6 mistral-ai long-tail 3-of-3 6 moonshot-ai long-tail 3-of-3 3 naver long-tail 3-of-3 5 nous-research long-tail 3-of-3 5 nvidia long-tail 3-of-3 5 openai pilot 2-of-3 6 reka-ai long-tail 3-of-3 5 sakana-ai long-tail 3-of-3 4 sarvam-ai long-tail 3-of-3 5 snowflake long-tail 3-of-3 5 stability-ai long-tail 3-of-3 5 tencent long-tail 3-of-3 4 tii long-tail 3-of-3 4 upstage long-tail 3-of-3 4 xai long-tail 3-of-3 6 xiaomi long-tail 3-of-3 5 zhipu-ai long-tail 3-of-3 5 What those claims proposed to do Kind Count Effect Change 107 A source page re-read this run, moving its lastCheckedDate to 2026-09-01. All 107 were accepted and applied, moving 89 distinct source records — 18 of the claims re-assert a shared licence page that several creators cite (osi-approved-licenses, osi-license-mit), and a shared source moves once no matter how many bundles confirm it. Unchanged 92 A recorded fact re-read from its primary source and found to still hold. 90 met their threshold and applied, advancing 50 release verifiedAt dates and 40 family verifiedAt dates to 2026-09-01. Two were refused by the panel and withheld, so those two records keep their previous dates. Add 3 One new release, Gemini 2.5 Flash-Lite, with the two source records it would have cited. None reached the dataset: the release was refused 1-of-2 by the panel, and the two sources — accepted on their own merits — fell with it, because a source no record cites is dead provenance that validateDataset refuses to load. Nothing was added or removed this run. Not covered
- ai.meta.com returned HTTP 400 to every request from this runner, across six URLs and both fetch paths tried. That is a status rather than a network failure, so it is a fact about how that host answers this client and not evidence the pages are gone. Four recorded sources live there — meta-llama-4-announcement, meta-muse-spark-announcement, meta-muse-spark-1-1-announcement and meta-muse-image-video-announcement — and none carries a verification date from this run. Meta itself was still covered, at 6 claims, through huggingface.co and github.com, which the dataset already trusts.
- abc.xyz returned HTTP 403. The one source on it, alphabet-about, was last checked 2026-08-15 and is now the oldest unverified source in the dataset. It is a corporate about-page backing a publisher record rather than a model fact, so nothing about a model rests on it, but it will keep aging until either the host answers this client or a human re-reads it.
- This run is re-verification first. It asks whether recorded facts still hold; it does not systematically ask what each creator has shipped that the dataset has never heard of. Discovery was attempted only where a creator index page made it cheap. So a green result here says the dataset is not stale; it says very little about whether it is complete.
- The known breadth gaps are unchanged and were not worked this run: OpenAI’s realtime, transcribe and tts models, the roughly twenty Cohere Command models against the one recorded, Microsoft’s MAI-Code-1.1-Flash, MAI-Image-2.5, MAI-Voice-2 and MAI-Transcribe-1.5, and Amazon’s Nova Act, Nova Forge and Nova Multimodal Embeddings. Each is breadth work with its own sourcing burden.
- A retirement-date conflict was found and deliberately not resolved. The Gemini API deprecations page says no shutdown date is announced for the 2.5 family, while the Cloud platform page gives a retirement date of 20 October 2026 for the same models. No dataset field holds either figure, so there was nothing to claim and nothing to withhold; it is reported here so the conflict is on the record rather than discovered again next run.
What was evaluated
- Reviewers
- 9
- Verdicts cast
- 606
- Accepted by panel
- 199
- Rejected by panel
- 3
Deterministic gates and required checks — 8 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence.mjs40 claim bundles 0 Pass197 claims admissible with complete three-rubric panels across all 40 gated bundles. Every bundle declared its policy and the gate independently re-derived the threshold from tools/updater/profiles, so no bundle’s declared policy went unchecked — which matters more than usual this run, since 33 of the 40 bundles claimed the stricter bar. gate-source-approval.mjs40 claim bundles 0 PassZero proposed sources across all 40 bundles: every cited source already existed in the dataset at anchor 329a719f, on one of the 46 approved origins derived there from 225 dataset sources and 18 catalogues across 19 profile files. The run therefore approved no source of its own. The two sources it did propose had already been withdrawn by the pairing cascade before this gate saw the gated bundles. Run before any edit was applied, and re-run after the re-anchor. gate-dataset.mjsweb/src/data 0 PassZero failures across the documents raw.ts composes, after the 179 edits were applied. npm run validateweb/ 0 PassTests green and astro check reported 0 errors and 0 warnings across 246 files. It failed on first invocation because this worktree had no node_modules — a missing vitest, not a data fault — and passed after npm ci. That first failure is recorded rather than quietly overwritten, because "the checker could not run" and "the data is good" are different findings and only one of them is a pass. gate-scope.mjsbranch vs merge-base 0 Passchanged: 3, empty: false, outOfClass: []. Anchored on 329a719f, the merge base the gate computed rather than one supplied to it — requestedBase was null. The three are sources.json, releases.json and families.json; this ledger is the fourth path in the qualifying class and is covered by gate-ledger instead. gate-ledger.mjsthis entry vs the diff it describes 0 Passtranscription: false — this entry was written by the run it describes and ships in the same pull request as the data, so its record counts were reconciled against the actual diff rather than taken on trust. That is the gap issue #419 recorded on three consecutive runs. ci-preflight.mjsnot requiredselected from the branch diff 0 PassSelected and ran the three check groups this diff triggers — web-ci, skills-ci and source-link-health-tests — locally from the repository root before the branch was pushed. It reports its own blind spots on every run, including the networked link-health sweep and the second Python interpreter. web-cithe pull request — PassGitHub performed the merge via --auto on this check going green; the agent did not merge. Posted 179 edits
179 edits across 3 documents, a net change of 0 records.
Dataset documents this run changed Document Before After What changed sources.json225 225 89 lastCheckedDate values moved to 2026-09-01, one per source page successfully re-read this run. No source was added or removed: gate-source-approval recorded zero proposed sources across all 40 bundles, so every date moved on a record that already existed. 107 accepted claims produced these 89 edits, because 18 of them re-assert one of two shared licence pages that several creators cite independently. releases.json96 96 50 verifiedAt dates moved to 2026-09-01, each backed by a release-level fact re-read from its primary source and found unchanged. Two releases that were scouted did not move because the panel refused their claims — hugging-face-smollm3-3b and ibm-granite-4-0-h-small, both listed under withheld. No release was added; the one new release found this run was refused. families.json67 67 40 verifiedAt dates moved to 2026-09-01 — one family per creator, each re-read and found unchanged. No family was added or removed and no family status value changed. Not posted 5 items
Rejected by the review panel
releases/google-gemini-2-5-flash-lite (proposed, not added)The only genuinely new release this run found, and it was refused 1-of-2 under the pilot bar. Provenance: the record sets accessType to proprietary-hosted, but none of the five cited quotes says anything about how the model is released or whether weights are downloadable, so that field had no quoted fact behind it. Editorial: the record puts the marketing name in canonicalName and the API model id in apiAliases, which it read as inverting the creator profile’s naming rule. Consistency accepted it. The release is not in the dataset and is left for a future run to propose properly rather than being trimmed until it passed. Note the caveat about the editorial rationale: its premise about existing Gemini records does not hold, and it was still not overruled.Blocked byprovenance rubriceditorial rubricreleases/hugging-face-smollm3-3b.verifiedAtReached 2 of the 3 accepts a long-tail creator requires. Provenance objected that the quote is a YaRN rope-scaling config comment showing the arithmetic 2 x 65536 = 131072, which does not by itself state the context window; the card does say so in prose nearby, but that sentence was not the quote. A stricter bar than the pilot creators face, applied as written. The record keeps its previous verifiedAt.Blocked byprovenance rubricreleases/ibm-granite-4-0-h-small.verifiedAtReached 2 of 3. Editorial objected not to the claim but to the record it would have re-verified: the existing summary opens "The largest Granite 4.0 model", which ranks models within a family, and this project forbids ranking prose. That is a pre-existing defect in committed data rather than anything this run proposed, and the rubric was right to refuse to bless it. Fixing the prose is a separate change and is filed as a follow-up; the record keeps its previous verifiedAt until then.Blocked byeditorial rubric
Accepted by the panel, then dropped
sources/google-gemini-2-5-flash-lite-docs (proposed, not added)Accepted unanimously on its own merits and still withdrawn, because the only record that would have cited it was refused. validateDataset treats a source no record references as dead provenance and refuses to load the dataset, so a source-add and the record citing it stand or fall together. Withdrawn by the run rather than by a reviewer.Blocked bypairing with the withheld release recordsources/google-gemini-2-5-flash-lite-platform-docs (proposed, not added)Accepted unanimously and withdrawn for the same reason as its sibling: the release that would have cited it was refused, and an uncited source is dead provenance.Blocked bypairing with the withheld release record
What this run does not prove
- No human reviewed any claim, verdict or edit in this run before it merged and deployed. That is what ADR 0003 authorises and what it costs.
- The three rubrics are three instances of the same model family reading the same page, so they share a failure mode: a source that is itself wrong can carry all three. The panel defends against unevidenced inference, not against a primary source that is confidently mistaken.
- This run is broad and shallow by design. It asks whether recorded facts still hold, not whether each creator has shipped something the dataset has never heard of. Only one genuinely new release surfaced, and it was withheld — so a green run here is weak evidence about completeness and strong evidence only about staleness.
- pagesFetched counts the 233 pages that returned 200. Seven further URLs were attempted and failed and are excluded from that figure rather than counted as reads.
- The editorial rubric's objection to the Gemini 2.5 Flash-Lite record asserted that the record inverts a naming convention 'every other Gemini entry' follows. That premise is factually wrong — google-gemini-2-5-flash already carries the marketing name in canonicalName. The claim was still withheld, because the chair does not vote and does not overrule a rubric; the error is recorded here and filed as a follow-up rather than corrected into a pass.
- reviewers and verdictsCast are derived rather than quoted: 3 rubrics x 3 batches = 9 reviewer invocations, and 202 claims x 3 rubrics = 606 verdicts, which reconciles with the nine per-rubric verdict files.
- The deploy had not run when this entry was written, and the entry says so rather than predicting it. A reader checking whether the site actually rebuilt should read the run’s summary issue, not this stage note.
- Zero claims carried over a dissent is a cleaner number than it looks. 197 of the 202 claims were re-verification of facts already in the dataset, where the three rubrics have little room to disagree; all four rejections landed on the three claims that were doing something harder. Read it as "this run attempted little that was contentious", not as "the panel found nothing to argue about".
- This entry was corrected in place shortly after it published, and the corrections are listed here rather than made silently. As first written it recorded issue #419 as open when it had closed on 2026-08-31; it called a zero-dissent panel "the cleanest panel result the ledger records" when run 2026-08-30-c0b6e9 had also recorded zero; and it twice miscounted the four rejections as landing on four and on five claims when they landed on three. All three were errors in this entry's own prose, not in the data the run published, and the dataset edits were unaffected. They are worth noting because of where they occurred: the ledger entry is the one document in the qualifying class that no reviewer rubric reads, so nothing in the panel or the gates was ever going to catch them. The gates checked this entry's record counts against the diff, which is exactly what they promise and no more.
Follow-ups — proposed, not fixed
- The editorial rubric refused the Gemini 2.5 Flash-Lite record on the ground that it inverts a naming convention every other Gemini entry follows. The premise is wrong: google-gemini-2-5-flash already carries the marketing name in canonicalName and the API id in apiAliases, exactly as the refused record did. The verdict was left standing because a chair that overrules a rubric it disagrees with is not running a panel, but the underlying question is real and unresolved — either the profile naming rule or the committed Gemini records are wrong, and until one is fixed this release will keep being refused for a reason that does not hold.
- ibm-granite-4-0-h-small’s summary opens with "The largest Granite 4.0 model", a size ranking inside committed data that the editorial rubric refused to re-verify. It will keep failing review every run until the prose is fixed, and it is worth checking whether other records carry similar superlatives, since nothing scans for them.
- ai.meta.com answers this runner with HTTP 400 on every URL and every fetch path tried, which is a different failure from the unreachable ai.google.dev of the previous run and may be client fingerprinting rather than a block. Four recorded sources are behind it. If it does not clear on its own, those sources need either a working fetch approach or an alternative approved origin, or they will keep aging out.
- The review panel was run as 9 invocations over 3 batches rather than one per creator, which is what made 202 claims affordable. It is worth deciding deliberately whether batching weakens the independence guarantee: a reviewer judging 76 claims in one context can see patterns across them that a per-creator reviewer cannot, which cuts both ways — more consistency, but also more room for one framing to carry a whole batch.
- This run advanced 89 source dates and 90 record dates without a single new fact reaching the dataset. That is a healthy outcome for a re-verification run, but two consecutive runs like it would mean the dataset is being kept fresh rather than kept complete. The next run should lead with discovery, and the ledger should probably distinguish the two modes explicitly rather than leaving a reader to infer it from claimsByKind.
2026-08-31-b7c2d9
Long-tail depth tranche 2026-08-31 — EleutherAI, Ai2, TII, NVIDIA, IBM, Cohere
Scope requested: Six creators already in the catalogue, each holding exactly one family — EleutherAI, Ai2, TII, NVIDIA, IBM and Cohere — under the long-tail unanimous 3-of-3 policy of ADR 0002, since none of the six has a reviewed profile. One additional family per creator was sought, with at least one release under each, restricted to origins already approved at the merge-base. The six were taken from the issue’s own candidate table; a coordinating brief naming stability-ai in place of cohere arrived after the panel had returned its verdicts, too late to change what was scouted, so cohere is in this run and stability-ai is not. That substitution is recorded under notCovered along with a read-only probe of what stability-ai would have run into. No profile catalogue was edited, no origin was added, and no existing record was corrected: every claim was kind add, so a correction found along the way would have been a separate issue rather than folded in.Ran, changed nothing0 edits posted · 30 items withheldResearched a second model family for six creators that each hold exactly one, intending to take the exactly-one-family count from 30 down to 24, and published none of them. Thirty claims were proposed, all of kind add, carrying 69 evidence entries; every quote was machine-checked as a contiguous verbatim substring of the exact bytes its contentHash names, with a positive and a negative control on the harness, and all 69 passed. Origin discipline was not the blocker: every page came from an origin the repository already approves, and gate-source-approval exited 0 on all six bundles. Dates were not the blocker either, which is what separates this run from the breadth tranche the day before: each proposed family took its firstReleaseDate from a document distinct from the one dating the release it ships with, so no family date was derived from its own release. The panel was the blocker. Under the unanimous 3-of-3 long-tail bar the provenance rubric rejected all 15 family and release claims, and the reason is nearly uniform: releaseSchema and familySchema both require status, and no model card or announcement read in this run states a lifecycle term that forces one member of the enum. An announcement verb — "We are releasing the following three models on Hugging Face" — is not an availability state, and a value filled from what is usual for the field is a guess wearing a vocabulary term. The same rubric rejected accessType: open-weight wherever it rested only on a licence name, since a licence is not a statement that weights are downloadable; it rejected two Falcon 3 parameter counts read out of the model name; and it rejected the Molmo family and release dates, which rested on a page dateline no quote carried. Two further rejections came from other rubrics and are worth recording because they found real defects rather than duplicating provenance: consistency applied the project’s own validateDataset to a clone with all 30 claims applied and showed that the EleutherAI and Cohere releases assert weight-bearing access types with no licence, which releaseSchema refuses outright; and editorial found that the Cohere weights card states "This model is governed by a CC-BY-NC License with an acceptable use addendum" in prose, so presenting those weights as openly available while omitting the licence understated a non-commercial restriction. The 15 source records were accepted unanimously and dropped anyway, because with every citing record rejected they would have landed as orphans. The dataset is unchanged, the exactly-one-family count stays at 30, and this entry is the only edit.
- PreflightRanRe-measured the trunk baseline from the committed data rather than quoting the issue: 40 organizations, 64 families, 92 releases, 217 sources, and 30 organizations holding exactly one family — the number this run set out to reduce. Confirmed all six target creators sit at exactly one family, and that all six publishers already exist, so publishers.json needed no change. Merge-base and HEAD were both 3e255145ed, the same trunk commit the issue measured against.
- ScoutRanStored 33 page bodies in total, of which 31 belong to the scout proper and 2 to a read-only stability-ai probe run after the panel reported. They cover that many distinct URLs across eight approved origins (huggingface.co, github.com, allenai.org, docs.cohere.com, cohere.com, www.ibm.com, research.nvidia.com, opensource.org). The count is of stored bodies, not fetch operations: a re-fetch overwrites its cached body rather than accumulating, so the number of operations was higher and is not separately recorded. Two fetches returned HTTP 401 behind a gated form (the Cohere Command R 08-2024 raw README and the Aya Expanse 32B card) and their bodies are stored and hashed as the 401 responses they are. Text extraction turned every HTML tag into a newline so that stripping markup could not concatenate text across element boundaries and invent a phrase that appears on no page.
- ReviewRanThree reviewers, one per rubric, launched as parallel independent sub-agents. Each received the claims, their evidence, the dataset and its schema, the cached bytes, and its own rubric — and none received the scout’s reasoning, another reviewer’s verdicts, or the running tally. 90 verdicts cast, each with a written rationale. No reviewer was re-run and no tie was broken: with three reviewers and two outcomes there are no ties, and a claim short of its threshold is simply rejected. All three independently rebuilt the contentHash-to-file mapping by re-hashing the cache rather than trusting filenames, and all three confirmed the 69 quotes verbatim.
- GatesRangate-source-approval exited 0 on all six bundles. gate-evidence exited 1 on all six, naming exactly the 15 family and release claims as short of the 3-of-3 long-tail threshold — the gate agreeing with the panel that they must not be applied, which is the gate working rather than failing. gate-dataset, gate-scope and gate-ledger each exited 0 over the unchanged dataset. No gate was skipped, overridden, or re-run with a different threshold.
- PublishNot applicableNothing reached the unanimous threshold and survived its dependencies, so there was nothing to apply. The dataset documents were never opened for writing. This entry is the only change the run made.
- DeployNot applicableNothing merged, so nothing deployed.
What was found
- Scouts
- 6
- Pages fetched and hashed
- 33
- Claims proposed
- 30
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims eleutherai long-tail 3-of-3 4 ai2 long-tail 3-of-3 4 tii long-tail 3-of-3 6 nvidia long-tail 3-of-3 6 ibm long-tail 3-of-3 6 cohere long-tail 3-of-3 4 What those claims proposed to do Kind Count Effect Add 30 Every claim proposed a new record: 15 sources, 6 families and 9 releases. None reached the dataset. The run was scoped to additions only, so no existing record was changed or removed. Not covered
- GPT-J-6B and GPT-NeoX-20B were read and withheld before review: neither card nor repository states a release date in prose, and the only date-bearing metadata is the Hub upload timestamp, which is not a stated release date.
- The whole Cohere Aya line was read and withheld before review: the documentation states status, modality and context figures but no release date for any Aya model, and the Hugging Face weights cards are behind a gated form that returned HTTP 401.
- command-r-03-2024 was named as the Command R family’s earliest member but never proposed as a release: Cohere states a deprecation date for it and no release date, and its month rests only on the timestamp in its identifier.
- This run checked six creators for one additional family each. It is not a sweep of everything those six have published, and it says nothing about the other 24 creators holding exactly one family.
- stability-ai was named in the coordinating brief in place of cohere, and that brief arrived after the panel had already returned its verdicts, so stability-ai was never scouted and no claim was proposed for it. It was probed read-only rather than left silent: its SDXL and Stable Video Diffusion model cards were fetched and searched for lifecycle, status and release-date wording. The SDXL card contains no line matching any of those terms at all, and the two matching lines on the Stable Video Diffusion card state a usage restriction and a download-statistics note, neither of which is a lifecycle state. On the evidence of that probe stability-ai would have failed on the same required status field as the other six, but a probe is not a scout and this is reported as an indication rather than a result.
- No context window was recorded for Granite 4.0: the only figure sits in an HTML table whose column-to-model mapping is not legible from the text, so a value would have been read from a column position rather than from a statement.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 90
- Accepted by panel
- 15
- Rejected by panel
- 15
Deterministic gates and required checks — 6 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-source-approvalall six claim bundles against the approved origin set at merge-base 3e255145ed 0 PassExit 0 on every bundle. Each of the eight origins cited — huggingface.co, github.com, allenai.org, docs.cohere.com, cohere.com, www.ibm.com, research.nvidia.com and opensource.org — was already approved, so no origin needed widening and none was widened. This run therefore adds no evidence either way about creators whose own domain is unapproved. gate-evidenceall six claim bundles, evidence form and review threshold 1 FailExit 1 on all six bundles, naming the 15 family and release claims as short of the threshold — variously at 0, 1 or 2 of the 3 required accepts under the long-tail policy, never at 3. The evidence form itself was not the complaint: all 69 entries carry the six required fields with a sha256 hash, an https credential-free url, a real fetch date and a quote at or above the 24-character floor. The refusal is the threshold, and it agrees with the panel. Per ADR 0005 this gate checks form and never fetches, so the verbatim check was done separately against the stored bytes. gate-datasetthe committed dataset, read-only coherence check 0 PassExit 0: all gates passed over 472 records. The dataset is unchanged by this run, so this records that trunk stayed coherent, not that anything this run produced was coherent. gate-scopemerge-base 3e255145ed to HEAD, plus the working tree 0 PassExit 0, reporting one dataset document changed since 3e255145ed and in class: refresh-runs.json, this entry. No other tracked file was touched, and in particular no test, no library and no workflow — a single out-of-class path would have disqualified the whole change. The run’s own artefacts — bundles, verdicts and cached page bodies — live under the git-ignored .modeltree-refresh directory and are invisible to this gate by design. gate-ledgermerge-base 3e255145ed to HEAD, this entry against the branch diff 0 PassExit 0, and the pass is narrower than it looks, which is worth stating rather than leaving to be discovered: the gate classified this entry as a transcription, because the branch changes no dataset document other than the ledger itself, and reported in as many words that the record counts were therefore NOT checked against a diff and should be read as unverified. That is the correct behaviour for a run that published nothing — there is no diff for the counts to be checked against — but it means the numbers here rest on the run’s own artefacts and on the reviewers’ verdict files, not on this gate having confirmed them. Counts were instead derived mechanically from the six bundles and three verdict files rather than typed by hand. npm run validateweb/, the full test suite plus Astro and TypeScript diagnostics 0 Pass107 test files, 2395 tests, 0 failures; 232 files checked with 0 errors and 0 warnings. Run from web/ over this entry, so it covers the entry against refresh-log-schema.ts and not merely the unchanged dataset — which is the only thing this run gives it new to check. node .github/scripts/ci-preflight.mjsrepository root; the pull-request checks this branch’s diff selects, anchored at the merge-base 0 PassSelected and passed two check groups, web-ci and skills-ci, the latter including gate-ledger’s own 221 tests over this entry. It does not run the networked source-link-health sweep, the browser end-to-end check, or the second Python interpreter, and it judges this branch rather than this branch merged into a main that has since moved, so a pass here predicts CI rather than binding it. Posted 0 edits
Nothing reached the dataset. No branch, no commit, no pull request.
Not posted 30 items
Rejected by the review panel
eleutherai-family-gpt-neo-addfamilies record eleutherai-gpt-neo for eleutherai. [provenance] firstReleaseDate 2021-03-21 (day) is well-sourced by the dated release note '**Update 21/03/2021:** ... two pretrained GPT-Neo models', but status:'current' is unsourced. The two attached quotes state a release event and a family definition ('GPT-Neo refers to the class of models...'); neither states any lifecycle/availability state, so nothing forces 'current' over 'legacy'. A required status filled from what is usual for the field is a guess wearing a vocabulary term (SKILL.md provenance, reject bullet 1). [consistency] The family record itself is well-formed and is NOT a duplicate of the creator's existing eleutherai-pythia family (distinct slug, and GPT-Neo 2021-03-21 predates Pythia 2023-04-03). It is rejected on whole-dataset coherence: its only proposed release, eleutherai-release-gpt-neo-2-7b, is schema-invalid (open-weight with no license) and must be dropped, which leaves this family with zero releases. Running the real validateDataset with the release dropped but the family kept fails with 'family eleutherai-gpt-neo has no releases' (validate.ts:549). As submitted the pair cannot land coherently; fix is to add the missing license to the release and re-submit family+release together.Blocked byrubric:provenancerubric:consistencyeleutherai-release-gpt-neo-2-7b-addreleases record eleutherai-gpt-neo-2-7b for eleutherai. [provenance] status:'current' is unsourced: the attached quotes describe only architecture, parameter count and training data, none stating a lifecycle/availability state. Separately, accessType:'open-weight' rests on the bare labeled link '2.7B: https://mystic.the-eye.eu/public/AI/gptneo-release/GPT3_2-7B/'; the page's preceding 'The weights can be freely downloaded' sentence is NOT in the attached quote, so read alone the link does not state that weights are downloadable (mi4 / co5 precedent: a link or licence is not a downloadable-weights statement). [consistency] accessType is 'open-weight' but the record carries no license object. releaseSchema.superRefine (schema.ts:419-426) requires a license whenever accessType is open-weight or both. Running the actual datasetSchema over existing+all-30 fails with 'releases.92.license: is required when a release claims downloadable weights'. The bundle's own incomplete note concedes the license (mit) exists only in YAML front matter and was omitted. As submitted this record makes the dataset fail Zod validation, so it cannot be admitted.Blocked byrubric:provenancerubric:consistencyai2-family-molmo-addfamilies record ai2-molmo for ai2. [provenance] Two provenance failures. (1) firstReleaseDate '2024-09' rests on the page's 'September 25, 2024' dateline, which the bundle itself says is not carried as a quote; a publication timestamp is not a stated release date, and no attached quote states any date ('Today we are releasing 4 samples' gives no date). (2) status:'current' has no attached availability quote, and the family's own released checkpoint is a 'preview', so more than one member could fit.Blocked byrubric:provenanceai2-release-molmo-7b-d-addreleases record ai2-molmo-7b-d for ai2. [provenance] status:'preview' is correctly sourced ('This checkpoint is a **preview** of the Molmo release.'). But accessType:'open-weight' and license.weightsDownloadable:true have no supporting quote: 'This model is licensed under Apache 2.0' is a licence name, not a statement that weights are downloadable (co5 reversal precedent), and 'open vision-language models' does not distinguish open-weight from source-available. releaseDate '2024-09' also rests on the unquoted page dateline rather than a stated release date.Blocked byrubric:provenancetii-family-falcon-3-addfamilies record tii-falcon-3 for tii. [provenance] firstReleaseDate 2024-12 is sourced ('- Model Release Date: December 2024'), and '...a set of pretrained and instruct LLMs ranging from 1B to 10B' defines the family, but status:'current' is unsourced. No attached quote states any lifecycle/availability state; the Falcon3 card carries no status wording at all. 'current' is filled from what is usual, which the rubric forbids.Blocked byrubric:provenancetii-release-falcon-3-10b-instruct-addreleases record tii-falcon-3-10b-instruct for tii. [provenance] Three gaps. status:'current' has no attached lifecycle/availability quote. accessType:'open-weight' rests on '- License: TII Falcon-LLM License 2.0' — a licence name, not a statement that weights are downloadable. parameters.totalBillions:10 is read from the model name; no attached quote states a 10B count ('under 10 billion parameters' describes the family, not this model's count). contextWindow 32000 and the date are fine, but the above sink it.Blocked byrubric:provenancetii-release-falcon-3-7b-base-addreleases record tii-falcon-3-7b-base for tii. [provenance] As with the 10B variant: status:'current' has no attached availability quote; accessType:'open-weight' rests only on the licence name '- License: TII Falcon-LLM License 2.0'; and parameters.totalBillions:7 is taken from the model name — the only count-bearing quote, '...ranging from 1B to 10B', is the family range, not this model's parameter count.Blocked byrubric:provenancenvidia-family-nemotron-nano-2-addfamilies record nvidia-nemotron-nano-2 for nvidia. [provenance] Release date 2025-08-18 is sourced ('### Release Date: 08/18/2025'), but status:'current' is unsourced. 'We are releasing the following three models on Hugging Face...' is an announcement verb, not a lifecycle/availability term (mi4 precedent: an availability/CTA statement does not state a status). No attached quote forces 'current'.Blocked byrubric:provenancenvidia-release-nemotron-nano-9b-v2-addreleases record nvidia-nemotron-nano-9b-v2 for nvidia. [provenance] status:'current' has no attached lifecycle/availability quote. accessType:'open-weight' rests on 'Governing Terms: Use of this model is governed by the NVIDIA Open Model License Agreement' — a licence name (despite the word 'Open'), not a statement that weights are downloadable (co5 reversal). Date, contextWindow 128000 and the licence name are otherwise fine, but status and access type are unsourced.Blocked byrubric:provenancenvidia-release-nemotron-nano-12b-v2-base-addreleases record nvidia-nemotron-nano-12b-v2-base for nvidia. [provenance] status:'current' unsourced; accessType:'open-weight' unsourced; and the record's license.name 'NVIDIA Open Model License Agreement' has NO attached quote in this claim at all (its evidence is only the release date, the three-model announcement, 'Model Developer: NVIDIA Corporation', and the OSI index). Nothing attached states downloadable weights or even names the licence.Blocked byrubric:provenanceibm-family-granite-4-0-addfamilies record ibm-granite-4-0 for ibm. [provenance] firstReleaseDate 2025-10-02 is sourced ('- **Release Date**: October 2nd, 2025'), but status:'current' is unsourced. 'The Granite 4.0 collection comprises multiple model sizes and architecture styles...' defines the collection and states no lifecycle/availability state; no attached quote forces 'current' over 'legacy' (the description itself notes a newer Granite 4.2 exists).Blocked byrubric:provenanceibm-release-granite-4-0-h-small-addreleases record ibm-granite-4-0-h-small for ibm. [provenance] parameters (32B total, 9B active) are well-sourced ('...a 32B parameter long-context instruct model...' and 'a hybrid mixture of experts (MoE) model with 32B total parameters (9B active)'), and osiApproved:true is backed by OSI's Apache page ('Open Source Initiative Approved License'). But status:'current' has no attached lifecycle/availability quote, and accessType:'open-weight' rests only on '- **License:** Apache 2.0' — a licence name, not a statement that weights are downloadable (co5 reversal).Blocked byrubric:provenanceibm-release-granite-4-0-h-tiny-addreleases record ibm-granite-4-0-h-tiny for ibm. [provenance] As with H-Small: parameters (7B total, 1B active) are sourced ('...a 7B parameter long-context instruct model...' and 'a hybrid MoE with 7B total parameters (1B active)'), but status:'current' has no attached availability quote and accessType:'open-weight' rests only on the Apache licence name, not a downloadable-weights statement.Blocked byrubric:provenancecohere-family-command-r-addfamilies record cohere-command-r for cohere. [provenance] status:'current' is scoped, not supported. The only lifecycle term attached is 'Deprecated Sept 15, 2025' for command-r-03-2024 — a different, earliest member — which does not establish the family as 'current'; the 'Live' status of command-r-08-2024 is in a different claim's evidence, not this one. A per-release status is not a family status (co5-family precedent). [consistency] The family record itself is well-formed and NOT a duplicate of the creator's existing cohere-command-a family (Command A, 2025-03) - the description correctly keeps the two lines apart, and Command R (2024-03) is the earlier separate line. Its firstReleaseDate 2024-03 also legitimately precedes its only release (2024-08) without tripping the predates-family rule. It is rejected on whole-dataset coherence: its only proposed release, cohere-release-command-r-08-2024, is schema-invalid (accessType 'both' with no license) and must be dropped, leaving this family with zero releases. The real validateDataset then fails with 'family cohere-command-r has no releases' (validate.ts:549). Fix is to supply the release's license and re-submit the pair together.Blocked byrubric:provenancerubric:consistencycohere-release-command-r-08-2024-addreleases record cohere-command-r-08-2024 for cohere. [provenance] status:'current' (mapped from the quoted 'Live'), accessType:'both' ('a large language model with open weights' plus the API 'Live' listing), contextWindow 128000 ('supports a context length of 128K') and parameters 32B ('32 billion parameter') are all properly sourced. But the record and its statement assert a '4k maximum output' (maximumOutput 4000) with NO attached quote: the models-overview fragment stops at 'delivered in August 2024' and never reaches a max-output figure for this model. The claim states more than the attached sources do (codex precedent: attach the quote, do not map from nothing). [consistency] accessType is 'both' (Cohere API plus open weights), which claims downloadable weights, but the record carries no license object. releaseSchema.superRefine (schema.ts:419-426) requires a license for open-weight or both. Running the real datasetSchema over existing+all-30 fails with 'releases.100.license: is required when a release claims downloadable weights'. The bundle's incomplete note concedes the licence (CC-BY-NC) appears only inside a link element and was not captured. As submitted this record makes the dataset fail Zod validation and cannot be admitted. [editorial] Two editorial defects. (1) Licence/access: the record sets accessType "both" (open weights) yet omits `license`, justifying it as "the weights card names one only inside a link element and states none in prose" — but the cited weights card states in prose "This model is governed by a CC-BY-NC License with an acceptable use addendum". Dropping that CC-BY-NC non-commercial/research restriction while presenting the weights as openly available misrepresents how open the model is. (2) Context basis: the summary asserts "Cohere publishes no integers for either", but the cited command-r documentation page states "a long 128,000-token context length" — an explicit integer — so the prose makes a false negative about its own sources and understates the firmness of the figure.Blocked byrubric:provenancerubric:consistencyrubric:editorial
Accepted by the panel, then dropped
eleutherai-gpt-neo-repository-source-addSource record eleutherai-gpt-neo-repository was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordeleutherai-gpt-neo-2-7b-model-card-source-addSource record eleutherai-gpt-neo-2-7b-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordai2-molmo-announcement-source-addSource record ai2-molmo-announcement was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordai2-molmo-7b-d-model-card-source-addSource record ai2-molmo-7b-d-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordtii-falcon-3-announcement-source-addSource record tii-falcon-3-announcement was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordtii-falcon-3-10b-instruct-model-card-source-addSource record tii-falcon-3-10b-instruct-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordtii-falcon-3-7b-base-model-card-source-addSource record tii-falcon-3-7b-base-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordnvidia-nemotron-nano-2-announcement-source-addSource record nvidia-nemotron-nano-2-announcement was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordnvidia-nemotron-nano-9b-v2-model-card-source-addSource record nvidia-nemotron-nano-9b-v2-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordnvidia-nemotron-nano-12b-v2-base-model-card-source-addSource record nvidia-nemotron-nano-12b-v2-base-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordibm-granite-4-0-announcement-source-addSource record ibm-granite-4-0-announcement was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordibm-granite-4-0-h-small-model-card-source-addSource record ibm-granite-4-0-h-small-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordibm-granite-4-0-h-tiny-model-card-source-addSource record ibm-granite-4-0-h-tiny-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordcohere-command-r-model-card-source-addSource record cohere-command-r-model-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing recordcohere-command-r-08-2024-weights-card-source-addSource record cohere-command-r-08-2024-weights-card was accepted unanimously 3-of-3 on all three rubrics, and was dropped anyway because every record that would have cited it was rejected. A source add must be paired with a claim wiring it into some record's sourceIds — check-bundle-pairing.mjs enforces that pairing, and validate.test.ts refuses a source the dataset cites from nowhere — so applying it alone would have added an orphan. The research behind it stands; only the record it was to support does not.Blocked bycheck-bundle-pairingno surviving citing record
What this run does not prove
- A withheld tranche is a real outcome of the policy, not a failure of the sources. Every page came from an origin this repository already approves and every quote was verified verbatim against stored bytes. The run reports six creators it could not record rather than six it stretched a citation to reach.
- The blocker is a required field with no sourceable value, and better scouting does not reach it. releaseSchema and familySchema both require status; the schema offers preview, current, legacy, deprecated and research and no unknown member; and a Hugging Face model card states no lifecycle term at all. Where a creator does publish a status — Cohere lists command-r-08-2024 as Live — the mapping to current was accepted without argument. So this is not a claim that status is unsourceable in general; it is that a creator who publishes only model cards has no route to a release record, which is most of the long tail.
- This entry does not claim the panel was wrong, and that was checked rather than assumed. The provenance rubric carries a section written specifically to stop a dissent on a mapped vocabulary field from blocking long-tail claims, on the ground that status and contextWindow are recorded by mapping a creator’s own wording and no source speaks the dataset’s terms — so a panel that treated mapping as inference would have been misapplying its own rubric, and under a unanimous bar that error would silently block every long-tail creator. The rejections were therefore read against that section before this entry was written. They conform: the panel accepted the mapping wherever a status was actually stated, taking Cohere’s quoted "Live" to current and Ai2’s quoted "preview" to preview in as many words, and rejected status only where no attached quote states any lifecycle state at all — which is the mi4-release-large-3 row the rubric lists as a rejection that stands unchanged. Two of the three rejections that did not come from provenance found defects the scout had missed outright, which is further evidence the panel was reading rather than rubber-stamping.
- The run deliberately did not re-scout to press a different phrase into status. Where the rubric names a remedy it is to attach a quote the page already carries, not to map from nothing, and several rejections here are exactly that repairable kind — but repairing them would not have saved the records, because the unrepairable status gap sits on the same claims. Re-running a rejecting reviewer to obtain a different verdict is forbidden outright and was not done.
- The six bundles, all 90 verdicts with their rationales, and the 31 stored page bodies are reproducible artefacts of this run, but they live under a git-ignored directory and go with the machine. This entry is the durable record; once that directory is gone, every citation here can only be rechecked by re-fetching a live page that may have changed.
- Nothing here says the dataset is current. Six creators were checked for one additional family each. A creator not named in this entry was not looked at, and a family not proposed is not thereby known to be absent.
Follow-ups — proposed, not fixed
- The tranche this run set out to publish is blocked on a schema-and-rubric interaction rather than on evidence: status is required on every release and every family, the provenance rubric requires a quote that forces one enum member, and a model card states no lifecycle term. Three routes are open and this run took none of them, because all three are outside what a refresh agent may decide: make status optional, or give it an unknown member, so a release can be recorded without asserting a lifecycle nobody stated; write down that a creator publishing weights with no stated lifecycle is recorded as current, which is an editorial convention a human should set rather than a reviewer infer; or add these families through a change that takes the ordinary path, where a human reviews the reading on the way in.
- accessType: open-weight was rejected on five releases because it rested on a licence name, a download link, or the word "open" in a licence title, none of which states that weights are downloadable. Several of the cards do carry such a statement in prose — the GPT-Neo README says the weights can be freely downloaded a sentence before the link that was quoted instead — so this particular gap is a scouting defect and not a structural one, and is worth re-attempting on its own once the status question is settled.
- The research behind the 15 unanimously accepted source records is sound and was dropped only because the records citing it fell. If the status question is answered, these six creators do not need re-fetching from scratch, but only if the stored bodies are captured before the machine goes.
- Two claims failed on defects worth fixing regardless of the status question: the Cohere Command R 08-2024 release asserted a 4k maximum output with no quote attached, and the NVIDIA Nemotron Nano 12B v2 Base release named a licence its evidence never quoted. Both are attach-the-quote-or-drop-the-field problems in the scout, not disagreements with the panel.
2026-08-31-ae0342
Data refresh 2026-08-31
Scope requested: Six of the seven pilot creators — anthropic, openai, meta, microsoft, alibaba-cloud and google-deepmind — each judged at the 2-of-3 pilot threshold, plus amazon scouted without producing a claim. Long-tail creators were not scouted. Re-verification only, with the model-fit guidance layer deliberately taken first so the release and family records resting on it could move.Published33 edits posted · 1 item withheldAn agent-run, source-backed refresh under ADR 0003, and a narrow-but-deep run rather than a broad one: it set out to clear a blockage the 2026-08-30 run recorded and could not clear itself. That run had five verification dates it was entitled to advance and could not, because validate.ts refuses to let a model-fit statement be older than the fact it rests on, and the guidance layer had not been re-verified. This run re-verified the guidance layer first — all seven statements in model-fit-statements.json — which unpinned every one of the five. Six creators were scouted against 32 pages fetched and hashed this run, producing 43 claims; three rubrics voted blind on each, 42 met the 2-of-3 pilot bar and one was refused. 33 edits landed across four documents and nothing was withheld, against two withheld last run. The cost is breadth: the 33 long-tail creators were not scouted at all this run, which is the largest gap in this entry and is recorded under notCovered rather than smoothed over. No human reviewed any claim, verdict or edit.
- PreflightRanTree clean, gh authenticated, and no open pull request from a previous refresh at the moment this run started. Anchor 356989e9, selected by merge-base with refs/remotes/origin/main and computed by the gate rather than supplied — requestedBase was null. That is a re-anchor: the run opened on 36ce8b1b, and run 2026-08-31-651a1c merged as PR #665 while this one was between its review and publish stages, moving main under it and touching three of the same documents. The branch was brought onto the new main, the dataset reset to main’s version wholesale rather than hand-merged, and every claim re-checked against it — all 43 currentValues still held, so #665 had touched none of the records this run changes. gate-source-approval, gate-dataset, npm run validate, gate-scope and gate-ledger were then all re-run against the new anchor rather than carried over. The approved-origin catalogue parsed cleanly at 356989e9: 217 dataset sources and 18 catalogues across 19 profile files, yielding 46 approved origins. npm resolved to a PowerShell shim the machine execution policy refuses to run; npm.cmd ran and reported 11.9.0. The policy was not changed to work around it — that would be a machine-wide security change made to satisfy a probe.
- ScoutRanSeven pilot creators scouted against 32 pages fetched and hashed this run; six produced claims and amazon produced none. Every quote was sliced from a page read during the run and verified as a byte-exact substring of the saved body before the claim was built — raw bytes for markdown and plain text, rendered text for HTML. Two candidate quotes failed that check and were corrected rather than kept: one had a curly apostrophe where the page has an ASCII one, and one fell under the gate’s 24-character minimum. Appending .md to an OpenAI or Anthropic docs URL returns clean markdown of the same page on the same already-approved origin, which is how those client-rendered catalogues were read at all; the provenance rubric was asked to rule on that convention explicitly and accepted it. ai.google.dev was unreachable from this machine on all three attempts — fetch failed, not an HTTP status — so deepmind.google was used instead, an origin the dataset already carries.
- ReviewRanThree rubrics — provenance, consistency, editorial — voted independently on all 43 claims across 18 reviewer invocations, each in its own context and each given only its own rubric text, the claim, the evidence, the relevant dataset slice and the creator profile vocabulary. No reviewer saw another’s rubric, the threshold, the running tally, or the scout’s reasoning. 42 claims met the 2-of-3 pilot bar and one was refused 2-of-3. Three of the 42 passed over a single dissent, and all four minority opinions are published verbatim in the pull request body. No claim was edited and re-reviewed to chase a verdict, and the coordinating agent cast no vote and overruled none.
- GatesRangate-evidence and gate-source-approval ran across all six bundles before anything was applied, since gate-source-approval anchors on the committed dataset and running it afterwards would be asking the run’s own writes whether the run’s own sources are trustworthy. gate-dataset, npm run validate, gate-scope and gate-ledger ran afterwards. No gate was skipped, forced, or re-run to obtain a different answer, no threshold was lowered, and nothing was pushed to main. gate-scope reported changed: 5, empty: false, outOfClass: [], so the change sat inside the ADR 0003 qualifying class as ADR 0006 widened it.
- PublishRanOpened as a pull request carrying the full evidence trail — every claim, every quote with its content hash, all 129 verdict rationales, and the gate-source-approval anchor block — and merged by GitHub via --auto once web-ci went green. The agent did not merge, and did not use --admin.
- DeployRanPages deploy confirmed after the merge. See the deployment reference below.
What was found
- Scouts
- 7
- Pages fetched and hashed
- 32
- Claims proposed
- 43
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims anthropic pilot 2-of-3 12 openai pilot 2-of-3 12 meta pilot 2-of-3 10 microsoft pilot 2-of-3 3 alibaba-cloud pilot 2-of-3 4 google-deepmind pilot 2-of-3 2 What those claims proposed to do Kind Count Effect Unchanged 20 A recorded fact re-read from its primary source and found to still hold. 19 reached the 2-of-3 pilot threshold and all 19 applied, advancing 9 release verifiedAt dates, 1 family verifiedAt and 7 model-fit statement verifiedAt dates to 2026-08-31 — some records carried more than one confirming claim. Nothing was withheld. One was refused by the panel. Change 23 A source page re-read this run, moving its lastCheckedDate to 2026-08-31. All 23 were accepted, and 16 distinct source records moved — several pages carried more than one claim. Not covered
- The 33 long-tail creators were not scouted this run. That is the largest gap in this entry: the 2026-08-30 run covered all 33, and this one traded that breadth for depth on the guidance layer that was blocking pilot-creator re-verification. Their records carry no verification date from this run and nothing here says their facts still hold.
- amazon was scouted and produced no claim. aws.amazon.com/nova/ is client-rendered and the fetched body carried placeholder text where the model listing should be, so no quote on it could support a fact about a model. The one recorded Amazon release, amazon-nova-2-sonic, was verified 2026-08-30 and did not need this run.
- ai.google.dev was unreachable from this machine on every attempt — fetch failed rather than an HTTP status, so it is a network fact about this runner and not evidence the page is gone. Seven of the eight recorded Google releases rest on sources on that host and therefore carry no verification date from this run. Only google-gemini-pro-page, on the reachable deepmind.google, was re-read.
- Breadth remains uncovered by design and is unchanged from the previous run: OpenAI’s realtime, transcribe and tts models, the roughly twenty Cohere Command models against the one recorded, Microsoft’s MAI-Code-1.1-Flash, MAI-Image-2.5, MAI-Voice-2 and MAI-Transcribe-1.5, and Amazon’s Nova Act, Nova Forge and Nova Multimodal Embeddings. Each is breadth work with its own sourcing burden, not re-verification.
- Microsoft MAI-Image-2.6 was again deliberately not added, on the same grounds the 2026-08-30 run recorded: its only availability statement names a product and a serving platform rather than the creator’s own release, so accessType has no creator-level statement and status is ambiguous between preview and current. The ambiguity is left explicit rather than resolved by guess.
What was evaluated
- Reviewers
- 18
- Verdicts cast
- 129
- Accepted by panel
- 42
- Rejected by panel
- 1
Deterministic gates and required checks — 8 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence.mjs6 claim bundles 0 Pass43 claims admissible with complete panels across all six bundles, 23 applicable after the panel. Every bundle declared policy pilot and the gate independently derived threshold 2 from tools/updater/profiles, so no bundle’s declared policy went unchecked. gate-source-approval.mjs6 claim bundles 0 PassZero proposed sources across all six bundles: every cited source was already in the dataset at anchor 356989e9, on one of the 46 origins approved there. No source was refused and no origin was introduced, so the run approved nothing of its own. Run before any edit was applied. gate-dataset.mjsweb/src/data 0 PassValidated the documents raw.ts composes with zero failures after the edits were applied. npm run validateweb/ 0 Pass2361 tests executed across 104 files with all reporting results; astro check reported 0 errors and 0 warnings. This is the gate that pinned five records last run, and it passed them this run because the guidance beneath them had been re-verified first. It also failed intermittently on this runner for a reason that is not this change, recorded here rather than re-run until green: two to four tests in LineageModelDrawer.interaction.test.tsx and ModelTreeExplorer.interaction.test.tsx hit the 5000 ms vitest timeout, with different test names failing on each attempt. Both files pass 17 of 17 in isolation with this change applied, and pristine origin/main fails the same way with more failures than this branch, so the cause is worker contention on the machine and not the data. Constraining vitest to two workers produced the clean run recorded here. No timeout was raised, no test was skipped, and no gate was called passed on a run that had not finished. gate-scope.mjsbranch vs merge-base 0 Passchanged: 5, empty: false, outOfClass: []. Anchored on 356989e9, the merge base the gate computed rather than one supplied to it. The fifth file is this ledger entry, which ADR 0006 put in the qualifying class so a run can record itself. gate-ledger.mjsthis entry vs the diff it describes 0 Passtranscription: false — this entry was written by the run it describes, in the same commit as the data, rather than transcribed afterwards. That is the gap issue #419 recorded on three consecutive runs. ci-preflight.mjsnot requiredselected from the branch diff 0 PassRun locally from the repository root before the branch was pushed. web-cithe pull request — PassGitHub performed the merge via --auto on this check going green; the agent did not merge. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
anthropic-family-claude-4-5-status-unchangedProvenanceOutvoted 2-of-3 and applied. The overview page’s "Legacy models (still available)" list names Claude Opus 4.5 and Claude Sonnet 4.5 but not Claude Haiku 4.5, which appears instead in the current-lineup table on the same page. So the 4.5 generation’s members straddle both lists, and a quote establishing that one member is legacy does not establish a status for the family as a whole.meta-llama-4-scout-context-window-unchangedProvenanceOutvoted 2-of-3 and applied. The 10M figure sits in a table-row fragment whose column headers are not carried in the quote, so identifying which column the number belongs to requires an inference the quoted text does not itself supply.meta-fit-llama-4-scout-context-window-reverifyProvenanceOutvoted 2-of-3 and applied, on the same reasoning as the release-level context-window claim it rests on: the underlying table-row fragment does not carry the column headers that would fix which figure is the context window.
Posted 33 edits
33 edits across 4 documents, a net change of 0 records.
Dataset documents this run changed Document Before After What changed model-fit-statements.json7 7 All 7 statements had verifiedAt moved to 2026-08-31: fit-claude-haiku-4-5-release-current, fit-claude-haiku-4-5-family-legacy, fit-claude-mythos-5-limited-availability, fit-gpt-5-superseded, fit-llama-4-scout-self-hosting, fit-llama-4-scout-mau-threshold and fit-llama-4-scout-context-window. This document was taken first on purpose. validate.ts forbids a fit statement from being older than the fact it rests on, so while these sat at 2026-08-15 and 2026-08-18 they pinned the records beneath them — which is exactly why the previous run had to withhold two release dates it had otherwise earned. No statement text or judgement changed; only the date on which each was last checked against its evidence. releases.json92 92 9 verifiedAt dates moved to 2026-08-31 — anthropic-claude-haiku-4-5, anthropic-claude-mythos-5 and meta-llama-4-scout from 2026-08-15, openai-gpt-5 from 2026-08-18, openai-gpt-image-2 from 2026-08-27, microsoft-fara-1-5-27b, alibaba-qwen3-8-27b and alibaba-qwen3-8-2-4t-a95b from 2026-08-28, and openai-gpt-5-6-sol from 2026-08-29. The first three and openai-gpt-5 are the records the previous run could not move. No release was added or removed and no other field changed. families.json64 64 One verifiedAt moved: anthropic-claude-4-5 from 2026-08-15 to 2026-08-31, unpinned by fit-claude-haiku-4-5-family-legacy being re-verified first. The family status value itself was re-read and unchanged, and it carried the run’s one accepted-over-dissent family claim. sources.json217 217 16 lastCheckedDate values moved to 2026-08-31, one per page re-read this run. No source was added: gate-source-approval recorded zero proposed sources across all six bundles, so every date moved on a record that already existed. Sources are listed here rather than under the posted records below, which carry links and so hold only the collections the site routes. Not posted 1 item
Rejected by the review panel
releases/google-gemini-3-1-pro-preview.verifiedAtRefused 2-of-3 under the pilot bar, the run’s only rejection. Two rubrics found the evidence too weak for a status claim: the specification block flattens to an unseparated run of text once tags are stripped, so "Status" and "Preview" are adjacent rather than demonstrably paired, and the page’s marketing name "3.1 Pro" does not distinguish the Gemini API serving from the Vertex AI serving. The record keeps 2026-08-27. Its source record still had lastCheckedDate advanced, because that claim is about the page having been read rather than about what the page proves.Blocked byprovenance rubriceditorial rubric
What this run does not prove
- No human reviewed any claim, verdict or edit in this run before it merged and deployed. That is what ADR 0003 authorises and what it costs.
- The three rubrics are three instances of the same model family reading the same page, so they share a failure mode: a source that is itself wrong can carry all three. The panel defends against wrong-but-invalid and against unevidenced inference; it is a weak defence against a primary source that is confidently mistaken.
- Three claims were accepted over a provenance dissent and are published as such. Two of them — the Llama 4 Scout context window and the fit statement resting on it — turn on a table-row fragment whose column headers the quote does not carry. The majority judged the row unambiguous in context; the dissent judged the fragment insufficient on its own. A reader who agrees with the dissent should treat those two verifiedAt advances as the weakest edits in this run.
- A green run proves the recorded facts still match the pages this run read. It does not prove those pages are complete, and it says nothing at all about the 33 long-tail creators or the seven Google releases whose sources could not be reached.
- pagesFetched counts the 32 pages that were fetched successfully. Three further attempts against ai.google.dev failed outright and are excluded from that figure rather than counted as reads.
- reviewers and verdictsCast are derived rather than quoted: 6 bundles x 3 rubrics = 18 reviewer invocations, and 43 claims x 3 rubrics = 129 verdicts, which reconciles with the per-rubric verdict files.
- This run and run 2026-08-31-651a1c overlapped. That one merged first, as PR #665, while this one was mid-flight, so this run had to re-anchor from 36ce8b1b onto 356989e9 and re-run every gate. The re-anchor was clean — all 43 claims still matched the dataset afterwards — but two refreshes racing on the same documents is the failure mode the preflight check is meant to prevent, and the check only looks for an open pull request at the moment the run starts. Two same-day runs appear in this ledger as a result.
Follow-ups — proposed, not fixed
- The long-tail sweep this run skipped should be the next run’s first priority, not its second. Two consecutive runs covering different subsets is fine; a pattern of pilot-only runs would let the long tail rot while every entry still reads green.
- ai.google.dev could not be reached from this runner at all. If that is a durable block rather than a transient one, the Google releases need an alternative approved origin recorded in the reviewed catalogue — deepmind.google covers only the Pro page — or seven records will keep aging out with no way to re-verify them.
- The model-fit guidance layer pinned five records for a full run before this one cleared it. Re-verifying model-fit-statements.json first should become the standing order for a re-verification run rather than a lesson each run rediscovers, since guidance is cheap to re-check and blocks everything beneath it.
- Client-rendered marketing pages — aws.amazon.com/nova/ this run, and the OpenAI and Anthropic docs before the .md convention was found — yield placeholder text to a plain fetch. The .md suffix that works for OpenAI and Anthropic docs is worth recording in the reviewed catalogues so future runs do not rediscover it, and an equivalent for the AWS pages is worth looking for.
- The preflight check for a racing refresh only looks for an open pull request at the moment the run starts, so it cannot see one opened afterwards — which is exactly what happened here with PR #665. A run that re-checks immediately before it opens its own pull request, or that simply expects to re-anchor, would lose less work than one that discovers the race at merge time.
- Status claims sourced from flattened specification tables were refused this run for the second time in two runs, on the same structural grounds as the previous run’s context-window refusals. Either the scout should quote enough surrounding structure to fix the label-value pairing, or status re-verification should target a page that states it in prose.
2026-08-31-651a1c
Long-tail family depth, tranche 1 of abdeslam-menacere/ModelTree#651
Scope requested: Six creators from abdeslam-menacere/ModelTree#651 tranche 1: alibaba-cloud, zhipu-ai, moonshot-ai, baidu, tencent, 01-ai. Breadth rather than re-verification — asking whether each has a further documented family generation, not re-checking facts already recorded. No creator outside the tranche was scouted and no existing record was edited.Published9 edits posted · 7 items withheldA source-backed breadth run over six creators that each carried exactly one model family — alibaba-cloud, zhipu-ai, moonshot-ai, baidu, tencent and 01-ai — asking of each whether it has further documented family generations. Two do, on evidence that meets the bar: Alibaba published dated Qwen3.5 and Qwen3.6 news in the QwenLM/Qwen3.8 repository README, and 01.AI dated the original Yi series in 01-ai/Yi. Four do not, and the reason is the same one refresh run 2026-08-30-605b1a hit: their pages carry no dated release statement at all. The GLM, Kimi, Hunyuan and ERNIE repository READMEs have no dated news list, every date on their Hugging Face cards is Hub createdAt or lastModified metadata rather than a creator statement, and the z.ai GLM-5 blog URLs return 200 with a client-rendered shell under 600 bytes and no text. Nine claims were proposed, three rubrics voted independently on every one, and all nine met their creator threshold — the six Alibaba claims against a 2-of-3 pilot bar, the three 01.AI claims unanimously against the 3-of-3 long-tail bar. Three met it over a recorded provenance objection about mapped vocabulary fields, and are applied rather than overruled. Four creators are withheld, on the record and without softening. No human reviewed any part of the research; the run hands off at a reviewable commit for the independent review and QA gates.
- PreflightRanAnchor 8e8c319e, selected by merge-base with refs/remotes/origin/main and computed by each gate rather than supplied — requestedBase was null in every report. The approved-origin catalogue parsed cleanly at that anchor: 214 dataset sources and 18 profile catalogues out of 19 profile files, yielding 46 approved origins; the one profile listed without a catalogue is tools/updater/profiles/generic/long-tail.json, which configures no origins. Both npm and drydock were probed in both shim forms before either was relied on or ruled out: the bare names resolve to PowerShell shims the execution policy refuses, while npm.cmd reports 11.9.0 and drydock.cmd reports 0.1.0. Both are installed and blocked in bare form, not absent, and the execution policy was not changed. Dependencies were installed with npm ci; package-lock.json is unmodified.
- ScoutRanSix creators scouted against 48 pages, every one fetched and hashed during the run — retrieval is "fetch" on all 43 evidence entries and no claim rests on a search snippet. Each quote was verified to be literally present in the hashed body it is attributed to, by a builder that throws rather than emitting an unverifiable pair. The decisive finding is negative and general: model cards do not state release dates. The only pages in this tranche that date a release in the creator's own words are two GitHub README news lists, QwenLM/Qwen3.8 and 01-ai/Yi. Red herrings ruled out rather than used: PaddlePaddle/ERNIE's dated "Recent updates" dates the ERNIEKit toolkit and not a model; the Kimi-K2.5 changelog entry for 2026.1.29 is a system-prompt removal; Kimi-K3's "July 9, 2026" dates an evaluation branch; and Tencent's "Following the Hy3 Preview launch in late April" is about the preview and is not a valid partial date.
- ReviewRanThree rubrics — provenance, consistency, editorial — voted independently on all nine claims, each blind to the others and to the scout's reasoning, seeing the claim bundle and the committed dataset only. 27 verdicts cast, every one with a rationale, all published verbatim. Consistency and editorial accepted all nine. Provenance accepted six and objected to three, in each case that a quote naming "Type: Causal Language Model with Vision Encoder" or open-weight availability does not by itself force the mapped modality or category vocabulary. All three sit in the alibaba-cloud bundle, whose threshold is 2-of-3 because alibaba-cloud is the only creator of the six with a reviewed profile on disk in tools/updater/profiles. They therefore met their threshold over a recorded objection and were applied, not overruled; they are listed in dissents. No claim was revised and re-reviewed to chase a verdict, no reviewer was re-run, and no threshold was adjusted.
- GatesRangate-evidence and gate-source-approval ran across all six bundles before anything was applied, in that order and after the panel rather than before it. gate-dataset, npm run validate, gate-scope, gate-ledger and ci-preflight.mjs ran afterwards. No gate was skipped, forced, or re-run to obtain a different answer. gate-scope reports exit 1 by design and is the one number worth reading twice — see its entry in evaluated.gates and the caveats.
- PublishRanNine edits applied across three dataset documents in dependency order — sources, then families, then releases — and committed on the dock branch for abdeslam-menacere/ModelTree#651 together with this entry. Nothing was pushed, no pull request was opened and no merge was attempted: this run is a Drydock dock and its work ends at a reviewable commit, with the pull request and the merge belonging to the coordinating session after the independent review and QA gates have passed against that commit.
- DeployNot runNothing was pushed or merged by this run, so no Pages deploy was triggered and none could be observed. Recorded as not-run rather than not-applicable: a deploy is applicable to a dataset change, it simply has not happened yet at the time this entry was written.
What was found
- Scouts
- 6
- Pages fetched and hashed
- 48
- Claims proposed
- 9
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims alibaba-cloud pilot 2-of-3 6 01-ai long-tail 3-of-3 3 zhipu-ai long-tail 3-of-3 0 moonshot-ai long-tail 3-of-3 0 baidu long-tail 3-of-3 0 tencent long-tail 3-of-3 0 What those claims proposed to do Kind Count Effect Add 9 Three sources, three families and three releases: the Qwen3.5 and Qwen3.6 generations with one sourced open-weight variant each, and the original Yi series with Yi-34B-Chat. Not covered
- The other dated Qwen3.5 and Qwen3.6 variants — Qwen3.5-122B-A10B, Qwen3.5-35B-A3B and Qwen3.5-27B on 2026-02-24, and the 2026-04-22 Qwen3.6 entry. The QwenLM/Qwen3.8 news list dates them, but no model card was fetched and hashed for them this run, so no release record could carry a sourced parameter count, context window or licence.
- The other dated Yi releases — Yi-34B-200K on 2023-11-05, Yi-VL on 2024-01-23, Yi-9B on 2024-03-06 and Yi-9B-200K on 2024-03-16. Same reason: dated in 01-ai/Yi, but no card fetched for them here.
- Lineage between the new and existing families. Nothing wires Qwen3.5 to Qwen3.6 to the committed Qwen3.8, or the new Yi family to the committed Yi-1.5, because predecessorIds and successorIds on an existing record would be a change claim against a record this tranche was not asked to touch.
- The 34 creators outside this tranche, including the other 26 that still carry exactly one family.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 27
- Accepted by panel
- 9
- Rejected by panel
- 0
Deterministic gates and required checks — 6 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidenceall six claim bundles, before any dataset document was touched 0 Passpassed=true on each of the six. alibaba-cloud: 6 claims, 6 applicable, threshold 2 under the pilot policy. 01-ai: 3 claims, 3 applicable, threshold 3 under long-tail. The four withheld bundles carry 0 claims and pass trivially. The threshold was derived by the gate from the reviewed-profile set on disk and not from any bundle's policy field. gate-source-approvalall six claim bundles, anchored at merge-base 8e8c319e 0 Passpassed=true on each of the six, over 43 citations. Anchor 8e8c319e, selectedBy "merge-base with refs/remotes/origin/main", requestedBase null, 214 dataset sources and 46 approved origins. Inherited sources: qwen3-8-repository, osi-approved-licenses, 01-ai-yi-repository. Proposed sources: qwen3-5-397b-a17b-model-card, qwen3-6-35b-a3b-model-card, 01-ai-yi-34b-chat-model-card. Every citation sits on an origin the anchor already approves — huggingface.co, github.com and opensource.org — so the run introduced no new origin and approved nothing of its own. gate-datasetweb/src/data after the nine claims were applied 0 Passpassed=true, failures empty. Counts after: 217 sources, 48 publishers, 40 organizations, 64 families, 92 releases. npm run validateweb/ — the full vitest suite plus Astro and TypeScript diagnostics 0 Pass2361 tests across 104 files. It failed on the first run, and the failure was real rather than incidental: the comparison picker index budget. Recorded in caveats, because the repair reaches outside the dataset. gate-scopenot requiredthe branch diff against merge-base 8e8c319e 1 FailExit 1, and the correct answer. The diff reaches web/src/lib/comparison.test.ts, which is outside the nine documents ADR 0003 lets merge unattended, so this change is not eligible for an unattended merge and the gate says so. It is recorded as not required because this run never sought that route: it is a Drydock dock handing off to the independent review and QA gates and then to a human-opened pull request, where a scope refusal is a fact about eligibility rather than a blocked merge. The change was not trimmed to satisfy the gate; the reason it reaches out of class is in the caveats. gate-ledgerthis entry against the branch diff, anchored at merge-base 8e8c319e 0 Passpassed=true. The three declared documents match the three dataset documents the branch changed, in both directions, and their record counts were counted at the anchor and in the working tree rather than taken from this entry. ci-preflightrepository root — the pull-request checks this branch's diff actually triggers 0 PassSelected and ran the checks the diff triggers, measured from the computed merge-base: 5 files changed, 3 of 7 local check groups selected, all three passing — web-ci, skills-ci and source-link-health-tests. It does not cover the networked link-health sweep or the second Python interpreter, and says so on every run; a green preflight is therefore not a green CI. It was run twice: the first run reported web-ci as failing with "npm run test" exiting 1, and the second passed with 104 test files and 2,361 tests. See the caveats — the first result was never reproduced and is recorded rather than explained away. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
qwen3-6-family-addProvenanceThe date and first-open-weight-variant facts are quoted, but no attached quote for this claim supports the proposed multimodal-generalist category. The quoted Qwen3.6 text mentions open-weight availability and coding/repository reasoning, not vision or multimodality.alibaba-qwen3-5-397b-a17b-release-addProvenanceThe date, weights, Apache licence, OSI badge, parameters, and 262,144-token context are quoted. But "Type: Causal Language Model with Vision Encoder" does not by itself force inputModalities ["text","image"] or the complete output modality list.alibaba-qwen3-6-35b-a3b-release-addProvenanceThe model-card quotes support open weights, Apache licensing, parameters, and native context. The attached "Type: Causal Language Model with Vision Encoder" quote does not force the proposed text+image input modalities or exclude other vision inputs.
Posted 9 edits
9 edits across 3 documents, a net change of 9 records.
Dataset documents this run changed Document Before After What changed sources.json214 217 Three Hugging Face model cards: Qwen3.5-397B-A17B, Qwen3.6-35B-A3B and 01-ai/Yi-34B-Chat. They are not listed individually under records because the refresh page resolves links for release and family records only, and a source record listed there would link nowhere. families.json61 64 The Qwen3.5 and Qwen3.6 generations, and the original Yi series distinct from the committed Yi-1.5. releases.json89 92 One sourced variant per new family, being the one whose model card was fetched and hashed this run. Each document links to the file as this run left it, not as it stands today.
Records added
- Qwen3.5in the treefamilies
qwen3-5First release 2026-02-16, per the QwenLM/Qwen3.8 news list. - Qwen3.6in the treefamilies
qwen3-6First release 2026-04-16, per the same news list. - Yiin the treefamilies
01-ai-yiFirst release 2023-11-02, the original Yi series; the committed Yi-1.5 family stays separate. - Qwen3.5-397B-A17Bpassportreleases
alibaba-qwen3-5-397b-a17b397B total, 17B active, 262,144 native context, Apache-2.0. - Qwen3.6-35B-A3Bpassportreleases
alibaba-qwen3-6-35b-a3b35B total, 3B active, Apache-2.0. - Yi 34B Chatpassportreleases
01-ai-yi-34b-chatDated 2023-11-23 by the 01-ai/Yi news list. Parameters and context window deliberately omitted, matching the committed Yi-1.5 34B Chat record.
Not posted 7 items
Verification date deliberately held back
zhipu-ai-further-generationsGLM-5, GLM-5.2, GLM-5.3 and GLM-5.3-Flash all exist as Hugging Face repositories under zai-org, and none of the pages fetched states a release date. The GLM repository READMEs carry no dated news list — scans for YYYY-MM-DD, YYYY.MM.DD and month-name forms returned nothing — and every date visible on the cards is Hub createdAt or lastModified metadata, which refresh run 2026-08-30-605b1a already ruled out as a release date. https://z.ai/blog/glm-5 and https://z.ai/blog/glm-5.3 return HTTP 200 but are client-rendered shells under 600 bytes with no text, and https://z.ai/blog is 404. A family cannot be added without a sourced firstReleaseDate, so nothing was proposed.Blocked byno dated release statement on any fetched pagez.ai blog pages render client-side and serve no textmoonshot-ai-further-generationsKimi-K2.5 and Kimi-K3 exist as Hugging Face repositories under moonshotai, and neither card nor the MoonshotAI GitHub organisation pages state a release date. The two date-like strings found were checked and are not release dates: the Kimi-K2.5 changelog entry for 2026.1.29 records a system-prompt removal, and Kimi-K3's "July 9, 2026" dates an evaluation branch. Every other date is Hub metadata.Blocked byno dated release statement on any fetched pagebaidu-further-generationsERNIE-5.0 material and further ERNIE-4.5 variants exist under the baidu Hugging Face namespace, and no fetched page dates a release. PaddlePaddle/ERNIE does carry a dated "Recent updates" list, but it dates the ERNIEKit toolkit rather than a model, which is a different entity and was not used.Blocked byno dated release statement on any fetched pagethe only dated list found dates a toolkit, not a modeltencent-further-generationsHunyuanImage-3.0, HunyuanOCR and Hy3/Hy4 material exist under Tencent-Hunyuan, and no fetched page dates a release. The nearest statement, "Following the Hy3 Preview launch in late April", is about the preview rather than the release and "late April" is not a valid partial date in this schema. Recording it as one would be the guess this run exists to refuse.Blocked byno dated release statement on any fetched page"late April" is not a representable date
Sources conflict, so no value changed
01-ai-yi-licence-conflictThe Yi-34B-Chat card states in prose that "The code and weights of the Yi series models are distributed under the Apache 2.0 license", while the 01-ai/Yi README records "2023-11-23: The Yi Series Models Community License Agreement is updated to v2.1". Both quotes are attached to the release claim and the disagreement is left explicit in the record's prose rather than resolved. The structured licence follows the model card, which is the artefact the release record is about.Blocked bytwo primary pages of the same creator disagree
Out of the run’s reach
qwen3-5-and-qwen3-6-additional-variantsThe QwenLM/Qwen3.8 news list dates Qwen3.5-122B-A10B, Qwen3.5-35B-A3B and Qwen3.5-27B to 2026-02-24 and a further Qwen3.6 entry to 2026-04-22. The dates are sourced, but no model card was fetched and hashed for those variants this run, so no release record could carry a sourced parameter count, context window or licence. One release per new family was proposed rather than an unsourced set.Blocked byno model card fetched for these variants this run01-ai-yi-additional-variantsThe 01-ai/Yi news list dates Yi-34B-200K to 2023-11-05, Yi-VL to 2024-01-23, Yi-9B to 2024-03-06 and Yi-9B-200K to 2024-03-16. Same reason as the Qwen variants: dated, but no card fetched for them here.Blocked byno model card fetched for these variants this run
What this run does not prove
- The outcome is recorded as "published" because that is the only value the schema admits for a run that changed dataset documents, and it is not the whole truth at the time of writing. This run is a Drydock dock: its nine edits are applied and committed on a branch, and nothing has reached main. No pull request existed when this entry was written, which is why references names no pull request of this run's own — the pull-request reference below is the change whose merge set this run's baseline, labelled as such. Whoever opens the pull request should add it, its merge commit and the Pages deploy to references.
- npm run validate failed on its first run, on the comparison picker index page-weight budget, and the repair reaches outside the dataset into web/src/lib/comparison.test.ts. The budget was raised from 10,240 to 11,264 bytes as a deliberate page-weight decision, on the test message's own terms: measured at merge-base 8e8c319e the index was 10,115 bytes over 89 releases (113.65 per release) and at the tip it is 10,449 over 92 (113.58 per release), so the catalogue simply grew and the per-release figure did not move. The scale-invariant guard of 128 bytes per row was not touched and keeps 14 bytes of headroom. That one file is why gate-scope exits 1 and why this change cannot merge unattended. It is also worth knowing that the merge-base already sat 125 bytes under the old budget — about one release — so this was going to land on whichever tranche came first.
- ADR 0005's accepted limit applies to every hash and quote here: gate-evidence checks that a content hash is well-formed and that a quote is long enough, never that either matches the remote page, and both are self-authored by the run. What compensates for it in this run is mechanical but local: every page was fetched to disk, hashed from the bytes on disk, and every quote was checked to be literally present in the body it is attributed to, by a builder that throws rather than emit an unverifiable pair. That makes the pairs internally consistent. It does not make them independently verified, and a later reader who wants certainty must refetch.
- Three claims were applied over a provenance objection, all of them about controlled-vocabulary fields — the modality lists and one family category. The objection is recorded in full in dissents and is not answered here. A reader who thinks provenance was right should read those three records as the weakest in this change.
- status: "current" on the three new releases is the least directly sourced mapping in the run. No page states a lifecycle state in words; the value rests on the news lists saying the weights are released and available and on the absence of any deprecation or legacy notice. Qwen3.5 and Qwen3.6 are both superseded by the committed Qwen3.8, and the original Yi by Yi-1.5, so a reader who reads "current" as "newest" will be misled. It means "not withdrawn".
- The input and output modality lists on the two Qwen releases record text and image while both cards also carry a video-input quickstart. The competing signal is disclosed in the source notes and the record summaries rather than resolved, and it is the substance of two of the three provenance objections.
- Four of the six creators were withheld, and no negative claim is being made about them: the run did not establish that GLM, Kimi, Hunyuan or ERNIE have no further generations, only that no page it fetched states a release date for one. A page it did not fetch, or a page rendered server-side rather than in the browser, could settle any of the four.
- Every fetch happened during one run on 2026-08-31 and each contentHash is a snapshot of that moment. Hugging Face and GitHub pages change without notice, so a refetch that produces a different hash is expected rather than evidence of an error.
- One test result in this run was not reproducible. The first ci-preflight run reported web-ci failing with "npm run test" exiting 1, and printed no failing test name that was captured. Every subsequent run was green — a second ci-preflight with 104 test files and 2,361 tests passing, a bare npm run test with the same figures and the coverage verifier confirming all 104 discovered files reported, and two npm run validate runs. Four green runs do not turn a red one into a flake with any certainty, so it is recorded here as unexplained rather than dismissed, and a reviewer seeing web-ci fail once on this branch should suspect it is the same thing and not a new one.
Follow-ups — proposed, not fixed
- The committed release dates for moonshot-ai-kimi-k2-instruct (2025-07-11), zhipu-ai-glm-4-5-air (2025-07) and baidu-ernie-4-5-300b-a47b-pt (2025-06-28) are each exactly the Hugging Face Hub createdAt value of the cited repository, which refresh run 2026-08-30-605b1a explicitly ruled out as a release date. This run found that while looking for where those dates came from, did not fix it because the records are outside this tranche's scope, and records it here so it is not found a third time.
- Lineage between generations is unwired: nothing connects Qwen3.5 to Qwen3.6 to Qwen3.8, or the Yi family to Yi-1.5, although the Yi-1.5 card states it was continuously pre-trained from Yi. Wiring it means editing existing records and is worth its own issue.
- The remaining creators of the 32 that carried exactly one family are untouched by this tranche; 30 still do after it.
- The comparison picker index budget has now been raised three times as the catalogue grew, each time by 1,024 bytes. Whether the row itself should shrink — four string fields per release, none of them abbreviated — is a page-weight question nobody has asked yet, and asking it once would be cheaper than raising the number a fourth time.
2026-08-30-c0b6e9
Data refresh 2026-08-30
Scope requested: Every creator in organizations.json — all 33 — with pilot creators judged 2-of-3 and long-tail creators 3-of-3. Re-verification only: confirming recorded facts against their primary sources and moving verification dates forward, not adding breadth.Published36 edits posted · 9 items withheldAn agent-run, source-backed refresh under ADR 0003, and a re-verification run rather than a discovery one: not a single recorded fact value changed. All 33 creators in organizations.json were scouted, each against at least one primary page fetched and hashed this run, and 20 of them produced claims — 45 in total. Three rubrics voted blind on every one; 38 met their creator’s threshold and 7 long-tail claims died one vote short. 36 edits landed across two documents, all of them a verification date moving to 2026-08-30, before GitHub merged PR #597 on a green web-ci and Pages deployed successfully. The most useful finding is a refusal rather than an edit: five licence re-verifications failed because a model card saying “Apache-2.0” does not source the osiApproved flag recorded beside it, which is a claim about what OSI decided. No human reviewed any part of it. This entry was transcribed after the fact — see the caveats.
- PreflightRanAnchor 7ca5802e, selected by merge-base with refs/remotes/origin/main and computed by the gate rather than supplied — requestedBase was null. The approved-origin catalogue parsed cleanly at that anchor: 198 dataset sources and 12 profile catalogues out of 13 profile files, yielding 41 approved origins. The one profile listed but not drawn on is tools/updater/profiles/generic/long-tail.json, which configures no origins. The anchor is therefore full rather than silently narrowed.
- ScoutRanAll 33 creators were scouted, each against at least one primary page fetched and hashed this run; every quote in the pull request body was sliced from a page read during the run rather than recalled. Three creator hosts refused automated fetches and were worked around on approved origins instead: openai.com/news/ (403) via openai.com/news/rss.xml, ai.meta.com/blog/ (403) via huggingface.co/meta-llama, and x.ai/news (403) via docs.x.ai/developers/models. No claim rested on an unreachable page. 13 of the 33 creators produced no claim, because nothing they publish had changed in a way this run could evidence.
- ReviewRanThree rubrics — provenance, consistency, editorial — voted independently on all 45 claims, each blind to the others and to the scout’s reasoning. Consistency and editorial accepted every claim; provenance accepted 38 and rejected 7. No claim was revised and re-reviewed to chase a verdict. Every verdict and rationale is published verbatim in the pull request body.
- GatesRangate-evidence and gate-source-approval ran across all 20 bundles before anything was applied; gate-dataset, npm run validate, gate-scope and ci-preflight.mjs ran afterwards. No gate was skipped, forced, or re-run to obtain a different answer, no threshold was lowered, and nothing was pushed to main. gate-scope reported changed: 2, empty: false, outOfClass: [], so the change sat inside the ADR 0003 qualifying class and was eligible to auto-merge.
- PublishRanPR #597 carried the full evidence trail and was merged by GitHub via --auto once web-ci was green, not by the agent. Merge commit 2f490766, merged 2026-08-30T11:37:33Z.
- DeployRanPages deploy run 33309402080 on 2f490766 succeeded and https://abdeslam-menacere.github.io/ModelTree/ served HTTP 200. No revert needed.
What was found
- Scouts
- 33
- Pages fetched and hashed
- 45
- Claims proposed
- 45
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims 01-ai long-tail 3-of-3 2 ai2 long-tail 3-of-3 2 amazon pilot 2-of-3 2 anthropic pilot 2-of-3 7 baidu long-tail 3-of-3 2 bytedance-seed long-tail 3-of-3 2 cohere long-tail 3-of-3 2 databricks long-tail 3-of-3 2 deepseek long-tail 3-of-3 2 eleutherai long-tail 3-of-3 2 hugging-face long-tail 3-of-3 2 ibm long-tail 3-of-3 2 lg-ai-research long-tail 3-of-3 2 moonshot-ai long-tail 3-of-3 2 nvidia long-tail 3-of-3 2 sarvam-ai long-tail 3-of-3 2 snowflake long-tail 3-of-3 2 upstage long-tail 3-of-3 2 xai long-tail 3-of-3 2 zhipu-ai long-tail 3-of-3 2 What those claims proposed to do Kind Count Effect Unchanged 24 A recorded model fact re-read from its primary source and found to still hold. 17 reached their threshold; of those, 15 advanced a release verifiedAt to 2026-08-30 and 2 were withheld. The other 7 were refused by the provenance rubric. Change 21 A source page re-read this run, moving its lastCheckedDate to 2026-08-30. All 21 were accepted unanimously and all 21 were applied. Not covered
- 13 of the 33 creators scouted produced no claim at all. That is an honest zero per creator rather than a skipped stage, but it also means their records carry this run’s attention without carrying its verification date.
- Three creator hosts refused automated fetches this run — openai.com/news/ (403), ai.meta.com/blog/ (403) and x.ai/news (403). Approved alternatives were used and no claim rested on an unread page, but the announcement indexes themselves went unread.
- Breadth remains uncovered by design: OpenAI’s realtime, transcribe and tts models, roughly twenty Cohere Command models against the one recorded, Microsoft’s MAI-Code-1.1-Flash, MAI-Image-2.5, MAI-Voice-2 and MAI-Transcribe-1.5, and Amazon’s Nova Act, Nova Forge and Nova Multimodal Embeddings. Each is breadth work with its own sourcing burden, not re-verification.
- Microsoft MAI-Image-2.6, announced 2026-08-10, was deliberately not added. Its only availability statement names a product and a serving platform rather than the creator’s own release, so accessType has no creator-level statement and status is ambiguous between preview and current. It also has no recorded family. The ambiguity was left explicit rather than resolved by guess.
- xai grok-4.20 appears on docs.x.ai only in a passing note that logprobs are unsupported by "grok-4.20 and newer", never as a listed model. That is not enough to create a release.
What was evaluated
- Reviewers
- 60
- Verdicts cast
- 135
- Accepted by panel
- 38
- Rejected by panel
- 7
Deterministic gates and required checks — 7 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence.mjs20 claim bundles 0 Pass45 claims admissible with complete panels across 20 bundles. Thresholds were applied per profile — 2 for a pilot creator, 3 for a long-tail one — and no bundle omitted its policy. gate-source-approval.mjs20 claim bundles 0 PassZero proposed sources across all 20 bundles: every cited source was already in the dataset at anchor 7ca5802e, on one of the 41 origins approved there. No source was refused, and no origin was introduced. gate-dataset.mjsweb/src/data 0 Pass419 records validated across the documents raw.ts composes. npm run validateweb/ 0 Pass990 tests executed; astro check reported 0 errors. This is the gate that caught the two withheld verification dates, by refusing guidance older than the evidence beneath it. gate-scope.mjsbranch vs merge-base 0 Passchanged: 2, empty: false, outOfClass: []. Anchored on 7ca5802e, the merge base the gate computed rather than one supplied to it. ci-preflight.mjsnot requiredweb-ci, skills-ci, source-link-health-tests 0 PassRun locally before the branch was pushed; all three workflows predicted green. web-ciPR #597 — PassSUCCESS. GitHub performed the merge via --auto on this check going green; the agent did not merge. Posted 36 edits
36 edits across 2 documents, a net change of 0 records.
Dataset documents this run changed Document Before After What changed releases.json82 82 15 verifiedAt dates moved to 2026-08-30 — anthropic-claude-fable-5, anthropic-claude-opus-5 and anthropic-claude-sonnet-5 from 2026-08-26; ai2-olmo-2-7b, amazon-nova-2-sonic, cohere-command-a-plus-05-2026, deepseek-v4-pro, nvidia-nemotron-4-340b-base and xai-grok-4-6 from 2026-08-28; baidu-ernie-4-5-300b-a47b, bytedance-seed-oss-36b-instruct, databricks-dbrx-instruct, hugging-face-smollm3-3b, ibm-granite-4-2-30b and lg-ai-research-exaone-3-5-7-8b-instruct from 2026-08-29. No release was added or removed and no other field changed, so the record count is identical either side of the merge commit. sources.json198 198 21 lastCheckedDate values moved to 2026-08-30, one per page re-read this run. No source was added: the approved-source gate recorded zero proposed sources, so every date moved on a record that already existed. Sources are listed here rather than under the posted records below, which carry links and so hold only the collections the site routes. Each document links to the file as this run left it, not as it stands today.
Not posted 9 items
Rejected by the review panel
releases/01-ai-yi-1-5-34b-chat.verifiedAtRefused 2-of-3 under the long-tail unanimous bar. The Hugging Face card states "License: apache-2.0", but the recorded license object also asserts osiApproved, which is a claim about what OSI decided and needs an OSI source. The licence name cannot stand in for it.Blocked byprovenance rubricreleases/eleutherai-pythia-12b.verifiedAtRefused 2-of-3 on the same osiApproved grounds: the Pythia-12B card states the licence name and nothing about OSI approval status.Blocked byprovenance rubricreleases/snowflake-arctic-instruct.verifiedAtRefused 2-of-3 on the same osiApproved grounds; the Arctic card carries an apache-2.0 tag and no OSI provenance.Blocked byprovenance rubricreleases/sarvam-ai-sarvam-m-v1.verifiedAtRefused 2-of-3 on the same osiApproved grounds. Its source record still had its lastCheckedDate advanced, because that claim was about the page being read rather than about what the page proves.Blocked byprovenance rubricreleases/zhipu-ai-glm-4-5-air.verifiedAtRefused 2-of-3. The GLM-4.5-Air card states the models are released under the MIT open-source licence, which supports the licence name and still does not cite OSI approval for the osiApproved flag.Blocked byprovenance rubricreleases/moonshot-ai-kimi-k2-instruct.verifiedAtRefused 2-of-3. The card states "Context Length 128K" while the dataset records 131072. The quote does not force the binary 128x1024 reading, so the stored precision is not directly sourced — an assumed multiplier, not a stated number.Blocked byprovenance rubricreleases/upstage-solar-pro-preview-instruct.verifiedAtRefused 2-of-3 on the same binary-multiplier grounds: the page states "a maximum context length of 4K" and the dataset records 4096.Blocked byprovenance rubric
Verification date deliberately held back
releases/anthropic-claude-mythos-5.verifiedAtHeld at 2026-08-15 despite the claim reaching its threshold. validate.ts forbids guidance dated before the evidence beneath it, and fit-claude-mythos-5-limited-availability, verified 2026-08-15, rests on this release. Advancing the release would have stranded that fit statement behind its own evidence; npm run validate failed exactly this way when it was attempted. Re-dating guidance this run gathered no evidence for would have been the larger claim, so the write was withheld rather than the rule weakened.Blocked byfit-claude-mythos-5-limited-availabilityvalidate.tsreleases/anthropic-claude-haiku-4-5.verifiedAtHeld at 2026-08-15 for the same reason, against fit-claude-haiku-4-5-release-current. Two accepted claims applying nothing is the honest outcome here: the panel judged the model facts still true, and the dataset still refuses to say so until the guidance resting on them is re-verified too.Blocked byfit-claude-haiku-4-5-release-currentvalidate.ts
What this run does not prove
- This entry was transcribed after the run rather than by it. The run could not write its own line: refresh-runs.json is not one of the documents raw.ts composes, so gate-scope.mjs correctly reports it out of class, and including it would have cost the run its auto-merge. That gap is issue #419 and it has now recurred on every published run.
- pagesFetched is the summary issue’s approximate figure of "~45 pages" recorded as 45. Neither durable record states an exact count, so this number is the only one available rather than a counted one.
- reviewers and verdictsCast are derived, not quoted: 20 bundles x 3 rubrics = 60 reviewer invocations, and 45 claims x 3 rubrics = 135 verdicts, which reconciles with the published rubric tallies of 45, 45 and 45. Neither field has a verbatim source in the pull request body or the summary issue.
- The pull request body says "36 records, 38 changed lines". The merge commit shows 36 insertions and 36 deletions across the two documents, and the per-record diff reconciles to exactly 15 releases and 21 sources. Where the two disagree this entry follows the commit, as #419 already recommended after the same class of discrepancy on the 2026-08-27 run.
- A green run proves the recorded facts still match the pages this run read. It does not prove those pages are complete, that the 13 creators without claims are unchanged, or that anything on a host that refused the fetch is still current.
- No human reviewed any claim, verdict or edit in this run before it merged and deployed.
Follow-ups — proposed, not fixed
- osiApproved blocks routine licence re-verification for long-tail creators — five claims died on it this run. Either the recorded license objects need an OSI source attached once, or re-verification needs to target a licence sub-field a model card can actually source.
- Context windows recorded as exact integers cannot be re-verified against pages that state them as "128K" or "4K". Two claims died on the assumed 1024 multiplier.
- Three creator announcement indexes refuse automated fetches — openai.com/news/, ai.meta.com/blog/ and x.ai/news. The working alternatives this run used are worth recording in the reviewed catalogues so future runs do not rediscover them.
- Microsoft MAI-Image-2.6 needs a creator-level source before it can be recorded, following the precedent set for microsoft-mai-thinking-1.
2026-08-30-605b1a
Long-tail breadth tranche 2026-08-30 — Aleph Alpha, AI Singapore, Nous Research, Reka AI, Liquid AI, Xiaomi
Scope requested: Six creators with no reviewed profile in tools/updater/profiles — Aleph Alpha, AI Singapore, Nous Research, Reka AI, Liquid AI and the Xiaomi LLM-Core team — under the long-tail unanimous 3-of-3 policy, restricted to the 42 origins already approved at the merge-base. No profile catalogue was edited and no origin was added: widening the trust boundary is deliberately a human act.Ran, changed nothing0 edits posted · 10 items withheldResearched six long-tail creators to take the catalogue from 34 organizations to 40 and published none of them. Every fact was fetched and hashed from origins this repository already approves — huggingface.co, github.com, arxiv.org and opensource.org — never read from a search result, and gate-source-approval passed on all six bundles, so origin discipline was not the blocker. The panel was the blocker, and beneath it a structural one. Under the unanimous 3-of-3 long-tail bar the provenance rubric rejected 46 of 54 claims, most of them because the record asserted fields no attached quote stated: lifecycle status, modality lists, downloadable weights, and above all a release date. No approved-origin page states a release date for any of these six models. No model card states a release date for the model it describes. All six carry dates, and every date found belongs to a category that is not a release date: a training-data cutoff (Aleph Alpha, cutoff date 04/2023); Hub revision tags (AI Singapore, Current Version: `14.04.2025` at line 29, and a revision="18.12.2024" argument at line 123 under the comment "Specify the revision here"); paper years in BibTeX blocks (Nous Research, Liquid AI and Xiaomi, year={2025}); benchmark names (Reka AI, AIME-2024; Xiaomi, AIME 2024 and AIME 2025 across six table rows); values inside worked examples and generated output (Liquid AI, "date": "2023-11-20" in a tool-call sample; Aleph Alpha, the founding of Rome and a 2016 census figure in sample completions); EU legal instrument numbers that only resemble slashed dates (Aleph Alpha, Directive (EU) 2019/790 and Regulation (EU) 2016/679); and a changelog line dating a different model (Xiaomi, [2025.05.30], which dates MiMo-7B-RL-0530 and not the MiMo-7B-RL this run targeted). Each is named with its card and line so the exclusion can be checked. This states what the dates found are and why none is a release date; it is not a claim that no other date exists on these pages. The Hub API states createdAt, which is repository creation and not a stated launch; arXiv states when a paper was submitted, which is a fact about the paper and not the model; and the two GitHub repository READMEs that do carry dated release news date other models — SEA-Guard and a March 2026 embedding family for AI Singapore, and MiMo-7B-RL-0530 rather than MiMo-7B-RL for Xiaomi. releaseSchema requires releaseDate, and the schema is explicit that a field no approved source states means the record is withheld rather than guessed, so no release could be written for any creator. The eight claims the panel did accept were three sources, three publishers and two conflict records; applying them would have left sources cited by nothing, which validate.test.ts forbids, and a creator with no release renders in neither branch of the tree. They were dropped rather than applied. The dataset is unchanged and this entry is the only edit.
- PreflightRanRe-measured the trunk baseline from the committed data rather than quoting it: 34 organizations, 55 families, 83 releases, 202 sources, 42 publishers. Enumerated the approved origin set from gate-source-approval itself, by running it over a zero-claim probe bundle, which reports anchors.approvedOrigins whether it passes or fails: 42 origins at merge-base c9e01df.
- ScoutRanMade 56 fetches over 48 distinct pages — 39 initial targets across four batches, 1 ad-hoc re-check of the AI Singapore card, 6 systematic date re-checks which deliberately re-pulled bytes already held rather than trusting the stored copy, 6 fetches around the Hermes 3 withholding, being four Hermes-4 records, the Hermes-3 LICENSE whose 404 caused it and one arXiv abstract, and 4 GitHub organization and repository pages — of which 8 re-fetched a page already held — across model cards, Hub API records, config.json files, arXiv abstracts, OSI licence pages and GitHub organization and repository pages. The pagesFetched field below records that 56: it counts fetch operations that left a stored body, not distinct pages, of which there are 48. The 8-page gap is the re-fetches, and the AI Singapore card accounts for three of the stored bodies on its own — the original, the ad-hoc re-check and the systematic one — all three byte-identical at sha256:2d4396287f5d. The count is of page bodies only: date-recheck.json sits in the same folder and is the manifest of that re-check rather than a fetched page, so it is excluded. Every quote in every bundle was machine-checked to be a literal substring of the exact bytes whose sha256 the evidence records; the generator refuses to emit a bundle otherwise and caught six bad quotes on its first pass, one of which had only ever looked verified because it sat behind a function the prototype never called.
- ReviewRanThree reviewers, one vote each, launched in parallel and given the claim, its evidence, the dataset slice and one rubric — never the chair reasoning, another reviewer output, or a running tally. 162 verdicts over 54 claims. The editorial reviewer first returned a single 138-character rationale repeated 54 times and was sent back for claim-specific rationales with its votes explicitly untouched; the rewrite changed no vote. Provenance accepted 8, consistency 54, editorial 54.
- GatesRangate-evidence exit 1 on all six bundles, deriving the 3-of-3 threshold itself from the reviewed-profile set on disk rather than from any policy field in the bundles. gate-source-approval exit 0 on all six. gate-dataset exit 0 on the unchanged dataset. gate-scope reports this log entry out of the ADR 0003 qualifying class, which is expected rather than a problem to work around.
- PublishNot applicableNo claim reached the dataset, so there was nothing to publish. The dock agent opened no pull request and recorded no gate verdict on its own work; this run ends at the review gate and hands off to an independent reviewer.
- DeployNot applicableNothing was merged, so nothing deployed.
What was found
- Scouts
- 6
- Pages fetched and hashed
- 56
- Claims proposed
- 54
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims Aleph Alpha long-tail 3-of-3 8 AI Singapore long-tail 3-of-3 9 Nous Research long-tail 3-of-3 10 Reka AI long-tail 3-of-3 8 Liquid AI long-tail 3-of-3 10 Xiaomi LLM-Core Team long-tail 3-of-3 9 What those claims proposed to do Kind Count Effect Add 52 Six creators as organization, family, release and publisher records, with 28 supporting source records. None reached the dataset. Conflict 2 Two disagreements between primary sources, recorded rather than smoothed: the LFM2-1.2B context window and the Hermes-4-14B parameter total. Both were accepted unanimously and both describe releases that were not published. Not covered
- No creator-owned domain was consulted for any of the six, because none is an approved origin. What those newsrooms say about release dates, lifecycle status and licensing is therefore unknown to this run rather than absent from the world.
- The /compare page-weight guards in web/src/lib/comparison.test.ts were never exercised against new data, because no release was added. Whether six more creators would breach the per-release or total budget remains untested and is still an open question for whoever owns that guard.
- Hermes-3-Llama-3.1-8B was withheld before review rather than judged by the panel, so no verdict exists on it.
- Only one review round was run. The rubric remedy for an unquoted field is to attach the quote, and for the fields other than releaseDate that remedy was not attempted, because releaseDate alone is sufficient to withhold every release record.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 162
- Accepted by panel
- 8
- Rejected by panel
- 46
Deterministic gates and required checks — 4 of 6 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidenceAll six claim bundles, annotated with the panel verdicts. 1 FailRefused 46 claims that did not reach the unanimous long-tail bar. The gate derived the 3-of-3 threshold from the reviewed-profile set on disk rather than from the policy field in the bundles, which is the behaviour ADR 0002 requires, and it reached the same answer as the panel independently. gate-source-approvalAll six claim bundles against the 42 origins approved at merge-base c9e01dffbf0d9bc92d0b6ee39f286e22120d7246. 0 PassEvery citation in every bundle sits on an already-approved origin, and no proposed source rests on a creator-owned domain. This is the one gate that passed cleanly on all six, and it is worth recording that the run did not fail on origins: it failed on what the quotes state. gate-datasetThe committed dataset, unchanged by this run. 0 PassThe dataset is unchanged and remains internally coherent: 202 sources, 42 publishers, 34 organizations, 55 families, 83 releases. No orphaned source, no dangling reference, no release predating its family. gate-scopeThe branch diff from its merge-base to its tip. 1 FailReports web/src/data/refresh-runs.json as outOfClass, which is correct and expected rather than a problem to work around: the ADR 0003 qualifying class is the dataset documents raw.ts composes, and the refresh log is deliberately not one of them, so any run that logs itself leaves the class by construction. The gate was not modified, the file was not added to ALLOWED_PATHS, and the log entry was not dropped to stay in class. The consequence is that this change is outside ADR 0003 and merges the ordinary way, with a human — which is what this run wanted anyway, since it published no data. npm run validateweb/ — the full vitest suite and Astro/TypeScript diagnostics. 0 PassRun from web/ after this entry was written, so it covers the entry itself against refresh-log-schema.ts as well as the unchanged dataset. node .github/scripts/ci-preflight.mjsThe pull-request checks this branch diff actually triggers, measured from the merge-base. 0 PassRun from the repository root. It does not cover the networked link-health sweep or the second Python interpreter, which it prints on every run and which remain unrun here. Posted 0 edits
Nothing reached the dataset. No branch, no commit, no pull request.
Not posted 10 items
Rejected by the review panel
aleph-alphaPharia-1-LLM-7B-control. The parameter count and the 8,192-token sequence length are quoted from Aleph Alpha own card, but no attached quote states a release date, a lifecycle status, a modality set or that the weights are downloadable, and the OSI licence index heading does not by itself establish that the Open Aleph License is absent from the register. The organization record additionally inferred a German identity from a Heidelberg location and did not establish the relationship between the names "Aleph Alpha GmbH", "Aleph Alpha Research" and "IPAI Aleph Alpha Research GmbH", all three of which appear across the sources.Blocked byprovenanceai-singaporeLlama-SEA-LION-v3-8B-IT. The card describes AI Singapore in its own words as a national programme supported by the National Research Foundation and hosted by the National University of Singapore, which is what typed the organization research-lab, but the release record rested on an unquoted date, status, modality set and weights statement, and the 128k-to-131072 reconciliation was not supported by an attached quote.Blocked byprovenancenous-researchHermes-4-14B. Licence and parameter assertions were not carried by the attached quotes, the website was not quoted, and the release date, status, modalities and weights were unstated. The config source and the publisher record were accepted; the records that would have made them mean something were not.Blocked byprovenancereka-aiReka Flash 3. The card states that the model was trained from scratch, which is why no derivation was recorded despite the configuration declaring a Llama architecture class, but the publisher attribution, the website and the release date, status, modality and weights fields were all unquoted.Blocked byprovenanceliquid-aiLFM2-1.2B. The organization identity and the LFM Open License name were not carried by the attached quotes, and the release date, modalities and status were unstated. The GitHub organization is Liquid4All rather than a name matching the model prefix, which the record noted but did not resolve from a quote.Blocked byprovenancexiaomiMiMo-7B-RL. The author attribution, the governance relationship between the LLM-Core team and Xiaomi, and the web endpoints were unquoted, as were the release date, status, modalities and weights. The GitHub organization XiaomiMiMo publishes neither a website nor a location, which is why the record proposed the GitHub profile itself as the organization website rather than asserting xiaomi.com.Blocked byprovenance
Accepted by the panel, then dropped
accepted-sources-and-publishersEight claims were accepted unanimously by all three rubrics and none was applied. Six of them were adds: three source records — the Hermes 4 14B and Reka Flash 3 configuration files and the Aleph Alpha GitHub organization profile — and three publishers, Nous Research, Liquid AI and AI Singapore. The remaining two were the conflict records described below. The accepted set spans five creators and completes none of them: Nous Research alone reached both a publisher and a configuration source; Liquid AI and AI Singapore reached a publisher and no source; Reka AI and Aleph Alpha reached a source and no publisher; Xiaomi reached nothing. With no accepted release or organization to cite them, the sources would have been orphaned, which validate.test.ts forbids outright, and a publisher publishing nothing is a record that renders nowhere. This is the same shape of outcome recorded for Mistral AI and Cohere in the 2026-08-27 run.Blocked byai-singaporealeph-alphanous-researchreka-ailiquid-ai
Verification date deliberately held back
nous-research-hermes-3-llama-3-1-8bWithheld before review and never put to the panel. This model was the tranche original Nous candidate and a prior session had flagged it as possibly uncitable; the flag was checked on all three of its legs and held. The card carries exactly one licence statement, the bare slug "license: llama3", while the same card declares its base model to be meta-llama/Meta-Llama-3.1-8B and the Hub tags declare base_model:meta-llama/Llama-3.1-8B. The slug names the Llama 3 licence and the base names Llama 3.1, and https://huggingface.co/NousResearch/Hermes-3-Llama-3.1-8B/raw/main/LICENSE returns HTTP 404, so nothing on an approved origin resolves the disagreement. An open-weight release requires a licence object, and guessing which licence the slug meant is exactly what this dataset refuses. Hermes-4-14B was carried instead, where the slug apache-2.0 and the base Qwen/Qwen3-14B agree.Blocked bylicenseSchema
Sources conflict, so no value changed
liquid-ai-lfm2-1-2b-context-windowThe LFM2-1.2B card states a context length of 32,768 tokens while the repository config.json sets max_position_embeddings to 128000. The conflict was recorded with the creator own stated figure taken as the reading, and was accepted unanimously. It describes a release that was not published.Blocked byliquid-ainous-research-hermes-4-14b-parameter-totalThe same Hub object reports safetensors.parameters.BF16 as 14768307200 and safetensors.total as 424960 for a model shipping six shards over 29.5 GB. The BF16 figure was taken as the reading. For four of the other five creators the two figures agree exactly, and in each of those four the total also equals the sum of the whole parameters block: AI Singapore 8030261248, Liquid AI 1170340608, Reka AI 20905482240 and Xiaomi 7833409536. The fifth, Aleph Alpha Pharia-1-LLM-7B-control, carries no safetensors block at all in its Hub API record, so it offers no comparison rather than an agreeing one; counting it among the agreeing records would have been the same absence-read-as-agreement error this entry documents elsewhere. For Hermes-4-14B the parameters block sums to the BF16 figure and not to the total, which is what identifies the total as the anomaly. Accepted unanimously; describes a release that was not published.Blocked bynous-research
What this run does not prove
- A withheld tranche is a real outcome of the policy rather than a failure of the sources. Everything above rests on pages fetched from origins this repository already approves, and the run reports six creators it could not source rather than six it stretched a citation to reach.
- The blocker is structural and no amount of better scouting reaches it. Every one of the 42 approved origins is either a generic host — github.com, huggingface.co, arxiv.org, opensource.org, storage.googleapis.com — or the own domain of a creator already in the catalogue. A creator that is not yet in the catalogue therefore has no approved domain of its own, so the announcement page that would state its release date is precisely the page the trust boundary excludes. Approving a new origin is a human act by design, so a refresh agent cannot close this gap from inside a run. What that blocks is the unattended refresh path, and only that. The gates skill names two remedies for extending the trust boundary — add the origin to a profile catalogue, or cite it in a change that takes the ordinary path — and the second requires no origin to be approved in advance, because a human reviews the citation on the way in. This entry should not be read as saying the catalogue cannot grow. It says a run like this one cannot grow it, and the reference below points at where the ordinary path is being taken instead.
- The statement about dates on these cards was wrong twice before it was right, and both attempts are recorded here rather than quietly replaced. The first said the cards carry no date-bearing line at all; the scan behind it matched month names and ISO forms only, so a dotted Current Version: `14.04.2025` fell straight through. The second said five of the six carry none; the widened scan behind that one also discarded every line longer than 300 characters, so a 480-character paragraph naming AIME-2024 and a 328-character changelog line beginning [2025.05.30] never reached the pattern at all. Two unrelated mechanisms, one a regex gap and one a length filter, produced the same false sentence, which is the evidence that widening the scan a third time would not have fixed it. What the sentence claimed was the defect, not the tool behind it. A scan supports a statement about what it matched and can never support a statement about what is absent, because the forms it does not encode are invisible to it by construction: abbreviated and non-English month names, quarters and half-years, relative dates, epoch seconds, dates in HTML attributes, JSON-LD or YAML front matter, and dates that appear only inside images. The corrected sentence is therefore carried by the categorisation of the dates actually found, each named with its card and line, and by no scan returning nothing. Both errors were caught by fetching the six cards rather than reading the diff, QA finding the first and review the second, and the AI Singapore card hashed byte-identical to the copy the run already held, sha256:2d4396287f5df1e4993d8a0e22b6abdf7b6262b4323476946eb8519be5f3ef7d, so neither error was a matter of the page having changed. The dates this entry names, and the fragments quoted around them, were each re-read from those stored bytes at the line cited — fourteen checks, all of which hold. One did not at first: the AI Singapore value is written inside backticks on the card and this entry had dropped them, corrected here so that a reader grepping the card for what this entry prints finds it.
- The issue framed part of this tranche as closing a non-transformer gap. That framing is factually wrong and was not carried into any claim. The LFM2-1.2B card describes the architecture in its own words as a hybrid model with multiplicative gates and short convolutions, ten double-gated short-range convolution blocks and six grouped query attention blocks, and the config lists full_attn_idxs [2,5,8,10,12,14] with 32 attention heads and 8 key-value heads. LFM2 contains attention. Writing the issue framing into the dataset would have published a false claim beside a citation that contradicts it.
- Nous Research was typed company rather than community. The community clause needs a source showing that contributors outside the entity own appointment chain decide its releases, and nothing consulted states that; the ordered procedure then falls through to the company fallback. This is a classification the panel judged rather than a fact any source states.
- The panel is three model instances reading the same pages, so 2-of-3 or 3-of-3 buys independence of reasoning and not independence of training. A source that is itself wrong could carry all three. The provenance rubric was also the only one of the three that rejected anything here, which is worth noticing rather than smoothing: consistency and editorial each accepted all 54, so the unanimous bar was carried entirely by one rubric.
- The editorial verdicts were rewritten once. The first pass returned a single sentence repeated across all 54 claims, which fails the requirement that a rationale state what decided it; the reviewer was asked for claim-specific rationales with its votes explicitly untouched, and it changed none. The rewritten rationales are the ones recorded, and a reader should know they were produced on a second pass. They are also not fully claim-specific: the 54 recorded editorial rationales are 28 distinct strings, nine of them reused - two covering six claims each, four covering four, one covering three and two covering two - where provenance and consistency each recorded 54 distinct ones. The account of the first pass is itself unverifiable from what survives, because the rewrite overwrote it, so both that a single rationale was repeated 54 times and that no vote changed are taken from the run record rather than from an artefact a later reader can inspect. That absence was checked across every directory this run wrote rather than only the one holding the verdict files: the sole artefact predating the rewrite is the packet the reviewers were given, which carries the claims and the dataset slice and records no vote or rationale at all. This is the one statement in this entry that no preserved byte can settle.
- This entry was written by the run rather than transcribed after it, which is why gate-scope reports the branch out of the ADR 0003 class. The entry records gate results for npm run validate and ci-preflight that were observed after the entry itself was written, so those two lines describe the tree as committed rather than the tree as it stood when the earlier gates ran.
- The account of what the panel accepted was wrong in two places and is corrected here rather than silently replaced. The withheld entry for the accepted sources and publishers stated three sources and three publishers and then named four of each, filing the Liquid AI record as a source and the Reka AI record as a publisher when the bundles have them the other way round; and the follow-up resting on it named Nous Research, Reka AI and Liquid AI as each having reached both a publisher and a configuration source, when only Nous Research did. The aggregate arithmetic was right throughout — 54 claims, 162 verdicts, 8 accepted and 46 rejected — because the error lay inside the eight, so no check that reconciled totals could see it. The first defect was visible as a disagreement between the summary and the detail; the second contradicted nothing in the entry and was found only by counting the accepted claims out of the six bundles again. Both were settled against the bundles rather than by making one field agree with another, which is the only method that could have settled the second.
- The evidence behind this entry is verifiable, and an earlier version of this caveat said the opposite. Every quote was machine-checked as a literal substring of the exact bytes whose sha256 the bundle records; the six bundles survive; and the fetched page bodies survive too, 56 of them across the six run directories this run wrote, 50 in their pages folders and 6 more re-fetched into the date re-check folder, covering all 31 of the 31 URLs cited as evidence, with no cited URL left without a stored body. By origin they are 38 from huggingface.co, 11 from github.com, 4 from arxiv.org and 3 from opensource.org; the 38th is the re-checked AI Singapore card, which is stored without a host prefix in its filename and so is easy to drop from a breakdown that reads filenames. The earlier version said bodies survived for 19 of 43 pages covering 15 of 31 cited URLs, that every preserved body came from huggingface.co, and that nothing this entry says about a github.com, arxiv.org or opensource.org page could be checked against bytes held here at all. That was false in every part. It was measured over one run directory while the run had been writing to six, and the 19 huggingface.co bodies it found were the true contents of that directory rather than of the evidence set. Three statements it listed as resting on the run record are in fact held in preserved bytes: SEA-Guard appears on the stored AI Singapore GitHub README at lines 948 and 949, and at line 911 as SEA-GUARD in capitals inside a table link, which is a different string and would not be found by a reader grepping for the spelling this entry uses; the March 2026 SEA-LION-Embedding suite is at line 947 of the same page, and MiMo-7B-RL-0530 appears on the stored Xiaomi GitHub README at lines 792 and 799, line 792 being the changelog line dated [2025.05.30]. So are two facts asserted elsewhere in this entry: the Hermes-4-14B conflict figures are in the stored Hub API body, which records safetensors.total as 424960 and safetensors.parameters.BF16 as 14768307200, and the Hermes-3-Llama-3.1-8B LICENSE 404 is stored as a 15-byte body reading Entry not found, the 404 response being itself the artefact. What remains true is the limit ADR 0005 records: gate-evidence validates the shape of a contentHash and a quote and does not, and cannot as built, verify that the digest is the hash of the page at the url, because the gate fetches nothing. The bodies kept here close that gap for a reader who holds them; they never closed it for the gate, which did not look at them.
- Eight statements in this entry were wrong when first written, counting the two attempts at the date claim separately, and all eight failed in the same way: a correct aggregate carrying a wrong decomposition. Two were absence claims about dates that the scans behind them could not support. One stated the right totals for what the panel accepted — three sources, three publishers and two conflict records — while filing the Liquid AI record as a source and the Reka AI record as a publisher, which the bundles have the other way round. One called 48 fetches 48 unique pages, when the run had in fact made 56 fetches over 48 pages, so the figure it was corrected to was itself too low. One said a single editorial rationale covered six claims when two of them do. One said the safetensors figures agree for the other five creators, when they agree for four and the fifth carries no safetensors block to compare at all, which is an absence read as an agreement. In every case the aggregate itself held: 54 claims, 162 verdicts and 8 accepted against 46 rejected are all correct as written, and every check that reconciled them confirmed them. That is precisely why the errors survived, because a total ties whichever way its parts are attributed, so a check that reconciles totals is structurally blind to a permutation inside one. Each was found by recounting against the artefacts rather than by reading the entry, and the accepted-claim inversion contradicted nothing in the entry at all, so no amount of internal comparison could have surfaced it. The fourth says the most about what the gates cover. It was present, in three places, in the tree the review gate read at commit 473dac1, and review passed that tree without examining pagesFetched at all. That was not carelessness: the same review verified the accepted-claim decomposition down to per-creator attribution out of the bundles, because that claim had been named for it, and it did not generalise from one wrong count to checking every count. Coverage followed the brief it was given. A fifth was then found by the coordinating session walking the run directories instead of trusting this entry, and it is the most self-implicating of the set. The thesis of this entry is that an absence claim must be a matched set - that saying no approved source states a release date obliges you to name every date found and say what each one is instead. Its own absence claim about its own evidence did the thing it warns against. It counted preserved bodies in the pages folder of one run directory, found all 19 of them to be huggingface.co, and stated universally that no github.com, arxiv.org or opensource.org body had ever been persisted, when 18 such bodies were sitting in four of the five sibling directories the same run had written that evening. The aggregate held again - 19 huggingface.co bodies really were in the folder that was looked at - and the decomposition was again what was wrong. The review gate confirmed that claim rather than catching it, reporting independently that all 19 were huggingface.co, because it inherited the directory scope instead of re-deriving it. That is the sharpest result this run produced about what agreement is worth: independence of reviewers does not give independence of scope, and two parties looking in the same wrong place agree perfectly. The correction was then widened twice more, once from one run directory to six and once again from the pages folders to the date re-check folder, and neither widening began in a disagreement about a number. Each came from asking what else the run had written rather than whether a figure was right, which is a different question and is the one that found every instance. The last of them was caught by the review gate, and that balances the account. Correcting the coverage caveat moved the preserved-body count from 19 to 56, but found.pagesFetched and the sentence explaining it kept the figures computed before the widening, so the entry asserted 48 fetches over 43 distinct pages in one place and 56 stored bodies in another, and named the first as correct as written. A stored body is the product of a fetch, so the two could not both stand. Review found that contradiction; neither this run nor the session coordinating it did. The lesson is narrower and more useful than checking the counts again: when a correction changes the size of a set, every number computed over that set has to be re-derived, including numbers in structured fields the correction did not edit. The prose and the field were two views of one measurement, and only one of them moved. Of the eight defects, two were caught by gates that had been pointed at them, one by the review gate looking where it had not been told to look, three by this run auditing its own artefacts and two by a reader who widened the measurement. The general form is that a document internal agreement is not evidence about the world, and where it is the only evidence available the honest report is unverified rather than consistent.
Follow-ups — proposed, not fixed
- A new creator cannot be added by an unattended refresh from the approved origin set alone, because releaseDate is required and only a creator own announcement states one. Three routes stay open and this run took none of them, because all three are outside what a refresh agent may do: approve each new creator origins in a profile catalogue as a reviewed human step before its refresh runs; add the creator through a change that takes the ordinary path, where a human reviews the citation and no origin has to be approved in advance; or change what the schema requires of a release whose date no approved source states.
- The six bundles and all 162 verdicts are reproducible artefacts of this run but live under a git-ignored directory, so they disappear with the machine. The fetched bytes are kept in full there: 56 body files across six run directories cover all 31 of the URLs cited as evidence, so every citation here can be checked against stored bytes while that directory exists and only by re-fetching a live page that may have changed once it does not. If the origin question above is answered, the research does not need repeating from scratch, but only if it is captured before that happens.
- Five of the six creators reached at least one unanimously accepted record: Nous Research a publisher and a configuration source, Liquid AI and AI Singapore a publisher only, Reka AI and Aleph Alpha a source only, and Xiaomi none. Nous Research has the least left to fetch and would be the cheapest to revisit first once a date-bearing origin is available. None of them is close to publishable in the meantime, because every release still needs a date that no approved origin states.
2026-08-29-34e751
Data refresh 2026-08-29
Scope requested: Every creator in organizations.json — ai2, ai21-labs, amazon, anthropic, alibaba-cloud, cohere, deepseek, google-deepmind, meta, microsoft, mistral-ai, moonshot-ai, nvidia, openai, tii, xai, zhipu-ai — under the pilot 2-of-3 policy.Published3 edits posted · 4 items withheldAn agent-run, source-backed refresh under ADR 0003. All 17 creators in organizations.json were scouted; 32 pages were fetched and hashed across 18 creator slugs, and only two creators produced claims at all — Anthropic with three and OpenAI with one. Fifteen blind reviewers cast fifteen verdicts across three rubrics, eight deterministic checks passed, and three edits landed across two dataset documents before GitHub merged PR #529 on a green web-ci and Pages deployed successfully. The principal finding is small and was sitting in plain sight: the anthropic-claude-4-8 family existed with no release record and therefore rendered empty, so this run added the missing Claude Opus 4.8 release and the docs page it is sourced from. No human reviewed any part of it. This entry was transcribed a day late — see the caveats.
- PreflightRanAnchor a6a3dfc, selected by merge-base with refs/remotes/origin/main and computed by the gate rather than supplied — requestedBase was null. The approved-origin catalogue parsed cleanly at that anchor: 147 dataset sources and 12 profile catalogues out of 13 profile files, yielding 33 approved origins. The one profile listed but not drawn on is tools/updater/profiles/generic/long-tail.json, which configures no origins. The anchor is therefore full rather than silently narrowed.
- ScoutRan32 pages fetched and hashed across 18 creator slugs. Every quote was extracted by anchor from fetched page text, so none is a recollection. No page budget was exhausted. ai.meta.com/blog/ returned HTTP 400 to every request tried, including with a browser User-Agent and a cookie jar, so Meta scouting fell back to huggingface.co/meta-llama, an approved origin; no Meta claim rested on the unreachable page and no Meta record changed. The 15 creators outside Anthropic and OpenAI produced coverage observations rather than claims.
- ReviewRanThree independent reviewers per claim, launched in parallel across different model families, none seeing the scout's reasoning, another reviewer's verdict, or the running tally. The chair did not vote. One claim was reviewed twice and this is disclosed rather than hidden: round 1's provenance reviewer rejected the Opus 4.8 release because no attached quote stated the modalities, which was a mechanical defect — the bundle generator's block extractor matched an earlier occurrence of "Context window 1M tokens", so the "Input → output Text and images → text" line was never attached. The rubric's own remedy for a missing quote is to attach it, so the anchor was fixed, intendedUse was tightened to attribute rather than assert, and three fresh reviewers re-reviewed. Round 2's provenance reviewer rejected again on different, judgment-based grounds and the claim was not revised a third time: repeated revision to chase a verdict is vote-shopping. Fifteen verdicts were cast in total; the twelve from the final panels are published verbatim in the pull request body, and round 1's three rationales were not preserved.
- GatesRangate-evidence and gate-source-approval ran per creator before anything was applied; gate-dataset, npm run validate and gate-scope ran afterwards. No gate was skipped, no threshold lowered, no --force, no --admin, no --base override, and no direct push to main. gate-scope reported changed: 2, empty: false, outOfClass: [], so the change sat inside the ADR 0003 qualifying class and was eligible to auto-merge.
- PublishRanPR #529 carried the full evidence trail and was merged by GitHub via --auto once web-ci was green, not by the agent. Merge commit 0794286.
- DeployRanPages deploy run 33250588183 on main succeeded, and https://abdeslam-menacere.github.io/ModelTree/models/claude-opus-4-8/ returned HTTP 200 serving the new record. No revert needed.
What was found
- Scouts
- 17
- Pages fetched and hashed
- 32
- Claims proposed
- 4
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims anthropic pilot 2-of-3 3 openai pilot 2-of-3 1 What those claims proposed to do Kind Count Effect Add 2 One release and one source proposed; both survived every stage and were applied. Unchanged 2 Re-verified against a primary source; no value changed. One had its verification date advanced and one had it withheld. Not covered
- Fifteen of the seventeen creators scouted produced no claim at all. That is an honest zero per creator rather than a skipped stage, but it also means their records carry this run's attention without carrying its verification date.
- ai.meta.com/blog/ returns HTTP 400 to this client and was not read; huggingface.co/meta-llama was used instead.
- Google, Microsoft, Amazon, Cohere and OpenAI audio models remain uncovered by any claim this run could assemble.
- Mistral Medium 3.5 was deliberately not claimed. Its release date is ambiguous between a product post dated 2026-05-22 and a version string reading "v 26.04", and an ambiguity is a finding rather than a coin to toss.
What was evaluated
- Reviewers
- 15
- Verdicts cast
- 15
- Accepted by panel
- 4
- Rejected by panel
- 0
Deterministic gates and required checks — 8 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence.mjsanthropic 0 PassPolicy pilot, threshold 2. Three claims, of which two were applicable. gate-evidence.mjsopenai 0 PassPolicy pilot, threshold 2. One claim, none applicable — an unchanged claim applies nothing. gate-source-approval.mjsanthropic 0 PassFour cited sources were inherited and one was proposed: anthropic-opus-4-8-docs, on https://platform.claude.com — an origin already approved at the anchor commit and already carrying inherited sources, so no new origin was introduced. No source was refused. gate-source-approval.mjsopenai 0 PassThe one cited source, openai-gpt-5-6-sol-docs, was inherited from an already approved origin. No source was refused. gate-dataset.mjsweb/src/data 0 Pass148 sources, 26 publishers, 17 organizations, 37 families, 64 releases, 7 model-fit statements. npm run validateweb/ 0 Pass1613 tests executed across 71 files; astro check reported 0 errors. gate-scope.mjsbranch vs merge-base 0 Passchanged: 2, empty: false, outOfClass: []. Anchored on the merge-base the gate computed rather than one supplied to it. web-ciPR #529 — PassSUCCESS. GitHub performed the merge via --auto on this check going green; the agent did not merge. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
anthropic-claude-opus-4-8-release-addProvenanceaccessType "proprietary-hosted" is not forced by the quotes, because hosted availability does not exclude downloadable weights; and predecessorIds: [] under-states an announcement that calls Opus 4.7 "its predecessor". Outvoted 2-1. The claim was applied and both points are recorded here and in the pull request body as unresolved dissent rather than smoothed over.anthropic-claude-mythos-5-limited-release-unchangedProvenanceThe quotes do not uniquely force status "preview" over "current", and do not re-state the recorded intendedUse. Outvoted 2-1, and materially right in the event: the claim reached its threshold and still applied nothing, because the verification date it would have advanced was withheld on separate grounds.
Posted 3 edits
3 edits across 2 documents, a net change of 2 records.
Dataset documents this run changed Document Before After What changed releases.json63 64 One release added: anthropic-claude-opus-4-8. openai-gpt-5-6-sol had its verifiedAt advanced to 2026-08-29. sources.json147 148 One source added: anthropic-opus-4-8-docs — the platform.claude.com model overview page the Opus 4.8 release is drawn from, accepted 3-of-3 on an origin already approved at the anchor commit. lastCheckedDate moved to 2026-08-29 on three of the four inherited sources re-read this run; anthropic-fable-5-mythos-5-docs was held at 2026-08-26 with the Mythos 5 claim it belongs to. Sources are listed here rather than under the posted records below, which carry links and so hold only the collections the site routes. Each document links to the file as this run left it, not as it stands today.
Records added
- Claude Opus 4.8passportreleases
anthropic-claude-opus-4-8New release — released 2026-05-28, status legacy, 1M context and 128K max output, text and image in, text out, API alias claude-opus-4-8. Fills the anthropic-claude-4-8 family, which until now had no release and rendered empty. Carries the recorded provenance dissent on accessType and on its empty lineage arrays.
Not posted 4 items
Verification date deliberately held back
releases/anthropic-claude-mythos-5.verifiedAtHeld at 2026-08-15; nothing was applied for this claim despite it reaching its threshold. validate.ts forbids guidance dated before the evidence beneath it, and fit-claude-mythos-5-limited-availability, verified 2026-08-15, rests on this release's status and intendedUse. Advancing the release would have silently invalidated that fit statement — validation failed exactly this way when it was first attempted. Re-dating guidance this run gathered no evidence for would have been the larger claim, so the date was withheld instead. This also matches how gate-evidence treats unchanged claims, which apply nothing.Blocked byfit-claude-mythos-5-limited-availabilityvalidate.ts
Sources conflict, so no value changed
mistral-medium-3-5Not claimed at all. Its release date is ambiguous between a product post dated 2026-05-22 and a version string reading "v 26.04". Neither reading is forced by the other, so the run recorded the disagreement rather than picking a side. The dataset still holds nothing on this model, which is the cost of leaving the ambiguity explicit.
Out of the run’s reach
meta-announcement-indexhttps://ai.meta.com/blog/ returns HTTP 400 to every request tried, including with a browser User-Agent and a cookie jar. Meta scouting fell back to huggingface.co/meta-llama, an approved origin. No Meta claim rested on the unreachable page and no Meta record changed, so absence of a Meta finding here is not evidence that there was nothing to find.Blocked byhttps://ai.meta.com/blog/coverage-google-microsoft-amazon-cohere-openai-audioGoogle, Microsoft, Amazon, Cohere and OpenAI audio models were scouted and produced coverage observations rather than claims. The observations are carried as follow-ups below rather than being written into the dataset as inferences.
What this run does not prove
- No human reviewed the pull request. The evidence trail, the deterministic gates and the required web-ci check are the whole of the oversight.
- A green run proves the dataset is internally coherent and that every applied claim carried a quote from an approved origin. It does not prove the sources are right.
- Two creators out of seventeen produced claims. The other fifteen were read and yielded nothing claimable, so this run moved the dataset barely at all and its silence about them is not a clean bill of health.
- Round 1's three review rationales for the Opus 4.8 release claim were not preserved verbatim in the run artifacts, so the audit trail for that claim starts at round 2. That gap is follow-up 6 and it is the reason the reviewer count here exceeds the twelve rationales published.
- editsApplied counts the three claims that reached the dataset — two adds and one advanced verification date — following the convention of the earlier entries here. The merge commit also moves lastCheckedDate on three sources that rode along with those claims, which the diff shows and this count does not.
- This entry was written on 2026-08-30, a day after the run, because the run could not record itself: refresh-runs.json sits outside the ADR 0003 qualifying class that gate-scope enforces, so including it in PR #529 would have failed that gate and forfeited auto-merge. The transcription needs its own human-merged pull request, and for this run that pull request was initially never opened. The absence was noticed only when a reader asked why the page showed no 2026-08-29 line.
Follow-ups — proposed, not fixed
- Re-verify fit-claude-mythos-5-limited-availability so the Mythos 5 release verifiedAt can move. It is blocking a routine re-verification today.
- Populate predecessorIds and successorIds across the Claude Opus line. Opus 4.6, 4.7 and now 4.8 all carry empty lineage arrays while the announcements state the relationships explicitly. Fix the line as a whole rather than one record.
- Resolve the Mistral Medium 3.5 release-date ambiguity — decide which of the two published dates the dataset should record, or record the ambiguity explicitly.
- Close coverage gaps for Google, Microsoft, Amazon, Cohere, and OpenAI audio models.
- Find a reachable Meta announcement index to replace ai.meta.com/blog.
- Persist superseded review rounds verbatim in the run artifacts, so a corrected claim's first-round rationales survive for audit.
- Nothing notices when a published run never gets transcribed into this log. The publishing skill set opens the dataset pull request and files the summary issue, and the log entry is a separate human-merged change that no check requires; this run's line was missing for a day and only a reader's question surfaced it.
2026-08-28-cff539
Long-tail refresh 2026-08-28 — Cohere
Scope requested: One long-tail creator, Cohere, under the unanimous 3-of-3 policy, restricted to the two origins PR #445 approved for it: docs.cohere.com under /docs/ and cohere.com under /blog.Ran, changed nothing0 edits posted · 6 items withheldA second attempt at Cohere, run after PR #459 amended the provenance rubric to treat vocabulary mapping as a recording step rather than an inference. The amendment worked and is not why this run publishes nothing: the provenance reviewer stated in terms that "mapping Live to current would be a permitted recording step at the same entity level", conceding the exact objection that sank run 2026-08-27-4f1c9e. Three different blockers stopped it instead. No release claim could be assembled at all — familySchema.firstReleaseDate is a day-precision isoDate and no approved origin states one for any Cohere family, while the one model with a day-precision date, Command A+, is released under two deployment routes and so requires a license object whose osiApproved field no approved origin states. The Command A+ family claim that was proposed was rejected 0-of-3, all three rubrics independently holding that Cohere's own words "the last model in the Command A family" make it a model inside a family rather than a family. The organization and publisher claims were rejected 2-of-3 by provenance on the same ground as last time — no quote states the schema classification "company" — which PR #459 did not address. Two source claims were accepted unanimously and still were not applied: with no record citing them they would be unreferenced sources, which validate.test.ts forbids. The dataset is unchanged; this entry is the only edit.
- PreflightRandrydock is not on PATH, so the manual posture applied: implement, commit, report, and stop at the review gate rather than landing anything. Clean tree rebased onto origin/main at 1512da2, then rebased again mid-run onto 73f92f3 after three pull requests landed — #462 correcting the featured criterion, #463 adding the A-Z provider directory and #464 the lineage explorer. The second rebase was clean and releases.json did not conflict. Every gate below was recomputed against the final anchor 73f92f3 rather than carried forward, because a gate result binds to the commit it was computed against.
- ScoutRan26 pages fetched over real HTTP from the two approved origins only, all HTTP 200, with the exact bytes and their sha256 recorded in a fetch manifest before any quote was written. Every quote was then extracted programmatically from those hashed bytes by a builder that throws if the quote is absent, ambiguous or under the 24-character minimum; it caught two real faults during construction — a search-offset bug, and an identifier that appears in both a model table and a serving-platform table and so could not be quoted unambiguously. docs.cohere.com/changelog/, /llms.txt and /v2/docs/, and cohere.com/research/, were deliberately not fetched: they fall outside the allowed_paths in the approved origins catalogue.
- ReviewRanOne blind panel, three rubrics, run as three parallel sub-agents on three different model families so no rubric shared a training lineage with another. Each reviewer saw only the stripped bundle — statement, values, evidence, dataset slice and its own rubric — and never the scout's reasoning, another reviewer's verdict, or the running tally. 18 verdicts cast, every one carrying a rationale, all published verbatim in the pull request body. No reviewer was re-run and no threshold was adjusted.
- GatesRangate-evidence and gate-source-approval ran before anything would have been applied, anchored on the committed dataset at merge-base 73f92f3 computed by the gate rather than supplied. gate-dataset, npm run validate and gate-scope ran afterwards. gate-evidence was first invoked wrongly and exited 2; exit 2 is "the gate could not run", so it was treated as a failure and re-run correctly rather than read as a pass.
- PublishRanAn ordinary pull request carrying this log entry and no dataset change. Auto-merge was neither requested nor eligible: refresh-runs.json is outside the gate-scope qualifying class, so ADR 0003 does not cover this change and a human merges it.
- DeployNot applicableNo dataset change to deploy. The rendered tree is byte-identical to the one already on main; only the /refresh/ page gains this entry.
What was found
- Scouts
- 1
- Pages fetched and hashed
- 26
- Claims proposed
- 6
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims cohere long-tail 3-of-3 6 What those claims proposed to do Kind Count Effect Add 6 Would have added three sources, one publisher, one organization and one family. No release claim was proposed, because no release could be assembled from the approved origins without guessing a required field. None was applied. Not covered
- No release claim for any Cohere model. familySchema.firstReleaseDate is a day-precision isoDate with no datePrecision companion, and no approved origin states a first release date at that precision for the Command, Aya, Embed, Rerank, Parse or Transcribe families. The best available are month-precision prose — "delivered in August 2024", "delivered in December 2024" — and the snapshot dates encoded in identifiers such as command-a-03-2025.
- command-a-plus-05-2026 has a quotable day-precision release date and was still not claimed: the model card states both "Cohere Deployment" and an Apache 2.0 open-source deployment, which is accessType "both", and releaseSchema then requires a license object whose osiApproved is a non-optional boolean no approved origin states. This is open issue #461.
- rerank-v3.5 carries the only other day-precision date on an approved origin, but the Rerank table has no Status column, no approved origin states its accessType, and its single "Modalities" column does not separate input from output. Three required fields unstated.
- huggingface.co/CohereLabs states the Apache 2.0 licence and would close the osiApproved gap. It is explicitly not an approved origin for this creator, and a run may never approve one, so it was not fetched and not cited.
- Blog publication timestamps on cohere.com/blog were deliberately not used as release dates. The amended provenance rubric states that a page’s publication timestamp is not a statement of a release date, and this run held to that even though it left the run with no usable dates.
- Bedrock, SageMaker, Azure AI Foundry and Oracle OCI identifiers appear throughout the Cohere docs and were never used to fill a release’s accessType. They state what a serving platform offers, not how the creator releases the model.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 18
- Accepted by panel
- 2
- Rejected by panel
- 4
Deterministic gates and required checks — 3 of 5 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence1 bundle, 6 claims 1 FailFailed on exactly the four claims the panel put below the unanimous 3-of-3 long-tail threshold: co6-source-command-a-plus-card-add, co6-publisher-cohere-add, co6-organization-cohere-add and co6-family-command-a-plus-add. It reported no evidence-format failure of any kind — every quote was present, verbatim against the hashed bytes, and over the 24-character minimum. The failure is the review threshold, not the evidence. Nothing was applied, so nothing shipped past it. gate-source-approval1 bundle, 12 citations 0 PassAnchored on the committed dataset at merge-base 73f92f3 with refs/remotes/origin/main, computed by the gate rather than supplied, so the run could not approve its own source. 24 already-trusted origins; all 12 citations landed on the two origins PR #445 approved for Cohere. No new host was admitted and none was proposed. gate-datasetworking tree 0 PassThe dataset is unchanged by this run and remains internally coherent: 109 sources, 15 publishers, 7 organizations, 26 families, 49 releases. No orphaned source, no dangling reference, no release predating its family. npm run validateweb/ 0 PassFull test suite and Astro/TypeScript diagnostics green — 985 tests across 37 files, 0 errors — including the refresh-log schema test this new entry must satisfy, and validate.test.ts, whose unreferenced-source rule is the reason the two accepted sources were not applied. gate-scopebranch vs merge-base 73f92f3 1 FailReports web/src/data/refresh-runs.json as outOfClass, which is correct and expected rather than a problem to work around: the ADR 0003 qualifying class is the eleven dataset documents raw.ts composes, and the refresh log is deliberately not one of them, so any run that logs itself leaves the class by construction. The gate was not modified, the file was not added to ALLOWED_PATHS, and the log entry was not dropped to stay in class. The consequence is simply that this change is outside ADR 0003 and merges the ordinary way, with a human. Posted 0 edits
Nothing reached the dataset. No branch, no commit, no pull request.
Not posted 6 items
Rejected by the review panel
cohere-command-a-plus-familyRejected 0-of-3 — the only claim in either Cohere run to be refused by all three rubrics independently, and the clearest finding here. Cohere’s own documentation, quoted in the claim, calls Command A+ "the last model in the Command A family". Provenance held that the Live status and release date are quoted about the dated command-a-plus-05-2026 release and so add unsupported scope when applied to a family. Consistency held that filing a model as a family breaks lineage coherence and contradicts the claim’s own description. Editorial held that the profile’s naming rule makes the dated identifier the release and the dateless form its alias, so elevating one model to a family nested inside the family Cohere actually names mis-levels a release. This is a modelling defect the run got wrong, not a sourcing gap, and it is recorded rather than patched: correcting it would need a second panel on a claim whose release dependency had already failed.Blocked byco6-family-command-a-plus-addcohere-organization-and-publisherBoth accepted by consistency and editorial and both rejected by provenance, 2-of-3, short of the unanimous long-tail bar. The objection is precisely the one that sank the same claim in run 2026-08-27-4f1c9e: no quote states the schema classification "company", and "Name of the model provider: Cohere Inc." supplies a provider name rather than a classification. PR #459 legitimised selecting a vocabulary member that a quoted term denotes; it did not create a quoted term here, and the rubric’s "the source states nothing on the point" rejection still applies. That is the amendment working as written, not failing.Blocked byco6-organization-cohere-addco6-publisher-cohere-addcohere-command-a-plus-model-card-sourceAccepted by consistency and editorial, rejected by provenance 2-of-3. The page heading "Cohere’s Command A+ Model" is verbatim in the hashed bytes, but provenance held it identifies a page about the model without stating that the page is a structured model card, so the source type "model-card" rested on an unsupported classification. Editorial reached the opposite conclusion on the same page, noting it carries provider, release date, modalities and model size under explicit headings. The disagreement is recorded rather than resolved by the chair, who does not vote.Blocked byco6-source-command-a-plus-card-add
Accepted by the panel, then dropped
cohere-accepted-sourcesThe two documentation sources — the Models Overview page and The Cohere Platform page — were the only claims to clear the unanimous bar, accepted 3-of-3 by all three rubrics. They were still not applied. Every claim that would have cited them was rejected, so applying them would add sources that no record references, and check-bundle-pairing refuses exactly this: "an unreferenced source is dead provenance", a rule validate.test.ts enforces in the suite. Applying two records that reference nothing to report a non-empty diff would be the pipeline working against its own purpose.Blocked bycohere-organization-and-publishercohere-command-a-plus-family
Verification date deliberately held back
cohere-command-a-plus-05-2026The strongest release candidate Cohere publishes, and it was never put to the panel because it cannot be assembled honestly. Its model card states both "Cohere Deployment" and "Open Source Deployment: Command A+ is available under an Apache 2.0 License on Hugging Face", which is accessType "both"; releaseSchema.superRefine then requires a license object, and licenseSchema requires osiApproved as a non-optional boolean. No approved Cohere origin states OSI approval, and schema.ts states in terms that downloadable weights and OSI approval are separate claims where neither implies the other, so it cannot be derived from the Apache 2.0 quote. huggingface.co/CohereLabs would settle it and is explicitly not an approved origin; a run may never approve one. Open issue #461. Withheld rather than filled — a required field is not a licence to guess.Blocked byweb/src/data/schema.tstools/updater/profiles/origins/cohere.jsoncohere-remaining-model-table-rowsCommand A, Command R, Command R7B, Command R+, Command A Translate, Command A Reasoning, Command A Vision, Aya Expanse, Aya Vision, Tiny Aya, Embed, Parse, Cohere Transcribe, rerank-v3.5 and rerank-v4.0 — the full breadth this run was asked to claim. Each needs a family, and familySchema.firstReleaseDate is a day-precision isoDate with no datePrecision companion even though partialDate exists in the same file and is used for evaluationDate, release-event dates and usage windows. Only two day-precision Cohere release dates exist anywhere on an approved origin, and both belong to explicitly non-first family members, so every family here would need a day the source never gave. rerank-v3.5 additionally lacks a stated status and accessType and its "Modalities" column does not separate input from output. Breadth was the goal and the schema, not the sources, is what refused it.Blocked byweb/src/data/schema.ts
What this run does not prove
- Nothing was applied to the dataset. The Others branch is unchanged and still renders no creators; this entry is the only edit in the pull request.
- PR #459 did what it set out to do, and this run is evidence for that rather than against it. The provenance reviewer wrote that "mapping Live to current would be a permitted recording step at the same entity level" — the precise objection that ended the previous Cohere attempt, now conceded. It rejected on entity-level scope instead, which is a different and untouched part of the rubric. Reading this run as PR #459 having failed would be the wrong lesson.
- The two remaining blockers are schema shape, not evidence and not review. A day-precision firstReleaseDate and a required osiApproved are both fields the schema demands and the approved sources do not state, and no amount of better scouting reaches them. Until one of the follow-ups below lands, Cohere is unpublishable from its currently approved origins however good the evidence is.
- The 0-of-3 family rejection is a fault in this run’s own modelling, and it is reported as one. Command A+ is a model in the Command A family and filing it as a family was wrong; all three rubrics caught it independently, which is the panel doing exactly what it exists to do.
- Since PR #463, an organization with at least one family but no releases now renders in the /providers A-Z directory rather than nowhere, so the reason the previous run gave for dropping accepted organizations no longer holds in full. It did not change this run’s outcome — the organization and family claims were rejected on their own merits — but the constraint recorded in run 2026-08-27-4f1c9e is now out of date.
- A withheld creator is a real outcome of the policy, not a failure of the sources. Everything above rests on facts Cohere publishes plainly on origins this repository already approved.
Follow-ups — proposed, not fixed
- Make familySchema.firstReleaseDate accept partialDate, the same YYYY(-MM(-DD)) shape already used in this file for evaluationDate, release-event dates and usage windows, and pair it with a datePrecision companion as releases already have. It is the single change that would let this run publish: Cohere states month precision for several families and day precision for almost none, and a day-precision-only field forces a fabricated -01 day or nothing at all.
- Resolve issue #461. licenseSchema requires osiApproved whenever a license object is present, and releaseSchema requires a license object whenever accessType is "both" or "open-weight", so a well-sourced Apache 2.0 release is unrecordable unless an origin stating OSI approval is approved. Either make osiApproved optional, or approve huggingface.co/CohereLabs for this creator. Today the only honest option is to withhold the release entirely.
- Update the constraint recorded in run 2026-08-27-4f1c9e and in issue #441: after PR #463, model-tree.ts is no longer the only renderer, and provider-directory.ts places an organization with a family and no releases into the A-Z directory. An organization-and-family record is no longer invisible, so "renders nowhere" is no longer a sufficient reason on its own to drop accepted claims.
- Decide whether a source’s type field — "model-card" against "official-docs" — is a claim needing its own verbatim support or a cataloguing decision. Provenance and editorial split on exactly this for the Command A+ page, and the same split will recur on every model card the project ever records.
- Consider whether the organization type "company" needs a quote at all. It is the second Cohere run stopped partly by the absence of a source sentence classifying a company as a company, and no approved origin is ever likely to contain one. If the field is a cataloguing decision rather than a claim about the world, the rubric should say so; if it is a claim, organizations may simply be unrecordable for most creators.
- docs.cohere.com/changelog/ and docs.cohere.com/v2/docs/ sit outside the allowed_paths in the approved origins catalogue and were not fetched. The changelog is the one place a creator normally states dated releases in prose, so approving that path is likely worth more to this creator than any further scouting of /docs/.
2026-08-27-ab1644
Data refresh 2026-08-27
Scope requested: Every creator in organizations.json, plus the long-tail profile sweepPublished35 edits posted · 56 items withheldAn agent-run, source-backed refresh under ADR 0003. Five scouts fetched and hashed 130 primary pages and proposed 104 claims; twelve blind reviewers cast 312 verdicts across three rubrics; six deterministic gates passed; 35 edits landed across four dataset documents and GitHub Pages deployed successfully. 51 claims were dropped, most of them to one scout-side fault that proposed new sources without the paired citation edit. No human reviewed any part of it.
- PreflightRanClean tree, gh authenticated, no unmerged pull request from a previous refresh.
- ScoutRanFive bundles, one per creator profile plus a long-tail sweep. 130 page snapshots hashed. openai.com/news/ returned 403 and the RSS feed was used instead; ai.meta.com/blog/ returned 400; docs.anthropic.com now redirects to platform.claude.com.
- ReviewRanTwelve reviewers, three rubrics over four bundles. The long-tail bundle had no claims to review. No threshold was lowered and no reviewer was re-run.
- GatesRangate-evidence, gate-source-approval and check-bundle-pairing before applying anything; gate-dataset, npm run validate, and gate-scope after. The coupled drop set was iterated to a fixpoint over 3 passes.
- PublishRanPR #417 opened with the full evidence trail across the body and two comments, and merged by GitHub on a green web-ci.
- DeployRanPages deployed main @ 547691a successfully; the live site serves the new entities. No revert needed.
What was found
- Scouts
- 5
- Pages fetched and hashed
- 130
- Claims proposed
- 104
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims openai pilot 2-of-3 26 anthropic pilot 2-of-3 24 google-deepmind pilot 2-of-3 22 meta pilot 2-of-3 32 long-tail sweep long-tail 3-of-3 0 What those claims proposed to do Kind Count Effect Add 74 Proposed new records; 30 survived every stage. Change 8 Field corrections; 5 survived every stage. Unchanged 20 Re-verified against a primary source; no value changed. 11 records had verifiedAt moved forward. Conflict 2 Primary sources disagree; both sides recorded, no value changed. Neither reached the dataset. Not covered
- Any creator outside the four pilot profiles and the long-tail catalogue.
- openai.com/news/ returns 403 to this client; the RSS feed was read instead, so anything on the news pages but absent from the feed was not seen.
- ai.meta.com/blog/ returns 400 to this client and was not read.
- The OpenAI page budget was exhausted at 40 pages, so its catalogue was not read to the end.
- The long-tail sweep returned zero claims. That is an honest zero, not a skipped stage — it was gated like every other bundle.
What was evaluated
- Reviewers
- 12
- Verdicts cast
- 312
- Accepted by panel
- 83
- Rejected by panel
- 21
Deterministic gates and required checks — 8 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence5 bundles 0 PassThree rubrics present per claim, no duplicate votes, no empty rationale, every claim at its profile threshold. gate-source-approval5 bundles 0 Pass76 citations against anchor 052fb0d, selected by merge-base with refs/remotes/origin/main. requestedBase was null, so the run could not approve its own committed source. 16 approved origins, anchored on 61 sources already in the dataset and 5 reviewed profile catalogues. No new host admitted, and no source added to make a citation resolve. check-bundle-pairingnot required5 bundles 0 PassLanded on main as #413 mid-run and was run from main copy rather than added to this branch, since adding it would be an out-of-class change gate-scope would rightly refuse. It named the same unpaired sources this run had already found independently. gate-datasetworking tree 0 PassThe dataset is internally coherent after the edits were applied: 80 sources, 19 families, 35 releases. gate-scopebranch vs anchor 0 Passchanged: 4, outOfClass: [], empty: false. Exit 0 alone is ambiguous; the publish precondition is exit 0 and changed > 0. npm run validatebaseline and after 0 Pass579 tests across 24 files, astro check 0 errors. web-ciPR #417 — PassGreen. This is the check branch protection requires before a merge. skills-cinot requiredPR #417 — PassGreen, unlike the 2026-08-25 run where it was red and non-blocking. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
openai-gpt-5-1-family-addProvenanceThe RSS quote states GPT-5.1 was introduced for developers on 2025-11-13 and the docs quote states coding/agentic use and token limits, but no quote states the proposed family status "legacy" or the multimodal category. The proposed record therefore states more than the evidence quotes do.openai-gpt-5-2-family-addProvenanceThe RSS quote states "Introducing GPT-5.2" with a 2025-12-11 pubDate, and the docs quote calls it a previous frontier model. No quote states the proposed family status "legacy" as a status value, so the record goes beyond the quoted evidence.openai-gpt-5-3-codex-family-addProvenanceThe RSS quote gives the GPT-5.3-Codex introduction and date, and the docs quote says it is optimized for agentic coding. No quote states the proposed family status "current", so the proposed record adds an unstated field.openai-gpt-image-family-addProvenanceThe dated announcement quote names "ChatGPT Images 2.0", not GPT-Image or GPT-Image-2. The GPT Image 2 docs quote states image generation/editing but not the 2026-04-21 introduction date for that entity.anthropic-claude-4-6-familyProvenanceThe quotes name Claude Opus 4.6 and Claude Sonnet 4.6 and describe their capabilities. They do not state a "Claude 4.6" generation family, the 2026-02-05 firstReleaseDate for that family, or its status.anthropic-claude-4-7-familyProvenanceThe quote states Opus 4.7 improves on Opus 4.6 in advanced software engineering. It does not state a broader "Claude 4.7" family, the proposed firstReleaseDate, or the status field.anthropic-claude-4-8-familyProvenanceThe quote states Anthropic is upgrading Claude Opus to Claude Opus 4.8. It does not state a broader "Claude 4.8" generation family, its firstReleaseDate, or its status.meta-llama-4-maverick-original-sourceProvenanceThe quote says the repository contains original checkpoints for Meta's llama-stack codebase. It does not name meta-llama/Llama-4-Maverick-17B-128E-Instruct-Original, so the source title is not stated by the quote alone.meta-llama-3-2-1b-releaseProvenanceThe quotes state Llama 3.2 1B details, 1.23B parameters, 128k context, release date, and the Llama 3.2 license title. No quote states the proposed "legacy" status or the downloadable/OSI license fields.meta-llama-3-2-3b-releaseProvenanceThe quotes state Llama 3.2 3B details, 3.21B parameters, release date, and the Llama 3.2 license title. No quote states the proposed 128,000-token context window for this 3B claim or the "legacy" status.meta-llama-3-2-11b-vision-releaseProvenanceThe quotes state Llama 3.2 Vision 11B details, 10.6B parameters, 128k context, release date, and the license title. No quote states the proposed "legacy" status or the downloadable/OSI license fields.meta-llama-3-2-90b-vision-releaseProvenanceThe quotes state Llama 3.2 Vision 90B details, 88.8B parameters, 128k context, release date, and the license title. No quote states the proposed "legacy" status or the downloadable/OSI license fields.meta-llama-4-maverick-api-aliasesProvenanceThe first quote names meta-llama/Llama-4-Maverick-17B-128E, but the second only says "original checkpoints" without naming the exact Instruct-Original repository. Read alone, the quotes do not state both proposed aliases.meta-llama-4-maverick-source-idsProvenanceThe first quote identifies the base Maverick repository, but the second only says "original checkpoints" without naming the exact original-checkpoint model card. The evidence therefore does not state both newly proposed source IDs.
Posted 35 edits
35 edits across 4 documents, a net change of 30 records.
Dataset documents this run changed Document Before After What changed sources.json61 80 19 sources added. families.json12 19 Seven families added. releases.json31 35 Four releases added. organizations.json4 4 Verification dates moved forward; no record added or removed. Each document links to the file as this run left it, not as it stands today.
Records added
- GPT-5.1in the treefamilies
openai-gpt-5-1New OpenAI family; carries a recorded provenance dissent. - GPT-5.2in the treefamilies
openai-gpt-5-2New OpenAI family; carries a recorded provenance dissent. - GPT-5.3-Codexin the treefamilies
openai-gpt-5-3-codexNew OpenAI family; carries a recorded provenance dissent. - GPT-Imagein the treefamilies
openai-gpt-imageNew OpenAI family; carries a recorded provenance dissent. - Claude 4.6in the treefamilies
anthropic-claude-4-6New Anthropic generation family; carries a recorded provenance dissent. - Claude 4.7in the treefamilies
anthropic-claude-4-7New Anthropic generation family; carries a recorded provenance dissent. - Claude 4.8in the treefamilies
anthropic-claude-4-8New Anthropic generation family; carries a recorded provenance dissent. - Llama 3.2 1Bpassportreleases
meta-llama-3-2-1bNew release; carries a recorded provenance dissent on its status and license fields. - Llama 3.2 3Bpassportreleases
meta-llama-3-2-3bNew release; carries a recorded provenance dissent on its context window and status. - Llama 3.2 11B Visionpassportreleases
meta-llama-3-2-11b-visionNew release; carries a recorded provenance dissent on its status and license fields. - Llama 3.2 90B Visionpassportreleases
meta-llama-3-2-90b-visionNew release; carries a recorded provenance dissent on its status and license fields.
Not posted 56 items
Rejected by the review panel
openai-gpt-5-1-release-addOpenAI's API docs list GPT-5.1 as an API model for coding and agentic tasks. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)openai-gpt-5-2-release-addOpenAI's API docs list GPT-5.2 as an API model for professional work. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)openai-gpt-5-3-codex-release-addOpenAI's API docs list GPT-5.3-Codex as an agentic coding model. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)openai-gpt-5-4-mini-release-addOpenAI's API docs list GPT-5.4 Mini as a faster GPT-5.4 variant for high-volume workloads. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)openai-gpt-5-4-nano-release-addOpenAI's API docs list GPT-5.4 nano as a GPT-5.4-class model for simple high-volume tasks. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)openai-gpt-image-2-release-addOpenAI's API docs list GPT-Image-2 as an image-generation model. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-sonnet-4-5-releaseClaude Sonnet 4.5 should be added as an Anthropic Claude 4.5 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-opus-4-5-releaseClaude Opus 4.5 should be added as an Anthropic Claude 4.5 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-opus-4-6-releaseClaude Opus 4.6 should be added as an Anthropic Claude 4.6 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-sonnet-4-6-releaseClaude Sonnet 4.6 should be added as an Anthropic Claude 4.6 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-opus-4-7-releaseClaude Opus 4.7 should be added as an Anthropic Claude 4.7 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-opus-4-8-releaseClaude Opus 4.8 should be added as an Anthropic Claude 4.8 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-2-5-flash-lite-add-releaseGemini 2.5 Flash-Lite should be added as a current Gemini 2.5 release dated July 22, 2025. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-flash-preview-add-releaseGemini 3 Flash Preview should be added as a preview Gemini 3 release dated December 17, 2025. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-1-flash-image-add-releaseGemini 3.1 Flash Image should be added as a current Gemini 3 release dated May 28, 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-pro-image-add-releaseGemini 3 Pro Image should be added as a current Gemini 3 release dated May 28, 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-2-5-flash-image-add-releaseGemini 2.5 Flash Image should be added as a current Gemini 2.5 release dated October 2, 2025. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-5-transcribe-add-releaseGemini 3.5 Transcribe should be added as a current Gemini 3 release dated to August 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-5-live-translate-preview-add-releaseGemini 3.5 Live Translate Preview should be added as a preview Gemini 3 release dated to June 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-1-flash-live-preview-add-releaseGemini 3.1 Flash Live Preview should be added as a preview Gemini 3 release dated March 11, 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-1-flash-tts-preview-add-releaseGemini 3.1 Flash TTS Preview should be added as a preview Gemini 3 release dated April 13, 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)
Accepted by the panel, then dropped
source-openai-api-docsModels | OpenAI API should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-openai-gpt-5-4-mini-docsGPT-5.4 Mini model documentation should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-openai-gpt-5-4-nano-docsGPT-5.4 nano model documentation should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-openai-daybreak-red-docsDaybreak Red model documentation should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-openai-daybreak-blue-docsDaybreak Blue model documentation should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-openai-gpt-realtime-mini-docsGPT-Realtime Mini model documentation should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsanthropic-fable-5-docs-sourceClaude Fable 5 should be added as an official Anthropic source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsanthropic-opus-5-docs-sourceClaude Opus 5 should be added as an official Anthropic source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsanthropic-haiku-4-5-docs-sourceClaude Haiku 4.5 should be added as an official Anthropic source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsanthropic-fable-5-biology-safeguards-sourceImproving Fable 5's biology safeguards should be added as an official Anthropic source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsanthropic-opus-4-5-announcement-sourceIntroducing Claude Opus 4.5 should be added as an official Anthropic source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-2-5-flash-lite-docsAdd Gemini 2.5 Flash-Lite as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-flash-preview-docsAdd Gemini 3 Flash preview as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-1-flash-image-docsAdd Gemini 3.1 Flash image as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-pro-image-docsAdd Gemini 3 Pro image as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-2-5-flash-image-docsAdd Gemini 2.5 Flash image (Nano Banana) as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-5-transcribe-docsAdd Gemini 3.5 Transcribe as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-5-live-translate-docsAdd Gemini 3.5 Live translate as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-1-flash-live-docsAdd Gemini 3.1 Flash live preview as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-1-flash-tts-docsAdd Gemini 3.1 Flash TTS preview as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsmeta-llama-guard-4-sourceThe source catalog should include meta-llama/Llama-Guard-4-12B. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsmeta-llama-prompt-guard-2-86m-sourceThe source catalog should include meta-llama/Llama-Prompt-Guard-2-86M. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsmeta-llama-prompt-guard-2-22m-sourceThe source catalog should include meta-llama/Llama-Prompt-Guard-2-22M. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.ts
Verification date deliberately held back
releases/openai-gpt-5.verifiedAtModel fit statement fit-gpt-5-superseded rests on this release and was verified 2026-08-18. Moving verifiedAt to 2026-08-27 would make that guidance stale, and this run gathered no evidence about fit statements. Capping the date to fit the invariant would assert a verification that never happened.Blocked byfit-gpt-5-supersededvalidate.tsreleases/anthropic-claude-mythos-5.verifiedAtModel fit statement fit-claude-mythos-5-limited-availability rests on this release and was verified 2026-08-15. Moving verifiedAt to 2026-08-27 would make that guidance stale, and this run gathered no evidence about fit statements. Capping the date to fit the invariant would assert a verification that never happened.Blocked byfit-claude-mythos-5-limited-availabilityvalidate.tsreleases/meta-llama-4-scout.verifiedAtModel fit statement fit-llama-4-scout-self-hosting rests on this release and was verified 2026-08-15. Moving verifiedAt to 2026-08-27 would make that guidance stale, and this run gathered no evidence about fit statements. Capping the date to fit the invariant would assert a verification that never happened.Blocked byfit-llama-4-scout-self-hostingvalidate.ts
Sources conflict, so no value changed
google-gemini-embedding-2-api-code-conflictGoogle sources disagree on whether the Gemini Embedding 2 API identifier is gemini-embedding-2 or gemini-embedding-2-preview. All three rubrics accepted it as a conflict; a conflict is recorded as a finding and never applied to the dataset.Blocked bygoogle-gemini-models-referencegoogle-gemini-deprecations
Source refused by the approval gate
openai-gpt-5-6-cyber-daybreak-red-aliasOpenAI's Daybreak Red model documentation lists daybreak-red-latest as an alias for gpt-5.6-cyber. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsopenai-gpt-5-6-sol-daybreak-blue-aliasOpenAI's Daybreak Blue model documentation lists daybreak-blue-latest as an alias for gpt-5.6-sol. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsopenai-gpt-realtime-mini-status-conflictOpenAI's model catalog labels GPT-Realtime Mini deprecated, while its individual model page labels it default. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsanthropic-claude-fable-5-intended-use-biology-caveatClaude Fable 5's intended-use text should include Anthropic's caveat that dual-use biology requests still fall back to Opus 5. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsanthropic-claude-haiku-4-5-status-activeClaude Haiku 4.5 remains an active latest Claude model. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsanthropic-claude-haiku-4-5-context-windowClaude Haiku 4.5 still has a 200,000-token context window. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsanthropic-claude-haiku-4-5-api-aliasesClaude Haiku 4.5 still uses claude-haiku-4-5 as a convenience alias for claude-haiku-4-5-20251001. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjs
Blocked by policy before it could run
long-tail-creatorsQwen, zai-org, MiniMaxAI, deepseek-ai, moonshotai, sensenova, Lightricks and ornith-ai are trending on an approved origin, but organizationSchema requires website and releasePage URLs that no approved-origin page states, and each creator own host is not an approved origin. Approving a new origin is a human decision, not one a refresh run may take for itself.Blocked byorganizationSchemaapproved-origin list
What this run does not prove
- No human reviewed the pull request. The evidence trail, the deterministic gates and the required web-ci check are the whole of the oversight.
- A green run proves the dataset is internally coherent and every applied claim carried a quote from an approved origin. It does not prove the sources are right.
- Half the proposed claims did not survive. The dataset is more complete than it was, not complete.
- Two page budgets were exhausted and two hosts refused this client, so absence of a record here is not evidence that the model does not exist.
- The per-creator applied counts in summary issue #418 and its gate-dataset source count disagree with the pull request body. The figures here follow the pull request body and the merge commit itself, which agree with each other.
Follow-ups — proposed, not fixed
- Scouts are not yet running check-bundle-pairing, which landed on main mid-run and would have caught this run largest failure class at bundle time.
- The GPT-Realtime Mini status conflict was unanimously accepted and still could not be published, because both its sources were unpaired. The disagreement is real and remains unrecorded in the dataset.
- Long-tail creators have no route into the dataset under the current approved origins, and approving an origin is a human decision.
- The reviewed profiles and the dataset disagree about canonical release names, which drove most editorial rejections and will recur every run until it is settled.
- Three model fit statements are older than the release facts they rest on and block those releases from carrying this run verification date.