What each refresh found, and what it did not publish.
ModelTree's dataset is refreshed by agents against primary sources, reviewed by an independent three-rubric panel, and gated deterministically. Every run is recorded here in full — including the runs that published nothing.
A run's working state is never committed. This page transcribes the durable record: the pull request body and the summary issue, both linked from each entry.
- Runs recorded
- 23
- Pages fetched
- 2,128
- Claims proposed
- 1,352
- Edits published
- 611
- Items withheld
- 439
Showing 11–16 of 16 runsfiltered by Published
Clear filter2026-08-31-ae0342
Data refresh 2026-08-31
Scope requested: Six of the seven pilot creators — anthropic, openai, meta, microsoft, alibaba-cloud and google-deepmind — each judged at the 2-of-3 pilot threshold, plus amazon scouted without producing a claim. Long-tail creators were not scouted. Re-verification only, with the model-fit guidance layer deliberately taken first so the release and family records resting on it could move.Published33 edits posted · 1 item withheldAn agent-run, source-backed refresh under ADR 0003, and a narrow-but-deep run rather than a broad one: it set out to clear a blockage the 2026-08-30 run recorded and could not clear itself. That run had five verification dates it was entitled to advance and could not, because validate.ts refuses to let a model-fit statement be older than the fact it rests on, and the guidance layer had not been re-verified. This run re-verified the guidance layer first — all seven statements in model-fit-statements.json — which unpinned every one of the five. Six creators were scouted against 32 pages fetched and hashed this run, producing 43 claims; three rubrics voted blind on each, 42 met the 2-of-3 pilot bar and one was refused. 33 edits landed across four documents and nothing was withheld, against two withheld last run. The cost is breadth: the 33 long-tail creators were not scouted at all this run, which is the largest gap in this entry and is recorded under notCovered rather than smoothed over. No human reviewed any claim, verdict or edit.
- PreflightRanTree clean, gh authenticated, and no open pull request from a previous refresh at the moment this run started. Anchor 356989e9, selected by merge-base with refs/remotes/origin/main and computed by the gate rather than supplied — requestedBase was null. That is a re-anchor: the run opened on 36ce8b1b, and run 2026-08-31-651a1c merged as PR #665 while this one was between its review and publish stages, moving main under it and touching three of the same documents. The branch was brought onto the new main, the dataset reset to main’s version wholesale rather than hand-merged, and every claim re-checked against it — all 43 currentValues still held, so #665 had touched none of the records this run changes. gate-source-approval, gate-dataset, npm run validate, gate-scope and gate-ledger were then all re-run against the new anchor rather than carried over. The approved-origin catalogue parsed cleanly at 356989e9: 217 dataset sources and 18 catalogues across 19 profile files, yielding 46 approved origins. npm resolved to a PowerShell shim the machine execution policy refuses to run; npm.cmd ran and reported 11.9.0. The policy was not changed to work around it — that would be a machine-wide security change made to satisfy a probe.
- ScoutRanSeven pilot creators scouted against 32 pages fetched and hashed this run; six produced claims and amazon produced none. Every quote was sliced from a page read during the run and verified as a byte-exact substring of the saved body before the claim was built — raw bytes for markdown and plain text, rendered text for HTML. Two candidate quotes failed that check and were corrected rather than kept: one had a curly apostrophe where the page has an ASCII one, and one fell under the gate’s 24-character minimum. Appending .md to an OpenAI or Anthropic docs URL returns clean markdown of the same page on the same already-approved origin, which is how those client-rendered catalogues were read at all; the provenance rubric was asked to rule on that convention explicitly and accepted it. ai.google.dev was unreachable from this machine on all three attempts — fetch failed, not an HTTP status — so deepmind.google was used instead, an origin the dataset already carries.
- ReviewRanThree rubrics — provenance, consistency, editorial — voted independently on all 43 claims across 18 reviewer invocations, each in its own context and each given only its own rubric text, the claim, the evidence, the relevant dataset slice and the creator profile vocabulary. No reviewer saw another’s rubric, the threshold, the running tally, or the scout’s reasoning. 42 claims met the 2-of-3 pilot bar and one was refused 2-of-3. Three of the 42 passed over a single dissent, and all four minority opinions are published verbatim in the pull request body. No claim was edited and re-reviewed to chase a verdict, and the coordinating agent cast no vote and overruled none.
- GatesRangate-evidence and gate-source-approval ran across all six bundles before anything was applied, since gate-source-approval anchors on the committed dataset and running it afterwards would be asking the run’s own writes whether the run’s own sources are trustworthy. gate-dataset, npm run validate, gate-scope and gate-ledger ran afterwards. No gate was skipped, forced, or re-run to obtain a different answer, no threshold was lowered, and nothing was pushed to main. gate-scope reported changed: 5, empty: false, outOfClass: [], so the change sat inside the ADR 0003 qualifying class as ADR 0006 widened it.
- PublishRanOpened as a pull request carrying the full evidence trail — every claim, every quote with its content hash, all 129 verdict rationales, and the gate-source-approval anchor block — and merged by GitHub via --auto once web-ci went green. The agent did not merge, and did not use --admin.
- DeployRanPages deploy confirmed after the merge. See the deployment reference below.
What was found
- Scouts
- 7
- Pages fetched and hashed
- 32
- Claims proposed
- 43
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims anthropic pilot 2-of-3 12 openai pilot 2-of-3 12 meta pilot 2-of-3 10 microsoft pilot 2-of-3 3 alibaba-cloud pilot 2-of-3 4 google-deepmind pilot 2-of-3 2 What those claims proposed to do Kind Count Effect Unchanged 20 A recorded fact re-read from its primary source and found to still hold. 19 reached the 2-of-3 pilot threshold and all 19 applied, advancing 9 release verifiedAt dates, 1 family verifiedAt and 7 model-fit statement verifiedAt dates to 2026-08-31 — some records carried more than one confirming claim. Nothing was withheld. One was refused by the panel. Change 23 A source page re-read this run, moving its lastCheckedDate to 2026-08-31. All 23 were accepted, and 16 distinct source records moved — several pages carried more than one claim. Not covered
- The 33 long-tail creators were not scouted this run. That is the largest gap in this entry: the 2026-08-30 run covered all 33, and this one traded that breadth for depth on the guidance layer that was blocking pilot-creator re-verification. Their records carry no verification date from this run and nothing here says their facts still hold.
- amazon was scouted and produced no claim. aws.amazon.com/nova/ is client-rendered and the fetched body carried placeholder text where the model listing should be, so no quote on it could support a fact about a model. The one recorded Amazon release, amazon-nova-2-sonic, was verified 2026-08-30 and did not need this run.
- ai.google.dev was unreachable from this machine on every attempt — fetch failed rather than an HTTP status, so it is a network fact about this runner and not evidence the page is gone. Seven of the eight recorded Google releases rest on sources on that host and therefore carry no verification date from this run. Only google-gemini-pro-page, on the reachable deepmind.google, was re-read.
- Breadth remains uncovered by design and is unchanged from the previous run: OpenAI’s realtime, transcribe and tts models, the roughly twenty Cohere Command models against the one recorded, Microsoft’s MAI-Code-1.1-Flash, MAI-Image-2.5, MAI-Voice-2 and MAI-Transcribe-1.5, and Amazon’s Nova Act, Nova Forge and Nova Multimodal Embeddings. Each is breadth work with its own sourcing burden, not re-verification.
- Microsoft MAI-Image-2.6 was again deliberately not added, on the same grounds the 2026-08-30 run recorded: its only availability statement names a product and a serving platform rather than the creator’s own release, so accessType has no creator-level statement and status is ambiguous between preview and current. The ambiguity is left explicit rather than resolved by guess.
What was evaluated
- Reviewers
- 18
- Verdicts cast
- 129
- Accepted by panel
- 42
- Rejected by panel
- 1
Deterministic gates and required checks — 8 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence.mjs6 claim bundles 0 Pass43 claims admissible with complete panels across all six bundles, 23 applicable after the panel. Every bundle declared policy pilot and the gate independently derived threshold 2 from tools/updater/profiles, so no bundle’s declared policy went unchecked. gate-source-approval.mjs6 claim bundles 0 PassZero proposed sources across all six bundles: every cited source was already in the dataset at anchor 356989e9, on one of the 46 origins approved there. No source was refused and no origin was introduced, so the run approved nothing of its own. Run before any edit was applied. gate-dataset.mjsweb/src/data 0 PassValidated the documents raw.ts composes with zero failures after the edits were applied. npm run validateweb/ 0 Pass2361 tests executed across 104 files with all reporting results; astro check reported 0 errors and 0 warnings. This is the gate that pinned five records last run, and it passed them this run because the guidance beneath them had been re-verified first. It also failed intermittently on this runner for a reason that is not this change, recorded here rather than re-run until green: two to four tests in LineageModelDrawer.interaction.test.tsx and ModelTreeExplorer.interaction.test.tsx hit the 5000 ms vitest timeout, with different test names failing on each attempt. Both files pass 17 of 17 in isolation with this change applied, and pristine origin/main fails the same way with more failures than this branch, so the cause is worker contention on the machine and not the data. Constraining vitest to two workers produced the clean run recorded here. No timeout was raised, no test was skipped, and no gate was called passed on a run that had not finished. gate-scope.mjsbranch vs merge-base 0 Passchanged: 5, empty: false, outOfClass: []. Anchored on 356989e9, the merge base the gate computed rather than one supplied to it. The fifth file is this ledger entry, which ADR 0006 put in the qualifying class so a run can record itself. gate-ledger.mjsthis entry vs the diff it describes 0 Passtranscription: false — this entry was written by the run it describes, in the same commit as the data, rather than transcribed afterwards. That is the gap issue #419 recorded on three consecutive runs. ci-preflight.mjsnot requiredselected from the branch diff 0 PassRun locally from the repository root before the branch was pushed. web-cithe pull request — PassGitHub performed the merge via --auto on this check going green; the agent did not merge. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
anthropic-family-claude-4-5-status-unchangedProvenanceOutvoted 2-of-3 and applied. The overview page’s "Legacy models (still available)" list names Claude Opus 4.5 and Claude Sonnet 4.5 but not Claude Haiku 4.5, which appears instead in the current-lineup table on the same page. So the 4.5 generation’s members straddle both lists, and a quote establishing that one member is legacy does not establish a status for the family as a whole.meta-llama-4-scout-context-window-unchangedProvenanceOutvoted 2-of-3 and applied. The 10M figure sits in a table-row fragment whose column headers are not carried in the quote, so identifying which column the number belongs to requires an inference the quoted text does not itself supply.meta-fit-llama-4-scout-context-window-reverifyProvenanceOutvoted 2-of-3 and applied, on the same reasoning as the release-level context-window claim it rests on: the underlying table-row fragment does not carry the column headers that would fix which figure is the context window.
Posted 33 edits
33 edits across 4 documents, a net change of 0 records.
Dataset documents this run changed Document Before After What changed model-fit-statements.json7 7 All 7 statements had verifiedAt moved to 2026-08-31: fit-claude-haiku-4-5-release-current, fit-claude-haiku-4-5-family-legacy, fit-claude-mythos-5-limited-availability, fit-gpt-5-superseded, fit-llama-4-scout-self-hosting, fit-llama-4-scout-mau-threshold and fit-llama-4-scout-context-window. This document was taken first on purpose. validate.ts forbids a fit statement from being older than the fact it rests on, so while these sat at 2026-08-15 and 2026-08-18 they pinned the records beneath them — which is exactly why the previous run had to withhold two release dates it had otherwise earned. No statement text or judgement changed; only the date on which each was last checked against its evidence. releases.json92 92 9 verifiedAt dates moved to 2026-08-31 — anthropic-claude-haiku-4-5, anthropic-claude-mythos-5 and meta-llama-4-scout from 2026-08-15, openai-gpt-5 from 2026-08-18, openai-gpt-image-2 from 2026-08-27, microsoft-fara-1-5-27b, alibaba-qwen3-8-27b and alibaba-qwen3-8-2-4t-a95b from 2026-08-28, and openai-gpt-5-6-sol from 2026-08-29. The first three and openai-gpt-5 are the records the previous run could not move. No release was added or removed and no other field changed. families.json64 64 One verifiedAt moved: anthropic-claude-4-5 from 2026-08-15 to 2026-08-31, unpinned by fit-claude-haiku-4-5-family-legacy being re-verified first. The family status value itself was re-read and unchanged, and it carried the run’s one accepted-over-dissent family claim. sources.json217 217 16 lastCheckedDate values moved to 2026-08-31, one per page re-read this run. No source was added: gate-source-approval recorded zero proposed sources across all six bundles, so every date moved on a record that already existed. Sources are listed here rather than under the posted records below, which carry links and so hold only the collections the site routes. Not posted 1 item
Rejected by the review panel
releases/google-gemini-3-1-pro-preview.verifiedAtRefused 2-of-3 under the pilot bar, the run’s only rejection. Two rubrics found the evidence too weak for a status claim: the specification block flattens to an unseparated run of text once tags are stripped, so "Status" and "Preview" are adjacent rather than demonstrably paired, and the page’s marketing name "3.1 Pro" does not distinguish the Gemini API serving from the Vertex AI serving. The record keeps 2026-08-27. Its source record still had lastCheckedDate advanced, because that claim is about the page having been read rather than about what the page proves.Blocked byprovenance rubriceditorial rubric
What this run does not prove
- No human reviewed any claim, verdict or edit in this run before it merged and deployed. That is what ADR 0003 authorises and what it costs.
- The three rubrics are three instances of the same model family reading the same page, so they share a failure mode: a source that is itself wrong can carry all three. The panel defends against wrong-but-invalid and against unevidenced inference; it is a weak defence against a primary source that is confidently mistaken.
- Three claims were accepted over a provenance dissent and are published as such. Two of them — the Llama 4 Scout context window and the fit statement resting on it — turn on a table-row fragment whose column headers the quote does not carry. The majority judged the row unambiguous in context; the dissent judged the fragment insufficient on its own. A reader who agrees with the dissent should treat those two verifiedAt advances as the weakest edits in this run.
- A green run proves the recorded facts still match the pages this run read. It does not prove those pages are complete, and it says nothing at all about the 33 long-tail creators or the seven Google releases whose sources could not be reached.
- pagesFetched counts the 32 pages that were fetched successfully. Three further attempts against ai.google.dev failed outright and are excluded from that figure rather than counted as reads.
- reviewers and verdictsCast are derived rather than quoted: 6 bundles x 3 rubrics = 18 reviewer invocations, and 43 claims x 3 rubrics = 129 verdicts, which reconciles with the per-rubric verdict files.
- This run and run 2026-08-31-651a1c overlapped. That one merged first, as PR #665, while this one was mid-flight, so this run had to re-anchor from 36ce8b1b onto 356989e9 and re-run every gate. The re-anchor was clean — all 43 claims still matched the dataset afterwards — but two refreshes racing on the same documents is the failure mode the preflight check is meant to prevent, and the check only looks for an open pull request at the moment the run starts. Two same-day runs appear in this ledger as a result.
Follow-ups — proposed, not fixed
- The long-tail sweep this run skipped should be the next run’s first priority, not its second. Two consecutive runs covering different subsets is fine; a pattern of pilot-only runs would let the long tail rot while every entry still reads green.
- ai.google.dev could not be reached from this runner at all. If that is a durable block rather than a transient one, the Google releases need an alternative approved origin recorded in the reviewed catalogue — deepmind.google covers only the Pro page — or seven records will keep aging out with no way to re-verify them.
- The model-fit guidance layer pinned five records for a full run before this one cleared it. Re-verifying model-fit-statements.json first should become the standing order for a re-verification run rather than a lesson each run rediscovers, since guidance is cheap to re-check and blocks everything beneath it.
- Client-rendered marketing pages — aws.amazon.com/nova/ this run, and the OpenAI and Anthropic docs before the .md convention was found — yield placeholder text to a plain fetch. The .md suffix that works for OpenAI and Anthropic docs is worth recording in the reviewed catalogues so future runs do not rediscover it, and an equivalent for the AWS pages is worth looking for.
- The preflight check for a racing refresh only looks for an open pull request at the moment the run starts, so it cannot see one opened afterwards — which is exactly what happened here with PR #665. A run that re-checks immediately before it opens its own pull request, or that simply expects to re-anchor, would lose less work than one that discovers the race at merge time.
- Status claims sourced from flattened specification tables were refused this run for the second time in two runs, on the same structural grounds as the previous run’s context-window refusals. Either the scout should quote enough surrounding structure to fix the label-value pairing, or status re-verification should target a page that states it in prose.
2026-08-31-651a1c
Long-tail family depth, tranche 1 of abdeslam-menacere/ModelTree#651
Scope requested: Six creators from abdeslam-menacere/ModelTree#651 tranche 1: alibaba-cloud, zhipu-ai, moonshot-ai, baidu, tencent, 01-ai. Breadth rather than re-verification — asking whether each has a further documented family generation, not re-checking facts already recorded. No creator outside the tranche was scouted and no existing record was edited.Published9 edits posted · 7 items withheldA source-backed breadth run over six creators that each carried exactly one model family — alibaba-cloud, zhipu-ai, moonshot-ai, baidu, tencent and 01-ai — asking of each whether it has further documented family generations. Two do, on evidence that meets the bar: Alibaba published dated Qwen3.5 and Qwen3.6 news in the QwenLM/Qwen3.8 repository README, and 01.AI dated the original Yi series in 01-ai/Yi. Four do not, and the reason is the same one refresh run 2026-08-30-605b1a hit: their pages carry no dated release statement at all. The GLM, Kimi, Hunyuan and ERNIE repository READMEs have no dated news list, every date on their Hugging Face cards is Hub createdAt or lastModified metadata rather than a creator statement, and the z.ai GLM-5 blog URLs return 200 with a client-rendered shell under 600 bytes and no text. Nine claims were proposed, three rubrics voted independently on every one, and all nine met their creator threshold — the six Alibaba claims against a 2-of-3 pilot bar, the three 01.AI claims unanimously against the 3-of-3 long-tail bar. Three met it over a recorded provenance objection about mapped vocabulary fields, and are applied rather than overruled. Four creators are withheld, on the record and without softening. No human reviewed any part of the research; the run hands off at a reviewable commit for the independent review and QA gates.
- PreflightRanAnchor 8e8c319e, selected by merge-base with refs/remotes/origin/main and computed by each gate rather than supplied — requestedBase was null in every report. The approved-origin catalogue parsed cleanly at that anchor: 214 dataset sources and 18 profile catalogues out of 19 profile files, yielding 46 approved origins; the one profile listed without a catalogue is tools/updater/profiles/generic/long-tail.json, which configures no origins. Both npm and drydock were probed in both shim forms before either was relied on or ruled out: the bare names resolve to PowerShell shims the execution policy refuses, while npm.cmd reports 11.9.0 and drydock.cmd reports 0.1.0. Both are installed and blocked in bare form, not absent, and the execution policy was not changed. Dependencies were installed with npm ci; package-lock.json is unmodified.
- ScoutRanSix creators scouted against 48 pages, every one fetched and hashed during the run — retrieval is "fetch" on all 43 evidence entries and no claim rests on a search snippet. Each quote was verified to be literally present in the hashed body it is attributed to, by a builder that throws rather than emitting an unverifiable pair. The decisive finding is negative and general: model cards do not state release dates. The only pages in this tranche that date a release in the creator's own words are two GitHub README news lists, QwenLM/Qwen3.8 and 01-ai/Yi. Red herrings ruled out rather than used: PaddlePaddle/ERNIE's dated "Recent updates" dates the ERNIEKit toolkit and not a model; the Kimi-K2.5 changelog entry for 2026.1.29 is a system-prompt removal; Kimi-K3's "July 9, 2026" dates an evaluation branch; and Tencent's "Following the Hy3 Preview launch in late April" is about the preview and is not a valid partial date.
- ReviewRanThree rubrics — provenance, consistency, editorial — voted independently on all nine claims, each blind to the others and to the scout's reasoning, seeing the claim bundle and the committed dataset only. 27 verdicts cast, every one with a rationale, all published verbatim. Consistency and editorial accepted all nine. Provenance accepted six and objected to three, in each case that a quote naming "Type: Causal Language Model with Vision Encoder" or open-weight availability does not by itself force the mapped modality or category vocabulary. All three sit in the alibaba-cloud bundle, whose threshold is 2-of-3 because alibaba-cloud is the only creator of the six with a reviewed profile on disk in tools/updater/profiles. They therefore met their threshold over a recorded objection and were applied, not overruled; they are listed in dissents. No claim was revised and re-reviewed to chase a verdict, no reviewer was re-run, and no threshold was adjusted.
- GatesRangate-evidence and gate-source-approval ran across all six bundles before anything was applied, in that order and after the panel rather than before it. gate-dataset, npm run validate, gate-scope, gate-ledger and ci-preflight.mjs ran afterwards. No gate was skipped, forced, or re-run to obtain a different answer. gate-scope reports exit 1 by design and is the one number worth reading twice — see its entry in evaluated.gates and the caveats.
- PublishRanNine edits applied across three dataset documents in dependency order — sources, then families, then releases — and committed on the dock branch for abdeslam-menacere/ModelTree#651 together with this entry. Nothing was pushed, no pull request was opened and no merge was attempted: this run is a Drydock dock and its work ends at a reviewable commit, with the pull request and the merge belonging to the coordinating session after the independent review and QA gates have passed against that commit.
- DeployNot runNothing was pushed or merged by this run, so no Pages deploy was triggered and none could be observed. Recorded as not-run rather than not-applicable: a deploy is applicable to a dataset change, it simply has not happened yet at the time this entry was written.
What was found
- Scouts
- 6
- Pages fetched and hashed
- 48
- Claims proposed
- 9
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims alibaba-cloud pilot 2-of-3 6 01-ai long-tail 3-of-3 3 zhipu-ai long-tail 3-of-3 0 moonshot-ai long-tail 3-of-3 0 baidu long-tail 3-of-3 0 tencent long-tail 3-of-3 0 What those claims proposed to do Kind Count Effect Add 9 Three sources, three families and three releases: the Qwen3.5 and Qwen3.6 generations with one sourced open-weight variant each, and the original Yi series with Yi-34B-Chat. Not covered
- The other dated Qwen3.5 and Qwen3.6 variants — Qwen3.5-122B-A10B, Qwen3.5-35B-A3B and Qwen3.5-27B on 2026-02-24, and the 2026-04-22 Qwen3.6 entry. The QwenLM/Qwen3.8 news list dates them, but no model card was fetched and hashed for them this run, so no release record could carry a sourced parameter count, context window or licence.
- The other dated Yi releases — Yi-34B-200K on 2023-11-05, Yi-VL on 2024-01-23, Yi-9B on 2024-03-06 and Yi-9B-200K on 2024-03-16. Same reason: dated in 01-ai/Yi, but no card fetched for them here.
- Lineage between the new and existing families. Nothing wires Qwen3.5 to Qwen3.6 to the committed Qwen3.8, or the new Yi family to the committed Yi-1.5, because predecessorIds and successorIds on an existing record would be a change claim against a record this tranche was not asked to touch.
- The 34 creators outside this tranche, including the other 26 that still carry exactly one family.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 27
- Accepted by panel
- 9
- Rejected by panel
- 0
Deterministic gates and required checks — 6 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidenceall six claim bundles, before any dataset document was touched 0 Passpassed=true on each of the six. alibaba-cloud: 6 claims, 6 applicable, threshold 2 under the pilot policy. 01-ai: 3 claims, 3 applicable, threshold 3 under long-tail. The four withheld bundles carry 0 claims and pass trivially. The threshold was derived by the gate from the reviewed-profile set on disk and not from any bundle's policy field. gate-source-approvalall six claim bundles, anchored at merge-base 8e8c319e 0 Passpassed=true on each of the six, over 43 citations. Anchor 8e8c319e, selectedBy "merge-base with refs/remotes/origin/main", requestedBase null, 214 dataset sources and 46 approved origins. Inherited sources: qwen3-8-repository, osi-approved-licenses, 01-ai-yi-repository. Proposed sources: qwen3-5-397b-a17b-model-card, qwen3-6-35b-a3b-model-card, 01-ai-yi-34b-chat-model-card. Every citation sits on an origin the anchor already approves — huggingface.co, github.com and opensource.org — so the run introduced no new origin and approved nothing of its own. gate-datasetweb/src/data after the nine claims were applied 0 Passpassed=true, failures empty. Counts after: 217 sources, 48 publishers, 40 organizations, 64 families, 92 releases. npm run validateweb/ — the full vitest suite plus Astro and TypeScript diagnostics 0 Pass2361 tests across 104 files. It failed on the first run, and the failure was real rather than incidental: the comparison picker index budget. Recorded in caveats, because the repair reaches outside the dataset. gate-scopenot requiredthe branch diff against merge-base 8e8c319e 1 FailExit 1, and the correct answer. The diff reaches web/src/lib/comparison.test.ts, which is outside the nine documents ADR 0003 lets merge unattended, so this change is not eligible for an unattended merge and the gate says so. It is recorded as not required because this run never sought that route: it is a Drydock dock handing off to the independent review and QA gates and then to a human-opened pull request, where a scope refusal is a fact about eligibility rather than a blocked merge. The change was not trimmed to satisfy the gate; the reason it reaches out of class is in the caveats. gate-ledgerthis entry against the branch diff, anchored at merge-base 8e8c319e 0 Passpassed=true. The three declared documents match the three dataset documents the branch changed, in both directions, and their record counts were counted at the anchor and in the working tree rather than taken from this entry. ci-preflightrepository root — the pull-request checks this branch's diff actually triggers 0 PassSelected and ran the checks the diff triggers, measured from the computed merge-base: 5 files changed, 3 of 7 local check groups selected, all three passing — web-ci, skills-ci and source-link-health-tests. It does not cover the networked link-health sweep or the second Python interpreter, and says so on every run; a green preflight is therefore not a green CI. It was run twice: the first run reported web-ci as failing with "npm run test" exiting 1, and the second passed with 104 test files and 2,361 tests. See the caveats — the first result was never reproduced and is recorded rather than explained away. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
qwen3-6-family-addProvenanceThe date and first-open-weight-variant facts are quoted, but no attached quote for this claim supports the proposed multimodal-generalist category. The quoted Qwen3.6 text mentions open-weight availability and coding/repository reasoning, not vision or multimodality.alibaba-qwen3-5-397b-a17b-release-addProvenanceThe date, weights, Apache licence, OSI badge, parameters, and 262,144-token context are quoted. But "Type: Causal Language Model with Vision Encoder" does not by itself force inputModalities ["text","image"] or the complete output modality list.alibaba-qwen3-6-35b-a3b-release-addProvenanceThe model-card quotes support open weights, Apache licensing, parameters, and native context. The attached "Type: Causal Language Model with Vision Encoder" quote does not force the proposed text+image input modalities or exclude other vision inputs.
Posted 9 edits
9 edits across 3 documents, a net change of 9 records.
Dataset documents this run changed Document Before After What changed sources.json214 217 Three Hugging Face model cards: Qwen3.5-397B-A17B, Qwen3.6-35B-A3B and 01-ai/Yi-34B-Chat. They are not listed individually under records because the refresh page resolves links for release and family records only, and a source record listed there would link nowhere. families.json61 64 The Qwen3.5 and Qwen3.6 generations, and the original Yi series distinct from the committed Yi-1.5. releases.json89 92 One sourced variant per new family, being the one whose model card was fetched and hashed this run. Each document links to the file as this run left it, not as it stands today.
Records added
- Qwen3.5in the treefamilies
qwen3-5First release 2026-02-16, per the QwenLM/Qwen3.8 news list. - Qwen3.6in the treefamilies
qwen3-6First release 2026-04-16, per the same news list. - Yiin the treefamilies
01-ai-yiFirst release 2023-11-02, the original Yi series; the committed Yi-1.5 family stays separate. - Qwen3.5-397B-A17Bpassportreleases
alibaba-qwen3-5-397b-a17b397B total, 17B active, 262,144 native context, Apache-2.0. - Qwen3.6-35B-A3Bpassportreleases
alibaba-qwen3-6-35b-a3b35B total, 3B active, Apache-2.0. - Yi 34B Chatpassportreleases
01-ai-yi-34b-chatDated 2023-11-23 by the 01-ai/Yi news list. Parameters and context window deliberately omitted, matching the committed Yi-1.5 34B Chat record.
Not posted 7 items
Verification date deliberately held back
zhipu-ai-further-generationsGLM-5, GLM-5.2, GLM-5.3 and GLM-5.3-Flash all exist as Hugging Face repositories under zai-org, and none of the pages fetched states a release date. The GLM repository READMEs carry no dated news list — scans for YYYY-MM-DD, YYYY.MM.DD and month-name forms returned nothing — and every date visible on the cards is Hub createdAt or lastModified metadata, which refresh run 2026-08-30-605b1a already ruled out as a release date. https://z.ai/blog/glm-5 and https://z.ai/blog/glm-5.3 return HTTP 200 but are client-rendered shells under 600 bytes with no text, and https://z.ai/blog is 404. A family cannot be added without a sourced firstReleaseDate, so nothing was proposed.Blocked byno dated release statement on any fetched pagez.ai blog pages render client-side and serve no textmoonshot-ai-further-generationsKimi-K2.5 and Kimi-K3 exist as Hugging Face repositories under moonshotai, and neither card nor the MoonshotAI GitHub organisation pages state a release date. The two date-like strings found were checked and are not release dates: the Kimi-K2.5 changelog entry for 2026.1.29 records a system-prompt removal, and Kimi-K3's "July 9, 2026" dates an evaluation branch. Every other date is Hub metadata.Blocked byno dated release statement on any fetched pagebaidu-further-generationsERNIE-5.0 material and further ERNIE-4.5 variants exist under the baidu Hugging Face namespace, and no fetched page dates a release. PaddlePaddle/ERNIE does carry a dated "Recent updates" list, but it dates the ERNIEKit toolkit rather than a model, which is a different entity and was not used.Blocked byno dated release statement on any fetched pagethe only dated list found dates a toolkit, not a modeltencent-further-generationsHunyuanImage-3.0, HunyuanOCR and Hy3/Hy4 material exist under Tencent-Hunyuan, and no fetched page dates a release. The nearest statement, "Following the Hy3 Preview launch in late April", is about the preview rather than the release and "late April" is not a valid partial date in this schema. Recording it as one would be the guess this run exists to refuse.Blocked byno dated release statement on any fetched page"late April" is not a representable date
Sources conflict, so no value changed
01-ai-yi-licence-conflictThe Yi-34B-Chat card states in prose that "The code and weights of the Yi series models are distributed under the Apache 2.0 license", while the 01-ai/Yi README records "2023-11-23: The Yi Series Models Community License Agreement is updated to v2.1". Both quotes are attached to the release claim and the disagreement is left explicit in the record's prose rather than resolved. The structured licence follows the model card, which is the artefact the release record is about.Blocked bytwo primary pages of the same creator disagree
Out of the run’s reach
qwen3-5-and-qwen3-6-additional-variantsThe QwenLM/Qwen3.8 news list dates Qwen3.5-122B-A10B, Qwen3.5-35B-A3B and Qwen3.5-27B to 2026-02-24 and a further Qwen3.6 entry to 2026-04-22. The dates are sourced, but no model card was fetched and hashed for those variants this run, so no release record could carry a sourced parameter count, context window or licence. One release per new family was proposed rather than an unsourced set.Blocked byno model card fetched for these variants this run01-ai-yi-additional-variantsThe 01-ai/Yi news list dates Yi-34B-200K to 2023-11-05, Yi-VL to 2024-01-23, Yi-9B to 2024-03-06 and Yi-9B-200K to 2024-03-16. Same reason as the Qwen variants: dated, but no card fetched for them here.Blocked byno model card fetched for these variants this run
What this run does not prove
- The outcome is recorded as "published" because that is the only value the schema admits for a run that changed dataset documents, and it is not the whole truth at the time of writing. This run is a Drydock dock: its nine edits are applied and committed on a branch, and nothing has reached main. No pull request existed when this entry was written, which is why references names no pull request of this run's own — the pull-request reference below is the change whose merge set this run's baseline, labelled as such. Whoever opens the pull request should add it, its merge commit and the Pages deploy to references.
- npm run validate failed on its first run, on the comparison picker index page-weight budget, and the repair reaches outside the dataset into web/src/lib/comparison.test.ts. The budget was raised from 10,240 to 11,264 bytes as a deliberate page-weight decision, on the test message's own terms: measured at merge-base 8e8c319e the index was 10,115 bytes over 89 releases (113.65 per release) and at the tip it is 10,449 over 92 (113.58 per release), so the catalogue simply grew and the per-release figure did not move. The scale-invariant guard of 128 bytes per row was not touched and keeps 14 bytes of headroom. That one file is why gate-scope exits 1 and why this change cannot merge unattended. It is also worth knowing that the merge-base already sat 125 bytes under the old budget — about one release — so this was going to land on whichever tranche came first.
- ADR 0005's accepted limit applies to every hash and quote here: gate-evidence checks that a content hash is well-formed and that a quote is long enough, never that either matches the remote page, and both are self-authored by the run. What compensates for it in this run is mechanical but local: every page was fetched to disk, hashed from the bytes on disk, and every quote was checked to be literally present in the body it is attributed to, by a builder that throws rather than emit an unverifiable pair. That makes the pairs internally consistent. It does not make them independently verified, and a later reader who wants certainty must refetch.
- Three claims were applied over a provenance objection, all of them about controlled-vocabulary fields — the modality lists and one family category. The objection is recorded in full in dissents and is not answered here. A reader who thinks provenance was right should read those three records as the weakest in this change.
- status: "current" on the three new releases is the least directly sourced mapping in the run. No page states a lifecycle state in words; the value rests on the news lists saying the weights are released and available and on the absence of any deprecation or legacy notice. Qwen3.5 and Qwen3.6 are both superseded by the committed Qwen3.8, and the original Yi by Yi-1.5, so a reader who reads "current" as "newest" will be misled. It means "not withdrawn".
- The input and output modality lists on the two Qwen releases record text and image while both cards also carry a video-input quickstart. The competing signal is disclosed in the source notes and the record summaries rather than resolved, and it is the substance of two of the three provenance objections.
- Four of the six creators were withheld, and no negative claim is being made about them: the run did not establish that GLM, Kimi, Hunyuan or ERNIE have no further generations, only that no page it fetched states a release date for one. A page it did not fetch, or a page rendered server-side rather than in the browser, could settle any of the four.
- Every fetch happened during one run on 2026-08-31 and each contentHash is a snapshot of that moment. Hugging Face and GitHub pages change without notice, so a refetch that produces a different hash is expected rather than evidence of an error.
- One test result in this run was not reproducible. The first ci-preflight run reported web-ci failing with "npm run test" exiting 1, and printed no failing test name that was captured. Every subsequent run was green — a second ci-preflight with 104 test files and 2,361 tests passing, a bare npm run test with the same figures and the coverage verifier confirming all 104 discovered files reported, and two npm run validate runs. Four green runs do not turn a red one into a flake with any certainty, so it is recorded here as unexplained rather than dismissed, and a reviewer seeing web-ci fail once on this branch should suspect it is the same thing and not a new one.
Follow-ups — proposed, not fixed
- The committed release dates for moonshot-ai-kimi-k2-instruct (2025-07-11), zhipu-ai-glm-4-5-air (2025-07) and baidu-ernie-4-5-300b-a47b-pt (2025-06-28) are each exactly the Hugging Face Hub createdAt value of the cited repository, which refresh run 2026-08-30-605b1a explicitly ruled out as a release date. This run found that while looking for where those dates came from, did not fix it because the records are outside this tranche's scope, and records it here so it is not found a third time.
- Lineage between generations is unwired: nothing connects Qwen3.5 to Qwen3.6 to Qwen3.8, or the Yi family to Yi-1.5, although the Yi-1.5 card states it was continuously pre-trained from Yi. Wiring it means editing existing records and is worth its own issue.
- The remaining creators of the 32 that carried exactly one family are untouched by this tranche; 30 still do after it.
- The comparison picker index budget has now been raised three times as the catalogue grew, each time by 1,024 bytes. Whether the row itself should shrink — four string fields per release, none of them abbreviated — is a page-weight question nobody has asked yet, and asking it once would be cheaper than raising the number a fourth time.
2026-08-30-c0b6e9
Data refresh 2026-08-30
Scope requested: Every creator in organizations.json — all 33 — with pilot creators judged 2-of-3 and long-tail creators 3-of-3. Re-verification only: confirming recorded facts against their primary sources and moving verification dates forward, not adding breadth.Published36 edits posted · 9 items withheldAn agent-run, source-backed refresh under ADR 0003, and a re-verification run rather than a discovery one: not a single recorded fact value changed. All 33 creators in organizations.json were scouted, each against at least one primary page fetched and hashed this run, and 20 of them produced claims — 45 in total. Three rubrics voted blind on every one; 38 met their creator’s threshold and 7 long-tail claims died one vote short. 36 edits landed across two documents, all of them a verification date moving to 2026-08-30, before GitHub merged PR #597 on a green web-ci and Pages deployed successfully. The most useful finding is a refusal rather than an edit: five licence re-verifications failed because a model card saying “Apache-2.0” does not source the osiApproved flag recorded beside it, which is a claim about what OSI decided. No human reviewed any part of it. This entry was transcribed after the fact — see the caveats.
- PreflightRanAnchor 7ca5802e, selected by merge-base with refs/remotes/origin/main and computed by the gate rather than supplied — requestedBase was null. The approved-origin catalogue parsed cleanly at that anchor: 198 dataset sources and 12 profile catalogues out of 13 profile files, yielding 41 approved origins. The one profile listed but not drawn on is tools/updater/profiles/generic/long-tail.json, which configures no origins. The anchor is therefore full rather than silently narrowed.
- ScoutRanAll 33 creators were scouted, each against at least one primary page fetched and hashed this run; every quote in the pull request body was sliced from a page read during the run rather than recalled. Three creator hosts refused automated fetches and were worked around on approved origins instead: openai.com/news/ (403) via openai.com/news/rss.xml, ai.meta.com/blog/ (403) via huggingface.co/meta-llama, and x.ai/news (403) via docs.x.ai/developers/models. No claim rested on an unreachable page. 13 of the 33 creators produced no claim, because nothing they publish had changed in a way this run could evidence.
- ReviewRanThree rubrics — provenance, consistency, editorial — voted independently on all 45 claims, each blind to the others and to the scout’s reasoning. Consistency and editorial accepted every claim; provenance accepted 38 and rejected 7. No claim was revised and re-reviewed to chase a verdict. Every verdict and rationale is published verbatim in the pull request body.
- GatesRangate-evidence and gate-source-approval ran across all 20 bundles before anything was applied; gate-dataset, npm run validate, gate-scope and ci-preflight.mjs ran afterwards. No gate was skipped, forced, or re-run to obtain a different answer, no threshold was lowered, and nothing was pushed to main. gate-scope reported changed: 2, empty: false, outOfClass: [], so the change sat inside the ADR 0003 qualifying class and was eligible to auto-merge.
- PublishRanPR #597 carried the full evidence trail and was merged by GitHub via --auto once web-ci was green, not by the agent. Merge commit 2f490766, merged 2026-08-30T11:37:33Z.
- DeployRanPages deploy run 33309402080 on 2f490766 succeeded and https://abdeslam-menacere.github.io/ModelTree/ served HTTP 200. No revert needed.
What was found
- Scouts
- 33
- Pages fetched and hashed
- 45
- Claims proposed
- 45
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims 01-ai long-tail 3-of-3 2 ai2 long-tail 3-of-3 2 amazon pilot 2-of-3 2 anthropic pilot 2-of-3 7 baidu long-tail 3-of-3 2 bytedance-seed long-tail 3-of-3 2 cohere long-tail 3-of-3 2 databricks long-tail 3-of-3 2 deepseek long-tail 3-of-3 2 eleutherai long-tail 3-of-3 2 hugging-face long-tail 3-of-3 2 ibm long-tail 3-of-3 2 lg-ai-research long-tail 3-of-3 2 moonshot-ai long-tail 3-of-3 2 nvidia long-tail 3-of-3 2 sarvam-ai long-tail 3-of-3 2 snowflake long-tail 3-of-3 2 upstage long-tail 3-of-3 2 xai long-tail 3-of-3 2 zhipu-ai long-tail 3-of-3 2 What those claims proposed to do Kind Count Effect Unchanged 24 A recorded model fact re-read from its primary source and found to still hold. 17 reached their threshold; of those, 15 advanced a release verifiedAt to 2026-08-30 and 2 were withheld. The other 7 were refused by the provenance rubric. Change 21 A source page re-read this run, moving its lastCheckedDate to 2026-08-30. All 21 were accepted unanimously and all 21 were applied. Not covered
- 13 of the 33 creators scouted produced no claim at all. That is an honest zero per creator rather than a skipped stage, but it also means their records carry this run’s attention without carrying its verification date.
- Three creator hosts refused automated fetches this run — openai.com/news/ (403), ai.meta.com/blog/ (403) and x.ai/news (403). Approved alternatives were used and no claim rested on an unread page, but the announcement indexes themselves went unread.
- Breadth remains uncovered by design: OpenAI’s realtime, transcribe and tts models, roughly twenty Cohere Command models against the one recorded, Microsoft’s MAI-Code-1.1-Flash, MAI-Image-2.5, MAI-Voice-2 and MAI-Transcribe-1.5, and Amazon’s Nova Act, Nova Forge and Nova Multimodal Embeddings. Each is breadth work with its own sourcing burden, not re-verification.
- Microsoft MAI-Image-2.6, announced 2026-08-10, was deliberately not added. Its only availability statement names a product and a serving platform rather than the creator’s own release, so accessType has no creator-level statement and status is ambiguous between preview and current. It also has no recorded family. The ambiguity was left explicit rather than resolved by guess.
- xai grok-4.20 appears on docs.x.ai only in a passing note that logprobs are unsupported by "grok-4.20 and newer", never as a listed model. That is not enough to create a release.
What was evaluated
- Reviewers
- 60
- Verdicts cast
- 135
- Accepted by panel
- 38
- Rejected by panel
- 7
Deterministic gates and required checks — 7 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence.mjs20 claim bundles 0 Pass45 claims admissible with complete panels across 20 bundles. Thresholds were applied per profile — 2 for a pilot creator, 3 for a long-tail one — and no bundle omitted its policy. gate-source-approval.mjs20 claim bundles 0 PassZero proposed sources across all 20 bundles: every cited source was already in the dataset at anchor 7ca5802e, on one of the 41 origins approved there. No source was refused, and no origin was introduced. gate-dataset.mjsweb/src/data 0 Pass419 records validated across the documents raw.ts composes. npm run validateweb/ 0 Pass990 tests executed; astro check reported 0 errors. This is the gate that caught the two withheld verification dates, by refusing guidance older than the evidence beneath it. gate-scope.mjsbranch vs merge-base 0 Passchanged: 2, empty: false, outOfClass: []. Anchored on 7ca5802e, the merge base the gate computed rather than one supplied to it. ci-preflight.mjsnot requiredweb-ci, skills-ci, source-link-health-tests 0 PassRun locally before the branch was pushed; all three workflows predicted green. web-ciPR #597 — PassSUCCESS. GitHub performed the merge via --auto on this check going green; the agent did not merge. Posted 36 edits
36 edits across 2 documents, a net change of 0 records.
Dataset documents this run changed Document Before After What changed releases.json82 82 15 verifiedAt dates moved to 2026-08-30 — anthropic-claude-fable-5, anthropic-claude-opus-5 and anthropic-claude-sonnet-5 from 2026-08-26; ai2-olmo-2-7b, amazon-nova-2-sonic, cohere-command-a-plus-05-2026, deepseek-v4-pro, nvidia-nemotron-4-340b-base and xai-grok-4-6 from 2026-08-28; baidu-ernie-4-5-300b-a47b, bytedance-seed-oss-36b-instruct, databricks-dbrx-instruct, hugging-face-smollm3-3b, ibm-granite-4-2-30b and lg-ai-research-exaone-3-5-7-8b-instruct from 2026-08-29. No release was added or removed and no other field changed, so the record count is identical either side of the merge commit. sources.json198 198 21 lastCheckedDate values moved to 2026-08-30, one per page re-read this run. No source was added: the approved-source gate recorded zero proposed sources, so every date moved on a record that already existed. Sources are listed here rather than under the posted records below, which carry links and so hold only the collections the site routes. Each document links to the file as this run left it, not as it stands today.
Not posted 9 items
Rejected by the review panel
releases/01-ai-yi-1-5-34b-chat.verifiedAtRefused 2-of-3 under the long-tail unanimous bar. The Hugging Face card states "License: apache-2.0", but the recorded license object also asserts osiApproved, which is a claim about what OSI decided and needs an OSI source. The licence name cannot stand in for it.Blocked byprovenance rubricreleases/eleutherai-pythia-12b.verifiedAtRefused 2-of-3 on the same osiApproved grounds: the Pythia-12B card states the licence name and nothing about OSI approval status.Blocked byprovenance rubricreleases/snowflake-arctic-instruct.verifiedAtRefused 2-of-3 on the same osiApproved grounds; the Arctic card carries an apache-2.0 tag and no OSI provenance.Blocked byprovenance rubricreleases/sarvam-ai-sarvam-m-v1.verifiedAtRefused 2-of-3 on the same osiApproved grounds. Its source record still had its lastCheckedDate advanced, because that claim was about the page being read rather than about what the page proves.Blocked byprovenance rubricreleases/zhipu-ai-glm-4-5-air.verifiedAtRefused 2-of-3. The GLM-4.5-Air card states the models are released under the MIT open-source licence, which supports the licence name and still does not cite OSI approval for the osiApproved flag.Blocked byprovenance rubricreleases/moonshot-ai-kimi-k2-instruct.verifiedAtRefused 2-of-3. The card states "Context Length 128K" while the dataset records 131072. The quote does not force the binary 128x1024 reading, so the stored precision is not directly sourced — an assumed multiplier, not a stated number.Blocked byprovenance rubricreleases/upstage-solar-pro-preview-instruct.verifiedAtRefused 2-of-3 on the same binary-multiplier grounds: the page states "a maximum context length of 4K" and the dataset records 4096.Blocked byprovenance rubric
Verification date deliberately held back
releases/anthropic-claude-mythos-5.verifiedAtHeld at 2026-08-15 despite the claim reaching its threshold. validate.ts forbids guidance dated before the evidence beneath it, and fit-claude-mythos-5-limited-availability, verified 2026-08-15, rests on this release. Advancing the release would have stranded that fit statement behind its own evidence; npm run validate failed exactly this way when it was attempted. Re-dating guidance this run gathered no evidence for would have been the larger claim, so the write was withheld rather than the rule weakened.Blocked byfit-claude-mythos-5-limited-availabilityvalidate.tsreleases/anthropic-claude-haiku-4-5.verifiedAtHeld at 2026-08-15 for the same reason, against fit-claude-haiku-4-5-release-current. Two accepted claims applying nothing is the honest outcome here: the panel judged the model facts still true, and the dataset still refuses to say so until the guidance resting on them is re-verified too.Blocked byfit-claude-haiku-4-5-release-currentvalidate.ts
What this run does not prove
- This entry was transcribed after the run rather than by it. The run could not write its own line: refresh-runs.json is not one of the documents raw.ts composes, so gate-scope.mjs correctly reports it out of class, and including it would have cost the run its auto-merge. That gap is issue #419 and it has now recurred on every published run.
- pagesFetched is the summary issue’s approximate figure of "~45 pages" recorded as 45. Neither durable record states an exact count, so this number is the only one available rather than a counted one.
- reviewers and verdictsCast are derived, not quoted: 20 bundles x 3 rubrics = 60 reviewer invocations, and 45 claims x 3 rubrics = 135 verdicts, which reconciles with the published rubric tallies of 45, 45 and 45. Neither field has a verbatim source in the pull request body or the summary issue.
- The pull request body says "36 records, 38 changed lines". The merge commit shows 36 insertions and 36 deletions across the two documents, and the per-record diff reconciles to exactly 15 releases and 21 sources. Where the two disagree this entry follows the commit, as #419 already recommended after the same class of discrepancy on the 2026-08-27 run.
- A green run proves the recorded facts still match the pages this run read. It does not prove those pages are complete, that the 13 creators without claims are unchanged, or that anything on a host that refused the fetch is still current.
- No human reviewed any claim, verdict or edit in this run before it merged and deployed.
Follow-ups — proposed, not fixed
- osiApproved blocks routine licence re-verification for long-tail creators — five claims died on it this run. Either the recorded license objects need an OSI source attached once, or re-verification needs to target a licence sub-field a model card can actually source.
- Context windows recorded as exact integers cannot be re-verified against pages that state them as "128K" or "4K". Two claims died on the assumed 1024 multiplier.
- Three creator announcement indexes refuse automated fetches — openai.com/news/, ai.meta.com/blog/ and x.ai/news. The working alternatives this run used are worth recording in the reviewed catalogues so future runs do not rediscover them.
- Microsoft MAI-Image-2.6 needs a creator-level source before it can be recorded, following the precedent set for microsoft-mai-thinking-1.
2026-08-29-34e751
Data refresh 2026-08-29
Scope requested: Every creator in organizations.json — ai2, ai21-labs, amazon, anthropic, alibaba-cloud, cohere, deepseek, google-deepmind, meta, microsoft, mistral-ai, moonshot-ai, nvidia, openai, tii, xai, zhipu-ai — under the pilot 2-of-3 policy.Published3 edits posted · 4 items withheldAn agent-run, source-backed refresh under ADR 0003. All 17 creators in organizations.json were scouted; 32 pages were fetched and hashed across 18 creator slugs, and only two creators produced claims at all — Anthropic with three and OpenAI with one. Fifteen blind reviewers cast fifteen verdicts across three rubrics, eight deterministic checks passed, and three edits landed across two dataset documents before GitHub merged PR #529 on a green web-ci and Pages deployed successfully. The principal finding is small and was sitting in plain sight: the anthropic-claude-4-8 family existed with no release record and therefore rendered empty, so this run added the missing Claude Opus 4.8 release and the docs page it is sourced from. No human reviewed any part of it. This entry was transcribed a day late — see the caveats.
- PreflightRanAnchor a6a3dfc, selected by merge-base with refs/remotes/origin/main and computed by the gate rather than supplied — requestedBase was null. The approved-origin catalogue parsed cleanly at that anchor: 147 dataset sources and 12 profile catalogues out of 13 profile files, yielding 33 approved origins. The one profile listed but not drawn on is tools/updater/profiles/generic/long-tail.json, which configures no origins. The anchor is therefore full rather than silently narrowed.
- ScoutRan32 pages fetched and hashed across 18 creator slugs. Every quote was extracted by anchor from fetched page text, so none is a recollection. No page budget was exhausted. ai.meta.com/blog/ returned HTTP 400 to every request tried, including with a browser User-Agent and a cookie jar, so Meta scouting fell back to huggingface.co/meta-llama, an approved origin; no Meta claim rested on the unreachable page and no Meta record changed. The 15 creators outside Anthropic and OpenAI produced coverage observations rather than claims.
- ReviewRanThree independent reviewers per claim, launched in parallel across different model families, none seeing the scout's reasoning, another reviewer's verdict, or the running tally. The chair did not vote. One claim was reviewed twice and this is disclosed rather than hidden: round 1's provenance reviewer rejected the Opus 4.8 release because no attached quote stated the modalities, which was a mechanical defect — the bundle generator's block extractor matched an earlier occurrence of "Context window 1M tokens", so the "Input → output Text and images → text" line was never attached. The rubric's own remedy for a missing quote is to attach it, so the anchor was fixed, intendedUse was tightened to attribute rather than assert, and three fresh reviewers re-reviewed. Round 2's provenance reviewer rejected again on different, judgment-based grounds and the claim was not revised a third time: repeated revision to chase a verdict is vote-shopping. Fifteen verdicts were cast in total; the twelve from the final panels are published verbatim in the pull request body, and round 1's three rationales were not preserved.
- GatesRangate-evidence and gate-source-approval ran per creator before anything was applied; gate-dataset, npm run validate and gate-scope ran afterwards. No gate was skipped, no threshold lowered, no --force, no --admin, no --base override, and no direct push to main. gate-scope reported changed: 2, empty: false, outOfClass: [], so the change sat inside the ADR 0003 qualifying class and was eligible to auto-merge.
- PublishRanPR #529 carried the full evidence trail and was merged by GitHub via --auto once web-ci was green, not by the agent. Merge commit 0794286.
- DeployRanPages deploy run 33250588183 on main succeeded, and https://abdeslam-menacere.github.io/ModelTree/models/claude-opus-4-8/ returned HTTP 200 serving the new record. No revert needed.
What was found
- Scouts
- 17
- Pages fetched and hashed
- 32
- Claims proposed
- 4
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims anthropic pilot 2-of-3 3 openai pilot 2-of-3 1 What those claims proposed to do Kind Count Effect Add 2 One release and one source proposed; both survived every stage and were applied. Unchanged 2 Re-verified against a primary source; no value changed. One had its verification date advanced and one had it withheld. Not covered
- Fifteen of the seventeen creators scouted produced no claim at all. That is an honest zero per creator rather than a skipped stage, but it also means their records carry this run's attention without carrying its verification date.
- ai.meta.com/blog/ returns HTTP 400 to this client and was not read; huggingface.co/meta-llama was used instead.
- Google, Microsoft, Amazon, Cohere and OpenAI audio models remain uncovered by any claim this run could assemble.
- Mistral Medium 3.5 was deliberately not claimed. Its release date is ambiguous between a product post dated 2026-05-22 and a version string reading "v 26.04", and an ambiguity is a finding rather than a coin to toss.
What was evaluated
- Reviewers
- 15
- Verdicts cast
- 15
- Accepted by panel
- 4
- Rejected by panel
- 0
Deterministic gates and required checks — 8 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence.mjsanthropic 0 PassPolicy pilot, threshold 2. Three claims, of which two were applicable. gate-evidence.mjsopenai 0 PassPolicy pilot, threshold 2. One claim, none applicable — an unchanged claim applies nothing. gate-source-approval.mjsanthropic 0 PassFour cited sources were inherited and one was proposed: anthropic-opus-4-8-docs, on https://platform.claude.com — an origin already approved at the anchor commit and already carrying inherited sources, so no new origin was introduced. No source was refused. gate-source-approval.mjsopenai 0 PassThe one cited source, openai-gpt-5-6-sol-docs, was inherited from an already approved origin. No source was refused. gate-dataset.mjsweb/src/data 0 Pass148 sources, 26 publishers, 17 organizations, 37 families, 64 releases, 7 model-fit statements. npm run validateweb/ 0 Pass1613 tests executed across 71 files; astro check reported 0 errors. gate-scope.mjsbranch vs merge-base 0 Passchanged: 2, empty: false, outOfClass: []. Anchored on the merge-base the gate computed rather than one supplied to it. web-ciPR #529 — PassSUCCESS. GitHub performed the merge via --auto on this check going green; the agent did not merge. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
anthropic-claude-opus-4-8-release-addProvenanceaccessType "proprietary-hosted" is not forced by the quotes, because hosted availability does not exclude downloadable weights; and predecessorIds: [] under-states an announcement that calls Opus 4.7 "its predecessor". Outvoted 2-1. The claim was applied and both points are recorded here and in the pull request body as unresolved dissent rather than smoothed over.anthropic-claude-mythos-5-limited-release-unchangedProvenanceThe quotes do not uniquely force status "preview" over "current", and do not re-state the recorded intendedUse. Outvoted 2-1, and materially right in the event: the claim reached its threshold and still applied nothing, because the verification date it would have advanced was withheld on separate grounds.
Posted 3 edits
3 edits across 2 documents, a net change of 2 records.
Dataset documents this run changed Document Before After What changed releases.json63 64 One release added: anthropic-claude-opus-4-8. openai-gpt-5-6-sol had its verifiedAt advanced to 2026-08-29. sources.json147 148 One source added: anthropic-opus-4-8-docs — the platform.claude.com model overview page the Opus 4.8 release is drawn from, accepted 3-of-3 on an origin already approved at the anchor commit. lastCheckedDate moved to 2026-08-29 on three of the four inherited sources re-read this run; anthropic-fable-5-mythos-5-docs was held at 2026-08-26 with the Mythos 5 claim it belongs to. Sources are listed here rather than under the posted records below, which carry links and so hold only the collections the site routes. Each document links to the file as this run left it, not as it stands today.
Records added
- Claude Opus 4.8passportreleases
anthropic-claude-opus-4-8New release — released 2026-05-28, status legacy, 1M context and 128K max output, text and image in, text out, API alias claude-opus-4-8. Fills the anthropic-claude-4-8 family, which until now had no release and rendered empty. Carries the recorded provenance dissent on accessType and on its empty lineage arrays.
Not posted 4 items
Verification date deliberately held back
releases/anthropic-claude-mythos-5.verifiedAtHeld at 2026-08-15; nothing was applied for this claim despite it reaching its threshold. validate.ts forbids guidance dated before the evidence beneath it, and fit-claude-mythos-5-limited-availability, verified 2026-08-15, rests on this release's status and intendedUse. Advancing the release would have silently invalidated that fit statement — validation failed exactly this way when it was first attempted. Re-dating guidance this run gathered no evidence for would have been the larger claim, so the date was withheld instead. This also matches how gate-evidence treats unchanged claims, which apply nothing.Blocked byfit-claude-mythos-5-limited-availabilityvalidate.ts
Sources conflict, so no value changed
mistral-medium-3-5Not claimed at all. Its release date is ambiguous between a product post dated 2026-05-22 and a version string reading "v 26.04". Neither reading is forced by the other, so the run recorded the disagreement rather than picking a side. The dataset still holds nothing on this model, which is the cost of leaving the ambiguity explicit.
Out of the run’s reach
meta-announcement-indexhttps://ai.meta.com/blog/ returns HTTP 400 to every request tried, including with a browser User-Agent and a cookie jar. Meta scouting fell back to huggingface.co/meta-llama, an approved origin. No Meta claim rested on the unreachable page and no Meta record changed, so absence of a Meta finding here is not evidence that there was nothing to find.Blocked byhttps://ai.meta.com/blog/coverage-google-microsoft-amazon-cohere-openai-audioGoogle, Microsoft, Amazon, Cohere and OpenAI audio models were scouted and produced coverage observations rather than claims. The observations are carried as follow-ups below rather than being written into the dataset as inferences.
What this run does not prove
- No human reviewed the pull request. The evidence trail, the deterministic gates and the required web-ci check are the whole of the oversight.
- A green run proves the dataset is internally coherent and that every applied claim carried a quote from an approved origin. It does not prove the sources are right.
- Two creators out of seventeen produced claims. The other fifteen were read and yielded nothing claimable, so this run moved the dataset barely at all and its silence about them is not a clean bill of health.
- Round 1's three review rationales for the Opus 4.8 release claim were not preserved verbatim in the run artifacts, so the audit trail for that claim starts at round 2. That gap is follow-up 6 and it is the reason the reviewer count here exceeds the twelve rationales published.
- editsApplied counts the three claims that reached the dataset — two adds and one advanced verification date — following the convention of the earlier entries here. The merge commit also moves lastCheckedDate on three sources that rode along with those claims, which the diff shows and this count does not.
- This entry was written on 2026-08-30, a day after the run, because the run could not record itself: refresh-runs.json sits outside the ADR 0003 qualifying class that gate-scope enforces, so including it in PR #529 would have failed that gate and forfeited auto-merge. The transcription needs its own human-merged pull request, and for this run that pull request was initially never opened. The absence was noticed only when a reader asked why the page showed no 2026-08-29 line.
Follow-ups — proposed, not fixed
- Re-verify fit-claude-mythos-5-limited-availability so the Mythos 5 release verifiedAt can move. It is blocking a routine re-verification today.
- Populate predecessorIds and successorIds across the Claude Opus line. Opus 4.6, 4.7 and now 4.8 all carry empty lineage arrays while the announcements state the relationships explicitly. Fix the line as a whole rather than one record.
- Resolve the Mistral Medium 3.5 release-date ambiguity — decide which of the two published dates the dataset should record, or record the ambiguity explicitly.
- Close coverage gaps for Google, Microsoft, Amazon, Cohere, and OpenAI audio models.
- Find a reachable Meta announcement index to replace ai.meta.com/blog.
- Persist superseded review rounds verbatim in the run artifacts, so a corrected claim's first-round rationales survive for audit.
- Nothing notices when a published run never gets transcribed into this log. The publishing skill set opens the dataset pull request and files the summary issue, and the log entry is a separate human-merged change that no check requires; this run's line was missing for a day and only a reader's question surfaced it.
2026-08-27-ab1644
Data refresh 2026-08-27
Scope requested: Every creator in organizations.json, plus the long-tail profile sweepPublished35 edits posted · 56 items withheldAn agent-run, source-backed refresh under ADR 0003. Five scouts fetched and hashed 130 primary pages and proposed 104 claims; twelve blind reviewers cast 312 verdicts across three rubrics; six deterministic gates passed; 35 edits landed across four dataset documents and GitHub Pages deployed successfully. 51 claims were dropped, most of them to one scout-side fault that proposed new sources without the paired citation edit. No human reviewed any part of it.
- PreflightRanClean tree, gh authenticated, no unmerged pull request from a previous refresh.
- ScoutRanFive bundles, one per creator profile plus a long-tail sweep. 130 page snapshots hashed. openai.com/news/ returned 403 and the RSS feed was used instead; ai.meta.com/blog/ returned 400; docs.anthropic.com now redirects to platform.claude.com.
- ReviewRanTwelve reviewers, three rubrics over four bundles. The long-tail bundle had no claims to review. No threshold was lowered and no reviewer was re-run.
- GatesRangate-evidence, gate-source-approval and check-bundle-pairing before applying anything; gate-dataset, npm run validate, and gate-scope after. The coupled drop set was iterated to a fixpoint over 3 passes.
- PublishRanPR #417 opened with the full evidence trail across the body and two comments, and merged by GitHub on a green web-ci.
- DeployRanPages deployed main @ 547691a successfully; the live site serves the new entities. No revert needed.
What was found
- Scouts
- 5
- Pages fetched and hashed
- 130
- Claims proposed
- 104
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims openai pilot 2-of-3 26 anthropic pilot 2-of-3 24 google-deepmind pilot 2-of-3 22 meta pilot 2-of-3 32 long-tail sweep long-tail 3-of-3 0 What those claims proposed to do Kind Count Effect Add 74 Proposed new records; 30 survived every stage. Change 8 Field corrections; 5 survived every stage. Unchanged 20 Re-verified against a primary source; no value changed. 11 records had verifiedAt moved forward. Conflict 2 Primary sources disagree; both sides recorded, no value changed. Neither reached the dataset. Not covered
- Any creator outside the four pilot profiles and the long-tail catalogue.
- openai.com/news/ returns 403 to this client; the RSS feed was read instead, so anything on the news pages but absent from the feed was not seen.
- ai.meta.com/blog/ returns 400 to this client and was not read.
- The OpenAI page budget was exhausted at 40 pages, so its catalogue was not read to the end.
- The long-tail sweep returned zero claims. That is an honest zero, not a skipped stage — it was gated like every other bundle.
What was evaluated
- Reviewers
- 12
- Verdicts cast
- 312
- Accepted by panel
- 83
- Rejected by panel
- 21
Deterministic gates and required checks — 8 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence5 bundles 0 PassThree rubrics present per claim, no duplicate votes, no empty rationale, every claim at its profile threshold. gate-source-approval5 bundles 0 Pass76 citations against anchor 052fb0d, selected by merge-base with refs/remotes/origin/main. requestedBase was null, so the run could not approve its own committed source. 16 approved origins, anchored on 61 sources already in the dataset and 5 reviewed profile catalogues. No new host admitted, and no source added to make a citation resolve. check-bundle-pairingnot required5 bundles 0 PassLanded on main as #413 mid-run and was run from main copy rather than added to this branch, since adding it would be an out-of-class change gate-scope would rightly refuse. It named the same unpaired sources this run had already found independently. gate-datasetworking tree 0 PassThe dataset is internally coherent after the edits were applied: 80 sources, 19 families, 35 releases. gate-scopebranch vs anchor 0 Passchanged: 4, outOfClass: [], empty: false. Exit 0 alone is ambiguous; the publish precondition is exit 0 and changed > 0. npm run validatebaseline and after 0 Pass579 tests across 24 files, astro check 0 errors. web-ciPR #417 — PassGreen. This is the check branch protection requires before a merge. skills-cinot requiredPR #417 — PassGreen, unlike the 2026-08-25 run where it was red and non-blocking. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
openai-gpt-5-1-family-addProvenanceThe RSS quote states GPT-5.1 was introduced for developers on 2025-11-13 and the docs quote states coding/agentic use and token limits, but no quote states the proposed family status "legacy" or the multimodal category. The proposed record therefore states more than the evidence quotes do.openai-gpt-5-2-family-addProvenanceThe RSS quote states "Introducing GPT-5.2" with a 2025-12-11 pubDate, and the docs quote calls it a previous frontier model. No quote states the proposed family status "legacy" as a status value, so the record goes beyond the quoted evidence.openai-gpt-5-3-codex-family-addProvenanceThe RSS quote gives the GPT-5.3-Codex introduction and date, and the docs quote says it is optimized for agentic coding. No quote states the proposed family status "current", so the proposed record adds an unstated field.openai-gpt-image-family-addProvenanceThe dated announcement quote names "ChatGPT Images 2.0", not GPT-Image or GPT-Image-2. The GPT Image 2 docs quote states image generation/editing but not the 2026-04-21 introduction date for that entity.anthropic-claude-4-6-familyProvenanceThe quotes name Claude Opus 4.6 and Claude Sonnet 4.6 and describe their capabilities. They do not state a "Claude 4.6" generation family, the 2026-02-05 firstReleaseDate for that family, or its status.anthropic-claude-4-7-familyProvenanceThe quote states Opus 4.7 improves on Opus 4.6 in advanced software engineering. It does not state a broader "Claude 4.7" family, the proposed firstReleaseDate, or the status field.anthropic-claude-4-8-familyProvenanceThe quote states Anthropic is upgrading Claude Opus to Claude Opus 4.8. It does not state a broader "Claude 4.8" generation family, its firstReleaseDate, or its status.meta-llama-4-maverick-original-sourceProvenanceThe quote says the repository contains original checkpoints for Meta's llama-stack codebase. It does not name meta-llama/Llama-4-Maverick-17B-128E-Instruct-Original, so the source title is not stated by the quote alone.meta-llama-3-2-1b-releaseProvenanceThe quotes state Llama 3.2 1B details, 1.23B parameters, 128k context, release date, and the Llama 3.2 license title. No quote states the proposed "legacy" status or the downloadable/OSI license fields.meta-llama-3-2-3b-releaseProvenanceThe quotes state Llama 3.2 3B details, 3.21B parameters, release date, and the Llama 3.2 license title. No quote states the proposed 128,000-token context window for this 3B claim or the "legacy" status.meta-llama-3-2-11b-vision-releaseProvenanceThe quotes state Llama 3.2 Vision 11B details, 10.6B parameters, 128k context, release date, and the license title. No quote states the proposed "legacy" status or the downloadable/OSI license fields.meta-llama-3-2-90b-vision-releaseProvenanceThe quotes state Llama 3.2 Vision 90B details, 88.8B parameters, 128k context, release date, and the license title. No quote states the proposed "legacy" status or the downloadable/OSI license fields.meta-llama-4-maverick-api-aliasesProvenanceThe first quote names meta-llama/Llama-4-Maverick-17B-128E, but the second only says "original checkpoints" without naming the exact Instruct-Original repository. Read alone, the quotes do not state both proposed aliases.meta-llama-4-maverick-source-idsProvenanceThe first quote identifies the base Maverick repository, but the second only says "original checkpoints" without naming the exact original-checkpoint model card. The evidence therefore does not state both newly proposed source IDs.
Posted 35 edits
35 edits across 4 documents, a net change of 30 records.
Dataset documents this run changed Document Before After What changed sources.json61 80 19 sources added. families.json12 19 Seven families added. releases.json31 35 Four releases added. organizations.json4 4 Verification dates moved forward; no record added or removed. Each document links to the file as this run left it, not as it stands today.
Records added
- GPT-5.1in the treefamilies
openai-gpt-5-1New OpenAI family; carries a recorded provenance dissent. - GPT-5.2in the treefamilies
openai-gpt-5-2New OpenAI family; carries a recorded provenance dissent. - GPT-5.3-Codexin the treefamilies
openai-gpt-5-3-codexNew OpenAI family; carries a recorded provenance dissent. - GPT-Imagein the treefamilies
openai-gpt-imageNew OpenAI family; carries a recorded provenance dissent. - Claude 4.6in the treefamilies
anthropic-claude-4-6New Anthropic generation family; carries a recorded provenance dissent. - Claude 4.7in the treefamilies
anthropic-claude-4-7New Anthropic generation family; carries a recorded provenance dissent. - Claude 4.8in the treefamilies
anthropic-claude-4-8New Anthropic generation family; carries a recorded provenance dissent. - Llama 3.2 1Bpassportreleases
meta-llama-3-2-1bNew release; carries a recorded provenance dissent on its status and license fields. - Llama 3.2 3Bpassportreleases
meta-llama-3-2-3bNew release; carries a recorded provenance dissent on its context window and status. - Llama 3.2 11B Visionpassportreleases
meta-llama-3-2-11b-visionNew release; carries a recorded provenance dissent on its status and license fields. - Llama 3.2 90B Visionpassportreleases
meta-llama-3-2-90b-visionNew release; carries a recorded provenance dissent on its status and license fields.
Not posted 56 items
Rejected by the review panel
openai-gpt-5-1-release-addOpenAI's API docs list GPT-5.1 as an API model for coding and agentic tasks. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)openai-gpt-5-2-release-addOpenAI's API docs list GPT-5.2 as an API model for professional work. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)openai-gpt-5-3-codex-release-addOpenAI's API docs list GPT-5.3-Codex as an agentic coding model. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)openai-gpt-5-4-mini-release-addOpenAI's API docs list GPT-5.4 Mini as a faster GPT-5.4 variant for high-volume workloads. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)openai-gpt-5-4-nano-release-addOpenAI's API docs list GPT-5.4 nano as a GPT-5.4-class model for simple high-volume tasks. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)openai-gpt-image-2-release-addOpenAI's API docs list GPT-Image-2 as an image-generation model. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-sonnet-4-5-releaseClaude Sonnet 4.5 should be added as an Anthropic Claude 4.5 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-opus-4-5-releaseClaude Opus 4.5 should be added as an Anthropic Claude 4.5 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-opus-4-6-releaseClaude Opus 4.6 should be added as an Anthropic Claude 4.6 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-sonnet-4-6-releaseClaude Sonnet 4.6 should be added as an Anthropic Claude 4.6 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-opus-4-7-releaseClaude Opus 4.7 should be added as an Anthropic Claude 4.7 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)anthropic-claude-opus-4-8-releaseClaude Opus 4.8 should be added as an Anthropic Claude 4.8 release. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-2-5-flash-lite-add-releaseGemini 2.5 Flash-Lite should be added as a current Gemini 2.5 release dated July 22, 2025. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-flash-preview-add-releaseGemini 3 Flash Preview should be added as a preview Gemini 3 release dated December 17, 2025. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-1-flash-image-add-releaseGemini 3.1 Flash Image should be added as a current Gemini 3 release dated May 28, 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-pro-image-add-releaseGemini 3 Pro Image should be added as a current Gemini 3 release dated May 28, 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-2-5-flash-image-add-releaseGemini 2.5 Flash Image should be added as a current Gemini 2.5 release dated October 2, 2025. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-5-transcribe-add-releaseGemini 3.5 Transcribe should be added as a current Gemini 3 release dated to August 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-5-live-translate-preview-add-releaseGemini 3.5 Live Translate Preview should be added as a preview Gemini 3 release dated to June 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-1-flash-live-preview-add-releaseGemini 3.1 Flash Live Preview should be added as a preview Gemini 3 release dated March 11, 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)google-gemini-3-1-flash-tts-preview-add-releaseGemini 3.1 Flash TTS Preview should be added as a preview Gemini 3 release dated April 13, 2026. Did not reach 2 of 3 accepts under the pilot policy.Blocked bypilot policy (2-of-3)
Accepted by the panel, then dropped
source-openai-api-docsModels | OpenAI API should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-openai-gpt-5-4-mini-docsGPT-5.4 Mini model documentation should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-openai-gpt-5-4-nano-docsGPT-5.4 nano model documentation should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-openai-daybreak-red-docsDaybreak Red model documentation should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-openai-daybreak-blue-docsDaybreak Blue model documentation should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-openai-gpt-realtime-mini-docsGPT-Realtime Mini model documentation should be added as an OpenAI official-docs source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsanthropic-fable-5-docs-sourceClaude Fable 5 should be added as an official Anthropic source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsanthropic-opus-5-docs-sourceClaude Opus 5 should be added as an official Anthropic source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsanthropic-haiku-4-5-docs-sourceClaude Haiku 4.5 should be added as an official Anthropic source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsanthropic-fable-5-biology-safeguards-sourceImproving Fable 5's biology safeguards should be added as an official Anthropic source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsanthropic-opus-4-5-announcement-sourceIntroducing Claude Opus 4.5 should be added as an official Anthropic source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-2-5-flash-lite-docsAdd Gemini 2.5 Flash-Lite as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-flash-preview-docsAdd Gemini 3 Flash preview as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-1-flash-image-docsAdd Gemini 3.1 Flash image as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-pro-image-docsAdd Gemini 3 Pro image as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-2-5-flash-image-docsAdd Gemini 2.5 Flash image (Nano Banana) as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-5-transcribe-docsAdd Gemini 3.5 Transcribe as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-5-live-translate-docsAdd Gemini 3.5 Live translate as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-1-flash-live-docsAdd Gemini 3.1 Flash live preview as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tssource-google-gemini-3-1-flash-tts-docsAdd Gemini 3.1 Flash TTS preview as an official Google AI source. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsmeta-llama-guard-4-sourceThe source catalog should include meta-llama/Llama-Guard-4-12B. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsmeta-llama-prompt-guard-2-86m-sourceThe source catalog should include meta-llama/Llama-Prompt-Guard-2-86M. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.tsmeta-llama-prompt-guard-2-22m-sourceThe source catalog should include meta-llama/Llama-Prompt-Guard-2-22M. Accepted by the panel, then dropped: validate.test.ts refuses a source no record cites, and the scout proposed this source without the paired change claim that would wire it into a record sourceIds.Blocked byvalidate.test.ts
Verification date deliberately held back
releases/openai-gpt-5.verifiedAtModel fit statement fit-gpt-5-superseded rests on this release and was verified 2026-08-18. Moving verifiedAt to 2026-08-27 would make that guidance stale, and this run gathered no evidence about fit statements. Capping the date to fit the invariant would assert a verification that never happened.Blocked byfit-gpt-5-supersededvalidate.tsreleases/anthropic-claude-mythos-5.verifiedAtModel fit statement fit-claude-mythos-5-limited-availability rests on this release and was verified 2026-08-15. Moving verifiedAt to 2026-08-27 would make that guidance stale, and this run gathered no evidence about fit statements. Capping the date to fit the invariant would assert a verification that never happened.Blocked byfit-claude-mythos-5-limited-availabilityvalidate.tsreleases/meta-llama-4-scout.verifiedAtModel fit statement fit-llama-4-scout-self-hosting rests on this release and was verified 2026-08-15. Moving verifiedAt to 2026-08-27 would make that guidance stale, and this run gathered no evidence about fit statements. Capping the date to fit the invariant would assert a verification that never happened.Blocked byfit-llama-4-scout-self-hostingvalidate.ts
Sources conflict, so no value changed
google-gemini-embedding-2-api-code-conflictGoogle sources disagree on whether the Gemini Embedding 2 API identifier is gemini-embedding-2 or gemini-embedding-2-preview. All three rubrics accepted it as a conflict; a conflict is recorded as a finding and never applied to the dataset.Blocked bygoogle-gemini-models-referencegoogle-gemini-deprecations
Source refused by the approval gate
openai-gpt-5-6-cyber-daybreak-red-aliasOpenAI's Daybreak Red model documentation lists daybreak-red-latest as an alias for gpt-5.6-cyber. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsopenai-gpt-5-6-sol-daybreak-blue-aliasOpenAI's Daybreak Blue model documentation lists daybreak-blue-latest as an alias for gpt-5.6-sol. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsopenai-gpt-realtime-mini-status-conflictOpenAI's model catalog labels GPT-Realtime Mini deprecated, while its individual model page labels it default. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsanthropic-claude-fable-5-intended-use-biology-caveatClaude Fable 5's intended-use text should include Anthropic's caveat that dual-use biology requests still fall back to Opus 5. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsanthropic-claude-haiku-4-5-status-activeClaude Haiku 4.5 remains an active latest Claude model. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsanthropic-claude-haiku-4-5-context-windowClaude Haiku 4.5 still has a 200,000-token context window. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjsanthropic-claude-haiku-4-5-api-aliasesClaude Haiku 4.5 still uses claude-haiku-4-5 as a convenience alias for claude-haiku-4-5-20251001. Rests on a source withdrawn earlier in this same run, so gate-source-approval refused it. Writing the missing citation by hand would have been an unreviewed dataset edit, so the coupled set was withheld together.Blocked bygate-source-approval.mjs
Blocked by policy before it could run
long-tail-creatorsQwen, zai-org, MiniMaxAI, deepseek-ai, moonshotai, sensenova, Lightricks and ornith-ai are trending on an approved origin, but organizationSchema requires website and releasePage URLs that no approved-origin page states, and each creator own host is not an approved origin. Approving a new origin is a human decision, not one a refresh run may take for itself.Blocked byorganizationSchemaapproved-origin list
What this run does not prove
- No human reviewed the pull request. The evidence trail, the deterministic gates and the required web-ci check are the whole of the oversight.
- A green run proves the dataset is internally coherent and every applied claim carried a quote from an approved origin. It does not prove the sources are right.
- Half the proposed claims did not survive. The dataset is more complete than it was, not complete.
- Two page budgets were exhausted and two hosts refused this client, so absence of a record here is not evidence that the model does not exist.
- The per-creator applied counts in summary issue #418 and its gate-dataset source count disagree with the pull request body. The figures here follow the pull request body and the merge commit itself, which agree with each other.
Follow-ups — proposed, not fixed
- Scouts are not yet running check-bundle-pairing, which landed on main mid-run and would have caught this run largest failure class at bundle time.
- The GPT-Realtime Mini status conflict was unanimously accepted and still could not be published, because both its sources were unpaired. The disagreement is real and remains unrecorded in the dataset.
- Long-tail creators have no route into the dataset under the current approved origins, and approving an origin is a human decision.
- The reviewed profiles and the dataset disagree about canonical release names, which drove most editorial rejections and will recur every run until it is settled.
- Three model fit statements are older than the release facts they rest on and block those releases from carrying this run verification date.
2026-08-26-d5f729
Data refresh 2026-08-26
Scope requested: Every creator in organizations.json, plus the long-tail profile sweepPublished63 edits posted · 9 items withheldAn agent-run, source-backed refresh under ADR 0003. Five scouts fetched and hashed 55 primary pages and proposed 71 claims; twelve blind reviewers cast 213 verdicts across three rubrics; four deterministic gates passed; 63 edits landed across five dataset documents and GitHub Pages deployed successfully. No human reviewed any part of it.
- PreflightRanClean tree, gh authenticated, no unmerged pull request from a previous refresh.
- ScoutRanFive bundles, one per creator profile plus a long-tail sweep. 55 page snapshots hashed.
- ReviewRanTwelve reviewers, three rubrics over four bundles. The long-tail bundle had no claims to review.
- GatesRangate-evidence and gate-source-approval before applying anything; gate-dataset, npm run validate, and gate-scope after.
- PublishRanPR #302 opened with the full evidence trail and merged by GitHub on a green web-ci.
- DeployRanPages deployed main @ 0e9867b successfully; the live site serves the new records. No revert needed.
What was found
- Scouts
- 5
- Pages fetched and hashed
- 55
- Claims proposed
- 71
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims openai pilot 2-of-3 14 anthropic pilot 2-of-3 18 google-deepmind pilot 2-of-3 19 meta pilot 2-of-3 20 long-tail sweep long-tail 3-of-3 0 What those claims proposed to do Kind Count Effect Add 25 New records. Change 15 Field corrections and counter updates. Unchanged 29 Re-verified against a primary source; no value changed. Conflict 2 Primary sources disagree; both sides recorded, no value changed. Not covered
- Any creator outside the four pilot profiles and the long-tail catalogue.
- www.llama.com content, deferred because the host now redirects to an unapproved origin.
- The long-tail sweep returned zero claims. That is an honest zero, not a skipped stage — it was gated like every other bundle.
What was evaluated
- Reviewers
- 12
- Verdicts cast
- 213
- Accepted by panel
- 71
- Rejected by panel
- 0
Deterministic gates and required checks — 6 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence5 bundles 0 PassThree rubrics present per claim, no duplicate votes, no empty rationale, every claim at its profile threshold. gate-source-approval5 bundles 0 Pass134 citations against anchor 13276b0, selected by merge-base with refs/remotes/origin/main. 20 inherited sources, 17 proposed, 16 approved origins, no new host admitted. gate-datasetworking tree 0 PassThe dataset is internally coherent after the edits were applied. gate-scopebranch vs anchor 0 Passchanged: 5, outOfClass: [], empty: false. Exit 0 alone is ambiguous; the publish precondition is exit 0 and changed > 0. npm run validatebaseline and after 0 Pass400/400 tests, 0 errors both before any change and on the branch. The baseline was captured first so a pre-existing red main could not be mistaken for damage from this run. web-ciPR #302 — Pass400/400 tests, 0 errors. This is the check branch protection requires before a merge. skills-cinot requiredPR #302 — FailRed, and not a required check, so GitHub merged on web-ci alone before it could be acted on. Filed as #304: gates.test.mjs pins TODAY to 2026-08-25 but asserts it against the live dataset, so any same-day refresh fails it. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
openai-add-gpt-5-6-cyber-releaseEditorial and entity boundariesreleaseDate was taken from an RSS pubDate, which dates the announcement item rather than the release.anthropic-unchanged-source-opus-5-announcement-checkedProvenanceThe quote supports the API identifier but not the pricing half of the same claim.meta-muse-video-releaseEditorial and entity boundariesA model described as "coming soon" was given a concrete day-precision releaseDate.
Posted 63 edits
63 edits across 5 documents, a net change of 24 records.
Dataset documents this run changed Document Before After What changed sources.json47 61 14 sources added. families.json11 12 One family added: meta-muse. releases.json22 31 Nine releases added. organizations.json4 4 One field corrected; no record added or removed. usage-observations.json2 2 Both readings updated; no record added or removed. Each document links to the file as this run left it, not as it stands today.
Records added
- Musein the treefamilies
meta-museNew Meta family. - GPT-5.6 Cyberpassportreleases
openai-gpt-5-6-cyberNew release; carries a recorded editorial dissent on its release date. - Claude Sonnet 5passportreleases
anthropic-claude-sonnet-5New release. - Gemini 3.5 Flashpassportreleases
google-gemini-3-5-flashNew release. - Gemini 3.6 Flashpassportreleases
google-gemini-3-6-flashNew release. - Gemini 3.7 Flashpassportreleases
google-gemini-3-7-flashNew release. - Muse Sparkpassportreleases
meta-muse-sparkNew release. - Muse Spark 1.1passportreleases
meta-muse-spark-1-1New release. - Muse Imagepassportreleases
meta-muse-imageNew release. - Muse Videopassportreleases
meta-muse-videoNew release; carries a recorded editorial dissent on its release date.
Not posted 9 items
Accepted by the panel, then dropped
openai-add-source-deprecationsAccepted by the panel, then dropped by the chair. validate.test.ts enforces that an unreferenced source is dead provenance, and no reviewed claim wired openai-deprecations into any record's sourceIds. Adding that edit would have meant authoring dataset content no reviewer saw.Blocked byvalidate.test.tsopenai-change-gpt-5-status-deprecatedDropped with its source. Keeping the status change alone records "deprecated" against sourceIds that do not state it — every supporting quote comes from openai-deprecations.Blocked byopenai-add-source-deprecationsopenai-change-gpt-4-1-nano-status-deprecatedDropped for the same reason as the GPT-5 status change: its only supporting quote comes from the source that could not be added.Blocked byopenai-add-source-deprecations
Verification date deliberately held back
releases/anthropic-claude-haiku-4-5The field was re-read and unchanged, but validate.ts requires a model-fit statement to be at least as fresh as the facts it reads. Bumping verifiedAt would strand a dependent statement, so it stays at 2026-08-15. Under-claiming a verification is honest; over-claiming is not.Blocked byfit-claude-haiku-4-5-release-currentfit-claude-haiku-4-5-family-legacyreleases/anthropic-claude-mythos-5Held at 2026-08-15 for the same freshness rule: nobody re-derived the dependent fit statement over the re-read field.Blocked byfit-claude-mythos-5-limited-availabilityfamilies/meta-llama-4Held at 2026-08-15 by a test-suite coupling rather than a data fact: model-fit.test.ts injects fixtures pinned at 2026-08-15 over the real seed records. Tests are not dataset documents, so this run may not edit them.Blocked bymodel-fit.test.ts
Sources conflict, so no value changed
families/anthropic-claude-4-5.statusTwo Anthropic pages disagree and neither changed a value. The models overview lists Claude Sonnet 4.5 under "Legacy models (still available)"; the deprecations page lists claude-sonnet-4-5-20250929 as Active, retirement not sooner than 29 September 2026. Conflicting data stays explicit rather than being smoothed over.releases/google-gemini-3-5-flash-lite.statusThe Gemini 3 developer guide states "All Gemini 3 models are currently in preview"; the platform docs give gemini-3.5-flash-lite launch stage GA, released 21 July 2026. Recorded with both sides quoted; no value changed.
Source refused by the approval gate
www.llama.comThe host now redirects to developer.meta.com, an origin neither the committed dataset nor any profile catalogue stands behind. The Meta scout deferred it rather than citing it, so the Muse claims rest on ai.meta.com instead. Admitting a new host is a human decision.Blocked bygate-source-approval
What this run does not prove
- The gates catch malformed, impossible, unreferenced, and boundary-violating data. They do not catch a well-formed claim that is simply wrong.
- Per ADR 0005 the evidence gate checks the form of a citation, never its remote content. A gate pass attests that the paperwork is complete, not that the page still says what it said.
- The three reviewers share a failure mode — three instances of one model family reading the same page — so a source that is itself wrong can carry all three.
- Auto-merge was armed correctly, but because skills-ci is not a required check GitHub merged on web-ci alone while skills-ci was red. No gate was skipped and no threshold lowered; it does mean an unattended run cannot currently stop on a non-required check going red.
Follow-ups — proposed, not fixed
- A scout proposing a status change must also propose the sourceIds change that carries its source. This cost three panel-accepted OpenAI claims.
- Fit-statement freshness blocks legitimate verification bumps on anthropic-claude-haiku-4-5, anthropic-claude-mythos-5, and meta-llama-4.
- www.llama.com now redirects to developer.meta.com, a new host that needs a human decision on the approved-origin list.
- Consider making skills-ci a required check, so a red non-required job can stop an unattended merge. Requiring a check is a branch-protection setting, so it is an owner action rather than a change this program can make.