What each refresh found, and what it did not publish.
ModelTree's dataset is refreshed by agents against primary sources, reviewed by an independent three-rubric panel, and gated deterministically. Every run is recorded here in full — including the runs that published nothing.
A run's working state is never committed. This page transcribes the durable record: the pull request body and the summary issue, both linked from each entry.
- Runs recorded
- 23
- Pages fetched
- 2,128
- Claims proposed
- 1,352
- Edits published
- 611
- Items withheld
- 439
Showing 1–10 of 16 runsfiltered by Published
Clear filter2026-09-13-bfee2c
Gemini 3.8 Flash clears the panel and is refused by the home byte ceiling for the second run running
Scope requested: Every creator in organizations.json plus the long-tail profile: 44 scouted, none unswept. 92 pages attempted across 2 fetch rounds, 88 returning content.Published1 edit posted · 4 items withheldA full sweep of all 44 creators. Only google-deepmind yielded claims: Gemini 3.8 Flash, GA on 2026-09-02, with its Gemini API model page and its Vertex AI model page as sources, plus a featured-rationale restate on Gemini 3.7 Flash whose current wording calls it the newest Flash tier. The three-rubric panel accepted all five claims and gate-evidence admitted the bundle at exit 0. The deterministic asset budget then refused the tranche: the home route measures 1,100,526 raw critical bytes on clean HEAD against a criticalMaxRaw of 1,105,000, and the three new records take it to 1,110,273 - over by 5,273. A gate cannot be outvoted by a panel majority, so the four records that carry bytes were withheld and only the zero-byte verification advance on the Gemini deprecations page was published. This is the second consecutive run to reach the same wall: 2026-09-11-cd4874 withheld the same records for the same reason.
- PreflightRanClean tree, gh authenticated, no refresh pull request from a previous run still open. Drydock probed both forms: the bare name is refused by the execution policy and drydock.cmd reports 0.1.0, so the CLI is installed and blocked rather than absent. npm probed the same way: bare refused, npm.cmd reports 11.9.0.
- ScoutRan92 pages attempted, 88 ok, 4 failed. Every quote is byte-verbatim from a fetched page, verified mechanically against the saved body by the bundle builder, which throws rather than emits when a quote is absent. Every page is sha256-hashed on the day. No search snippet was used as evidence. One bundle carried claims; 43 creators were scouted and yielded nothing.
- ReviewRanThree sub-agents, one per rubric, on three different models, run in parallel and blind to each other, to the scout reasoning and to the running tally; each saw only the claims, their evidence and the dataset slice they touch. 15 verdicts over 5 claims. The chair neither voted nor broke a tie, and no reviewer was re-run.
- GatesRangate-source-approval and gate-evidence ran pre-apply on a clean tree. gate-evidence passed at exit 0 and was controlled in the same session against two mutated bundles - one taken sub-threshold, one with retrieval set to search - which it refused at exit 1 apiece, so the pass is a reading rather than a blind instrument. gate-dataset, gate-reversals, gate-scope and gate-ledger passed on the applied tranche. npm run validate refused the full tranche on the asset budget, which is what withheld it.
- PublishRanTwo commits on one branch and one pull request: the dataset change first, then this ledger entry, which cannot be written until the pull request it must name exists. Merged by GitHub under --auto --squash once web-ci was green. No --admin, no --force, no skipped gate and no direct push to main.
- DeployNot runRecorded before the merge, so the Pages deploy had not yet run when this entry was written. Confirmed separately after merge; a failed deploy is reverted by pull request rather than left to freeze the published site.
What was found
- Scouts
- 1
- Pages fetched and hashed
- 92
- Claims proposed
- 5
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims openai pilot 2-of-3 0 anthropic pilot 2-of-3 0 google-deepmind pilot 2-of-3 5 meta pilot 2-of-3 0 xai long-tail 3-of-3 0 mistral-ai long-tail 3-of-3 0 deepseek long-tail 3-of-3 0 alibaba-cloud pilot 2-of-3 0 microsoft pilot 2-of-3 0 amazon pilot 2-of-3 0 cohere long-tail 3-of-3 0 ai2 long-tail 3-of-3 0 tii long-tail 3-of-3 0 nvidia long-tail 3-of-3 0 ai21-labs long-tail 3-of-3 0 zhipu-ai long-tail 3-of-3 0 moonshot-ai long-tail 3-of-3 0 eleutherai long-tail 3-of-3 0 lg-ai-research long-tail 3-of-3 0 snowflake long-tail 3-of-3 0 upstage long-tail 3-of-3 0 ibm long-tail 3-of-3 0 baidu long-tail 3-of-3 0 tencent long-tail 3-of-3 0 bytedance-seed long-tail 3-of-3 0 stability-ai long-tail 3-of-3 0 databricks long-tail 3-of-3 0 minimax long-tail 3-of-3 0 apple long-tail 3-of-3 0 hugging-face long-tail 3-of-3 0 01-ai long-tail 3-of-3 0 sakana-ai long-tail 3-of-3 0 sarvam-ai long-tail 3-of-3 0 naver long-tail 3-of-3 0 aleph-alpha long-tail 3-of-3 0 reka-ai long-tail 3-of-3 0 nous-research long-tail 3-of-3 0 liquid-ai long-tail 3-of-3 0 xiaomi long-tail 3-of-3 0 ai-singapore long-tail 3-of-3 0 kyutai long-tail 3-of-3 0 lelapa-ai long-tail 3-of-3 0 maritaca-ai long-tail 3-of-3 0 openbmb long-tail 3-of-3 0 What those claims proposed to do Kind Count Effect Add 3 The Gemini 3.8 Flash release and its two documentation sources. All three were accepted by the panel and then withheld by the home critical-path byte ceiling; none reached the dataset. Change 2 A featured-rationale restate on Gemini 3.7 Flash, withheld with the record it depends on, and a verification-date advance on the Gemini deprecations source, which is the one edit this run published. Not covered
- Pages behind an authentication wall or a client-side rendering shell were not executed; only the served bytes were read.
- Benchmark figures and vendor performance claims were not extracted at all this run, by policy: they are contested, versioned and not a fact about a release.
- Creators other than google-deepmind were swept for novel tokens against the committed dataset; a change that introduces no new token would not have been seen by that filter.
Degraded discovery channels
- openai — official-announcement: openai.com/news/ returned HTTP 403 with a 9,908 byte body carrying no article text. This channel has now been dark across four consecutive runs. Reported rather than worked around: no user-agent was spoofed and no cache was substituted for the source.
- xai — official-announcement: x.ai/news returned HTTP 403, so xai has no readable announcement channel this run.
- zhipu-ai — official-announcement: z.ai/blog returned HTTP 404. The channel recorded in the catalogue no longer resolves.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 15
- Accepted by panel
- 5
- Rejected by panel
- 0
Deterministic gates and required checks — 7 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-source-approvalgoogle-deepmind bundle, pre-apply on a clean tree 0 PassAnchored on merge-base 04116fc43d20d49dc58acbdeb2f66a53ae351ef1. Both proposed origins are already stood behind by the committed dataset. This run approved no new publisher. gate-evidencegoogle-deepmind bundle 0 Pass5 claims admissible, 5 applicable, 0 failures, policy derived as pilot from the reviewed-profile set on disk rather than read from the bundle. Controlled in the same session: a sub-threshold mutation and a retrieval=search mutation were each refused at exit 1, so the exit 0 discriminates. gate-datasetapplied change 0 PassZero failures. Referential integrity and entity boundaries hold across the changed document. gate-reversalsapplied change 0 PassPasses, and the pass is narrower than it looks: the 2026-09-06 rejection of these same Gemini 3.8 records sits in this gate's acknowledged blind spot, so it was not checked rather than checked and cleared. Disclosed here because not-looked and looked-and-found-nothing are not the same reading. npm run validatenot requiredweb/, full tranche 1 FailRecorded as a gate that did its job. 3186 of 3189 tests passed; the 3 failures were the home critical-path budget at 1,110,273 over a 1,105,000 ceiling, and the measuredRaw drift that follows from it. This is what withheld the tranche. No ceiling was raised and no drift allowance was widened. npm run validateweb/, published change 0 PassThe published edit is a two-character date advance and adds no bytes to any route. Exit code read from an unpiped invocation. gate-scopecommitted HEAD and working tree 0 PassOnly web/src/data documents moved. web/asset-budgets.json was deliberately not touched: gate-scope would permit regenerating its measuredRaw field, but this run was instructed to confine itself to the dataset, and regenerating it would not have unblocked the tranche, which is refused by a ceiling rather than by the drift guard. gate-ledgerthis entry against the diff it describes 0 Passtranscription false. The entry describes the diff that shipped, including what did not ship and why. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
google-gemini-3-8-flash-release-addEditorial and entity boundariesRejected while provenance and consistency accepted, so the claim cleared the 2-of-3 pilot bar at 2. The dissent is that the proposed summary calls the release the most intelligent Flash tier of the Gemini 3 family, which is ranking language in the dataset's own voice and cannot stay true once a later Flash model ships. The chair did not vote, did not break the tie and did not reword the claim after the panel had read it, because publishing amended text would publish something no reviewer reviewed. The record was withheld on bytes before the wording could reach the dataset, and the wording is carried to a follow-up.
Posted 1 edit
1 edit across 1 document, a net change of 0 records.
Dataset documents this run changed Document Before After What changed web/src/data/sources.json295 295 One field advanced on google-gemini-deprecations: lastCheckedDate 2026-09-11 to 2026-09-13, attested by a page refetched and hashed on the day, which still carries the gemini-3.8-flash row with its September 2, 2026 release date and no announced shutdown date. Accepted unanimously against a 2-of-3 pilot threshold. No record was added or removed, so the document count is unchanged. Each document links to the file as this run left it, not as it stands today.
Not posted 4 items
Accepted by the panel, then dropped
google-gemini-3-7-flash-featured-rationale-restateAccepted unanimously, then dropped because its replacement text names Gemini 3.8 Flash as the model 3.7 Flash is the documented efficiency fallback from, and that record was withheld. The panel read this claim as part of a tranche in which 3.8 lands; publishing it alone would publish a configuration nobody reviewed. The wording it replaces - calling 3.7 Flash the newest generally available Flash tier - remains true of the dataset as it now stands, though it is already false of the world.Blocked bygoogle-gemini-3-8-flashweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.md
Blocked by policy before it could run
google-gemini-3-8-flashAccepted 2-of-3 by the panel and admitted by gate-evidence, then refused by the home route critical-path budget. Measured on this tree: clean HEAD 1,100,526 raw critical bytes, the tranche 1,110,273, ceiling 1,105,000, so it is over by 5,273. Real headroom on clean HEAD is 4,474 bytes, not the 24,445 that asset-budgets.json implies, because its recorded measuredRaw of 1,080,555 is 19,971 bytes stale. No ceiling was raised to admit this record.Blocked byweb/asset-budgets.json criticalMaxRawweb/tests/build/asset-budgets.test.tsdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdgoogle-gemini-3-8-flash-docsWithheld with the release it exists to cite. A source registered for a record that is not in the dataset would be an unreferenced entry, so the pair is kept together rather than split to fit.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdgoogle-gemini-3-8-flash-platform-docsWithheld with the release it exists to cite, for the same reason as its sibling source.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.md
What this run does not prove
- The home route is effectively full. Clean HEAD measures 1,100,526 raw critical bytes against a 1,105,000 ceiling, leaving 4,474 bytes; the three withheld records cost 9,747. At roughly 3,250 rendered bytes per release record this dataset can absorb at most one further release on the home route before the ceiling refuses it, whatever the panel says.
- asset-budgets.json records measuredRaw for home as 1,080,555 while the tree measures 1,100,526. The drift of 19,971 is inside the 21,611 allowance, so it passes, but at 92% of it - a subsequent run adding as little as 1,640 bytes would fail the drift test rather than the ceiling, and would be told the ceiling was fine. ADR 0015 admits measuredRaw to the qualifying class and would have permitted regenerating it; this run was instructed to confine itself to web/src/data, and regenerating it would not have unblocked the tranche in any case, because what refuses the tranche is criticalMaxRaw, which ADR 0015 explicitly keeps out of class.
- These same Gemini 3.8 Flash records were rejected by a panel on 2026-09-06 (docs-source-add 1-of-3, release-add 0-of-3) and accepted by a later panel on 2026-09-11, which then withheld them on the same byte ceiling. This run is a third independent panel and reached the 2026-09-11 result. rejection-reversals.json does not annotate the 2026-09-06 rejections and cannot be written by a refresh run, because it sits outside gate-scope ALLOWED_PATHS by design so that a run cannot absolve itself.
- gate-reversals passes, but its check does not cover these rejections; they are 2 of the 186 of 204 it acknowledges it cannot see. A pass here is the absence of a finding, not a finding of absence.
- The editorial dissent on the release summary was not acted on by rewording, because the chair does not amend text the panel has already read. The wording is carried to a follow-up rather than quietly fixed.
- Only google-deepmind produced claims. The other 43 creators were swept by comparing fetched pages against the committed dataset for novel tokens; a change carrying no new token would not surface through that filter, so a nil result is weaker evidence than a positive one.
Follow-ups — proposed, not fixed
- The home route has 4,474 bytes of headroom and each new release record costs about 3,250 rendered bytes, so the next refresh that finds more than one release will be refused exactly as this one was. This is now the binding constraint on the whole pipeline and it needs a page-weight decision, not another withheld tranche.
- asset-budgets.json records a home measuredRaw 19,971 bytes below what the tree measures, which is 92% of the drift allowance. A refresh confined to web/src/data cannot correct it, so it will keep eating the allowance until someone regenerates it.
- The Gemini 3.8 Flash records carry a 2026-09-06 panel rejection that rejection-reversals.json does not annotate, and no refresh run can add that annotation. Two later panels have since accepted the same records, so the register is out of step with the reviewed record.
- The summary wording proposed for the Gemini 3.8 Flash release drew an editorial rejection for ranking language. Whoever lands that record should settle the wording before it ships rather than inheriting the text this run withheld.
2026-09-12-ebc58b
MAI-Code-1.1-Flash, MAI-Voice-2 and Granite 4.2 land; the nvidia bundle is refused
Scope requested: Every creator in organizations.json plus the long-tail profile: 44 scouted, none unswept. 155 pages attempted across 5 fetch rounds, 152 returning content.Published9 edits posted · 5 items withheldA full sweep of all 44 creators. Nine claims across microsoft and ibm were accepted unanimously by the three-rubric panel and applied: two Microsoft AI releases with their announcement sources, two IBM Granite 4.2 releases with their model-card sources, and a coverage note on the granite-4-2 family. The whole nvidia bundle was refused at the evidence gate: nvidia-nemotron-3-5-lightning-release reached 1 of the 3 accepts the long-tail policy requires, the provenance and editorial rubrics independently rejecting an unsourced status="current" that the model card behind it records as unknown, and the record is schema-invalid regardless because OpenMDW-1.1 carries no SPDX identifier. Its two unanimously accepted siblings cascaded out with it rather than landing a family with no releases.
- PreflightRanClean tree, gh authenticated, no refresh pull request from a previous run still open. Drydock probed both forms: the bare name is refused by the execution policy and drydock.cmd reports 0.1.0, so the CLI is installed and blocked rather than absent.
- ScoutRan155 pages attempted, 152 ok, 3 failed. Every quote byte-verbatim from a fetched page, every page sha256-hashed on the day. No search snippet was used as evidence. Three bundles carried claims; 41 creators were scouted and yielded nothing.
- ReviewRanThree sub-agents, one per rubric, run in parallel and blind to each other and to the scout reasoning; each saw only the claim, its evidence and the dataset slice it touches. 36 verdicts over 12 claims. Consistency accepted all 12; editorial and provenance each rejected the same nvidia release claim independently.
- GatesRanBoth bundle gates ran before anything was applied, on a clean tree, because gate-source-approval anchors on the committed dataset. gate-evidence refused the nvidia bundle at exit 1. gate-dataset, npm run validate, gate-scope and gate-ledger all passed on the applied tranche.
- PublishRanTwo commits on one branch and one pull request: the dataset tranche first, then this ledger entry, which cannot be written until the pull request it must name exists. Merged by GitHub under --auto --squash once web-ci was green.
- DeployNot runRecorded before the merge, so the Pages deploy had not yet run when this entry was written. Confirmed separately after merge; a failed deploy is reverted by pull request rather than left to freeze the published site on its previous build.
What was found
- Scouts
- 1
- Pages fetched and hashed
- 155
- Claims proposed
- 12
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims openai pilot 2-of-3 0 anthropic pilot 2-of-3 0 google-deepmind pilot 2-of-3 0 meta pilot 2-of-3 0 xai long-tail 3-of-3 0 mistral-ai long-tail 3-of-3 0 deepseek long-tail 3-of-3 0 alibaba-cloud pilot 2-of-3 0 microsoft pilot 2-of-3 4 amazon pilot 2-of-3 0 cohere long-tail 3-of-3 0 ai2 long-tail 3-of-3 0 tii long-tail 3-of-3 0 nvidia long-tail 3-of-3 3 ai21-labs long-tail 3-of-3 0 zhipu-ai long-tail 3-of-3 0 moonshot-ai long-tail 3-of-3 0 eleutherai long-tail 3-of-3 0 lg-ai-research long-tail 3-of-3 0 snowflake long-tail 3-of-3 0 upstage long-tail 3-of-3 0 ibm long-tail 3-of-3 5 baidu long-tail 3-of-3 0 tencent long-tail 3-of-3 0 bytedance-seed long-tail 3-of-3 0 stability-ai long-tail 3-of-3 0 databricks long-tail 3-of-3 0 minimax long-tail 3-of-3 0 apple long-tail 3-of-3 0 hugging-face long-tail 3-of-3 0 01-ai long-tail 3-of-3 0 sakana-ai long-tail 3-of-3 0 sarvam-ai long-tail 3-of-3 0 naver long-tail 3-of-3 0 aleph-alpha long-tail 3-of-3 0 reka-ai long-tail 3-of-3 0 nous-research long-tail 3-of-3 0 liquid-ai long-tail 3-of-3 0 xiaomi long-tail 3-of-3 0 ai-singapore long-tail 3-of-3 0 kyutai long-tail 3-of-3 0 lelapa-ai long-tail 3-of-3 0 maritaca-ai long-tail 3-of-3 0 openbmb long-tail 3-of-3 0 What those claims proposed to do Kind Count Effect Add 11 Four source registrations, four releases and three nvidia records proposed as new entries; nine of the eleven were applied. Change 1 A coverage note on the ibm-granite-4-2 family description, reflecting the two sizes added alongside it. Not covered
- Pages behind an authentication wall or a client-side rendering shell were not executed; only the served bytes were read.
- Benchmark figures and vendor performance claims were not extracted at all this run, by policy: they are contested, versioned and not a fact about a release.
Degraded discovery channels
- openai — official-announcement: openai.com/news/ returned HTTP 403, as did /index/gpt-4-1/ and /index/gpt-5-6/ - 3 of the run's 155 fetches. Reported rather than worked around: the run did not scrape past it or spoof a user agent. OpenAI ships no open weights, so the Hugging Face gap signal is structurally blind to it and this 403 left the run with no discovery channel for that creator at all.
- zhipu-ai — official-announcement: z.ai/blog/* served a 598 to 604 byte JavaScript shell carrying no article text. Body sizes across the run range from 598 B to 1.24 MB, so "too small to be an article" discriminates here on measurement rather than on assumption.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 36
- Accepted by panel
- 11
- Rejected by panel
- 1
Deterministic gates and required checks — 7 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-source-approvalmicrosoft, ibm and nvidia bundles, pre-apply on a clean tree 0 PassEvery claim rests on an origin the committed dataset or a reviewed profile catalogue already stands behind. This run approved no new publisher. gate-evidencemicrosoft and ibm bundles 0 PassEvery quote resolves byte-verbatim to a fetched page whose sha256 the bundle records, and every claim meets its policy threshold. gate-evidencenot requirednvidia bundle 1 FailRecorded as a gate that did its job. The bundle-wide verdict refused all three nvidia claims because one of them reached 1 of 3 accepts against a unanimous long-tail threshold; nothing from the bundle was applied. Marked not-required because this failure is the mechanism that withheld the work, not a blocker on the tranche that shipped. gate-datasetapplied tranche 0 PassZero failures. Referential integrity, entity boundaries and the family-has-release rule all hold across the three changed documents. npm run validateweb/ 0 Pass3189 of 3189 tests pass across 139 of 139 files with 0 not executed, and astro check reports 0 errors, 0 warnings, 0 hints. Exit code read from an unpiped invocation with a failing control for comparison, because a pipeline stage that stops early corrupts it. gate-scopecommitted HEAD and working tree 0 Passchanged=3, outOfClass empty. Only the three dataset documents moved in the data commit; this ledger entry is the one further in-class document, put in class by ADR 0006. gate-ledgerthis entry against the diff it describes 0 Passtranscription false, meaning the entry was written on a branch that genuinely carries the data change it reports rather than on an empty one where nothing could verify it. ci-preflightrepository root, checks selected from the branch diff 0 PassSelects the pull-request checks this diff actually triggers and runs them locally. Does not cover the networked link-health sweep or the second Python interpreter, so a green preflight is not a green CI on its own. Posted 9 edits
9 edits across 3 documents, a net change of 8 records.
Dataset documents this run changed Document Before After What changed web/src/data/sources.json291 295 Two Microsoft AI announcement posts and two IBM Granite 4.2 model cards registered as sources. web/src/data/releases.json124 128 MAI-Code-1.1-Flash, MAI-Voice-2, Granite 4.2 8B and Granite 4.2 3B, each with a publisher-stated date and so needing no dateBasis marker. web/src/data/families.json86 86 One description changed on ibm-granite-4-2 to note the sizes now covered. No family was added or removed. Each document links to the file as this run left it, not as it stands today.
Records added
- MAI-Code-1.1-Flashpassportreleases
microsoft-mai-code-1-1-flashDated 2026-08-11 from the announcement page, which states the date beside the headline and republishes it as JSON-LD datePublished. - MAI-Voice-2passportreleases
microsoft-mai-voice-2Registered from Microsoft AI's own announcement post, unanimous 3/3 against a 2-of-3 pilot threshold. - Granite 4.2 8Bpassportreleases
ibm-granite-4-2-8bFrom the IBM model card, unanimous 3/3 against the unanimous long-tail threshold. - Granite 4.2 3Bpassportreleases
ibm-granite-4-2-3bFrom the IBM model card, unanimous 3/3 against the unanimous long-tail threshold.
Not posted 5 items
Rejected by the review panel
nvidia-nemotron-3-5-lightning-releaseReached 1 of the 3 accepts a long-tail creator requires. The provenance and editorial rubrics rejected it independently and converged on the same reason: status="current" is stated by no quote, and the bare model card it rests on is recorded by schema.ts as status unknown. It is also schema-invalid on its own terms, failing validation on releases.128.license.spdxId because OpenMDW-1.1 carries no SPDX identifier. No reviewer was re-run and no vote was overruled.Blocked byrubric:provenancerubric:editorialgate-evidenceweb/src/data/schema.ts
Accepted by the panel, then dropped
nvidia-nemotron-3-5-lightning-sourceUnanimous 3/3, and still not applied. Landing a source without the release it pairs with fails check-bundle-pairing, and the bundle-wide evidence verdict refused the bundle whole.Blocked bynvidia-nemotron-3-5-lightning-releasecheck-bundle-pairingnvidia-nemotron-3-5-familyUnanimous 3/3, and still not applied. A family with no releases throws at validate.ts:599 and fails the family-has-release gate in gate-dataset, so accepting it without the refused release was never available.Blocked bynvidia-nemotron-3-5-lightning-releasegate-dataset:family-has-release
Verification date deliberately held back
platform-dated-release-candidatesGLM-5.3, GLM-5.3-Flash, LFM2.5, MiniMax-Music3, Hy4-preview, Qwen-Drive, VibeVoice-ASR and Gemini 3.8 Flash were found, and each has only a hosting-platform observation for its date rather than a publisher-stated one. date-basis-policy.test.ts requires such a record to carry dateBasis and then pins the complete set of dateBasis-bearing ids with toEqual, so adding any new one fails it, and the only repair is editing a test file. That is outside the qualifying class, so the candidates were withheld rather than the class widened. The four releases that did land are unaffected because each carries a publisher-stated date and needs no marker.Blocked byweb/src/data/date-basis-policy.test.tsdocs/adr/0003-an-agent-gated-data-refresh-may-auto-merge.md
Blocked by policy before it could run
bundle:nvidiaThe bundle-wide gate verdict is what the applier honours, so a single sub-threshold claim withheld all three. Recorded as policy rather than as a defect: the alternative is a panel majority on two claims overriding a deterministic gate on a third.Blocked bygate-evidence
What this run does not prove
- The gates catch malformed, impossible, unreferenced and boundary-violating data. They do not catch a claim that is well-formed and simply wrong, and neither does web-ci: a plausible wrong date passes both.
- All three reviewers ran claude-opus-4.8, so they share a failure mode. Three instances of one model reading one page can agree and be wrong together, and a source that is itself wrong can carry all three. The independent convergence of provenance and editorial on the nvidia rejection is evidence the panel can discriminate, not proof it always will.
- A 3189-test pass is a statement about the dataset being well-formed and internally consistent, not about it being true.
- Coverage of 44 of 44 creators means every creator was scouted, not that every release each creator shipped was found. A creator publishing only through a channel this run could not read - as OpenAI did, behind a 403 - can have shipped something this run reports nothing about.
- The asset-budget figures were left untouched because a real production build with this tranche applied passed all 44 assertions at 1.848% drift against a 2% guard. That is a measurement of this build on this machine, not a guarantee about the deployed one.
- This entry was written before the Pages deploy ran, so its deploy stage records not-run rather than a result. The deploy was confirmed separately after merge.
Follow-ups — proposed, not fixed
- OpenAI's official-announcement channel has now returned 403 across this run; a creator whose only discovery channel is dark cannot be swept meaningfully however green the run looks.
- The date-basis-policy test pins its expected set with toEqual, which makes every platform-dated candidate a stop for an agent run rather than a claim it can weigh. That coupling is worth a human decision, not a refresh's.
2026-09-11-cd4874
Full 44-creator sweep: 190 re-verifications published, every new record withheld at the home-route ceiling
Scope requested: All 44 creators scouted, so found.unswept is empty: nothing was passed off as unchanged that was merely not looked at. 7 creators carry a reviewed profile and ran at the 2-of-3 pilot threshold; the other 37 ran at the unanimous 3-of-3 long-tail threshold. Only web/src/data/sources.json changed, re-dating lastCheckedDate on 189 sources and correcting one notes field. No record was added, removed or re-identified, and the record count is unchanged at 291.Published190 edits posted · 82 items withheldA full sweep of all 44 creators, 477 pages fetched and hashed on 2026-09-11, 398 claims proposed and 1194 verdicts cast by blind three-rubric panels - exactly three per claim. The panel accepted 357 and rejected 41. gate-evidence then refused 19 bundles whole, all of them long-tail, costing 145 individually-accepted claims. Of what survived, the run applied 190 claims and withheld all 14 accepted new records: adding them puts the home route at 1106468 bytes against a criticalMaxRaw of 1105000, and raising a ceiling is the one move ADR 0015 puts outside the qualifying class. The tranche was therefore narrowed by a stated structural rule - re-verification only, no new records - rather than by choosing which new model to hide, which would have been an unrecorded editorial judgement. web/asset-budgets.json was not touched at all. A separate evidence-reproducibility checker, two-sided controlled, found 25 of 491 evidence items unreproducible across 18 claims; all 18 had already been rejected by the blind panel, so two independent instruments converged completely.
- PreflightRanClean worktree, gh authenticated, no competing refresh pull request open. Anchor d395d2c8fda531c8b633885d6fd56594b9788755 resolved and confirmed against a live git ls-remote read, then re-read at report time and found unchanged. A full baseline was taken before any edit - all gates at exit 0 and npm run validate at exit 0 with 139 files and 3189 tests - so any later failure is attributable to this run rather than inherited.
- ScoutRan477 pages fetched and hashed over the bytes received on 2026-09-11 across 7 parallel scouts covering all 44 creators. 26 channels degraded and are named per creator in found.degradedChannels rather than folded into a flat error list. gate-source-approval passed 44 of 44 bundles, so no claim rests on an unapproved origin.
- ReviewRan18 blind reviewer agents cast 1194 verdicts, exactly 398 x 3, so every claim received three independent reviewers and none was skipped. Each reviewer saw only the claim, its evidence, the relevant dataset slice, the creator profile and its own rubric - never another reviewer verdict, the running tally, or the scout reasoning. The chair cast no vote, broke no tie and re-ran no reviewer. 357 accepted, 41 rejected, 8 accepted over a recorded dissent.
- GatesRangate-source-approval 44/44, gate-evidence 25 pass and 19 refuse, gate-dataset exit 0 over 634 records, gate-scope exit 0 with outOfClass empty and one document changed. npm run validate exit 0 with 139/139 files and 3189/3189 tests and 0 Astro errors; ci-preflight.mjs exit 0 with web-ci, skills-ci and source-link-health-tests all passing. Exit codes were read from unpiped invocations after a piped read of validate reported 1 where the true status was 0 - the Select-Object early-termination hazard, not a test failure.
- PublishRanPublished as pull request #1150 with auto-merge enabled on a squash, gated by web-ci. The full per-claim evidence table and all 41 verbatim reviewer rationales are posted as comments on that pull request, because the body limit will not hold them and truncating a reviewer reasoning would defeat the point of publishing it. No bypass was used: no --force, no --admin, no skipped gate, no lowered threshold and no direct push to main.
- DeployNot runNot yet observed at the moment this entry was written, because the entry ships in the same pull request it describes and the Pages deploy runs on push to main afterwards. Recorded as not-run rather than as a pass: a deploy nobody has looked at is not a deploy that succeeded.
What was found
- Scouts
- 7
- Pages fetched and hashed
- 477
- Claims proposed
- 398
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims 01-ai long-tail 3-of-3 5 ai-singapore long-tail 3-of-3 20 ai2 long-tail 3-of-3 4 ai21-labs long-tail 3-of-3 12 aleph-alpha long-tail 3-of-3 3 alibaba-cloud pilot 2-of-3 17 amazon pilot 2-of-3 6 anthropic pilot 2-of-3 19 apple long-tail 3-of-3 7 baidu long-tail 3-of-3 4 bytedance-seed long-tail 3-of-3 2 cohere long-tail 3-of-3 17 databricks long-tail 3-of-3 2 deepseek long-tail 3-of-3 21 eleutherai long-tail 3-of-3 5 google-deepmind pilot 2-of-3 22 hugging-face long-tail 3-of-3 26 ibm long-tail 3-of-3 5 kyutai long-tail 3-of-3 4 lelapa-ai long-tail 3-of-3 5 lg-ai-research long-tail 3-of-3 5 liquid-ai long-tail 3-of-3 3 maritaca-ai long-tail 3-of-3 5 meta pilot 2-of-3 25 microsoft pilot 2-of-3 17 minimax long-tail 3-of-3 7 mistral-ai long-tail 3-of-3 18 moonshot-ai long-tail 3-of-3 8 naver long-tail 3-of-3 5 nous-research long-tail 3-of-3 3 nvidia long-tail 3-of-3 7 openai pilot 2-of-3 27 openbmb long-tail 3-of-3 12 reka-ai long-tail 3-of-3 3 sakana-ai long-tail 3-of-3 5 sarvam-ai long-tail 3-of-3 4 snowflake long-tail 3-of-3 1 stability-ai long-tail 3-of-3 5 tencent long-tail 3-of-3 8 tii long-tail 3-of-3 4 upstage long-tail 3-of-3 1 xai long-tail 3-of-3 7 xiaomi long-tail 3-of-3 3 zhipu-ai long-tail 3-of-3 9 What those claims proposed to do Kind Count Effect Change 312 Re-verification of facts already in the dataset, overwhelmingly lastCheckedDate on existing sources. 190 of these were applied; the rest were lost to bundle-level gate-evidence refusal or rejected by the panel. Add 85 Proposed new sources, families and releases. None reached the dataset: 14 accepted adds were withheld at the home-route asset ceiling, 7 were dropped on pinned test registries or their cascade, and the remainder were rejected or lost to bundle refusal. Conflict 1 One recorded conflict between primary sources, published as a finding rather than resolved by the pipeline. Not covered
- The 9 OpenAI announcement and news pages that returned HTTP 403, and the xAI news channel that failed outright. Those creators were swept through their remaining reachable channels only, so their coverage this run is partial and is reported as partial.
- No workaround was attempted against any channel that refused. A page that returns 403 is recorded as degraded rather than fetched by another route.
- The 19 bundles gate-evidence refused were scouted and reviewed in full; what they did not get is an applied change. They are not unswept, and they are named individually in withheld.
Degraded discovery channels
- ai-singapore — ai-singapore-gemma-sea-lion-v3-9b-docs-lowercase-404: https://github.com/aisingapore/sealion/blob/main/models/sea-lion-v3/gemma-sea-lion-v3-9b.md returned error
- ai-singapore — ai-singapore-llama-sea-lion-v3-70b-docs-lowercase-404: https://github.com/aisingapore/sealion/blob/main/models/sea-lion-v3/llama-sea-lion-v3-70b.md returned error
- cohere — cohere-transcribe-c4ai-model-card-probe: https://huggingface.co/CohereLabs/c4ai-transcribe-4b-v1 returned http-401
- cohere — hugging-face-cohere-transcribe-c4ai-hub-record-probe: https://huggingface.co/api/models/CohereLabs/c4ai-transcribe-4b-v1 returned http-401
- google-deepmind — google-gemini-3-8-flash-cyber-docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-cyber returned http-404
- google-deepmind — google-gemini-3-8-flash-cyber-platform-docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-8-flash-cyber returned http-404
- ibm — ibm-granite-4-0-h-small-model-card: https://huggingface.co/ibm-granite/granite-4.0-h-small returned http-429
- ibm — ibm-granite-4-0-h-tiny-model-card: https://huggingface.co/ibm-granite/granite-4.0-h-tiny returned http-429
- ibm — ibm-granite-4-2-30b-model-card: https://huggingface.co/ibm-granite/granite-4.2-30b returned http-429
- ibm — ibm-granite-4-2-technical-blog: https://huggingface.co/blog/ibm-granite/granite-4-2 returned http-429
- lg-ai-research — lg-ai-research-exaone-3-5-model-card: https://huggingface.co/LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct returned http-429
- mistral-ai — mistral-medium-3-5-hf-model-card-probe: https://huggingface.co/mistralai/Mistral-Medium-3.5-128B-Instruct-2604 returned http-401
- openai — openai-chatgpt-images-2-5-announcement: https://openai.com/index/introducing-chatgpt-images-2-5 returned http-403
- openai — openai-gpt-4-1-announcement: https://openai.com/index/gpt-4-1/ returned http-403
- openai — openai-gpt-5-4-announcement: https://openai.com/index/introducing-gpt-5-4/ returned http-403
- openai — openai-gpt-5-5-announcement: https://openai.com/index/introducing-gpt-5-5/ returned http-403
- openai — openai-gpt-5-6-launch: https://openai.com/index/gpt-5-6/ returned http-403
- openai — openai-gpt-5-6-price-update: https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/ returned http-403
- openai — openai-gpt-5-6-sol-preview: https://openai.com/index/previewing-gpt-5-6-sol/ returned http-403
- openai — openai-gpt-live-1-announcement: https://openai.com/index/introducing-gpt-live-1-in-the-api returned http-403
- openai — openai-news: https://openai.com/news/ returned http-403
- snowflake — snowflake-arctic-instruct-model-card: https://huggingface.co/Snowflake/snowflake-arctic-instruct returned http-429
- tencent — tencent-hunyuanimage-3-0-instruct-distil-license: https://huggingface.co/tencent/HunyuanImage-3.0-Instruct-Distil/raw/main/LICENSE returned http-404
- upstage — upstage-solar-pro-preview-model-card: https://huggingface.co/upstage/solar-pro-preview-instruct returned http-429
- xai — xai-news: https://x.ai/news returned failed
- zhipu-ai — zhipu-ai-glm-4-5-announcement: https://z.ai/blog/glm-4.5 returned empty-text
What was evaluated
- Reviewers
- 18
- Verdicts cast
- 1,194
- Accepted by panel
- 357
- Rejected by panel
- 41
Deterministic gates and required checks — 6 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-source-approvalAll 44 claim bundles 0 PassEvery cited origin is in the reviewed catalogue. 44 of 44 bundles at exit 0. gate-evidenceThe 25 bundles that passed, which are the only bundles contributing a record to this pull request 0 PassExit 0. Every claim carries a verbatim quote, a sha256 hash and a fetch date, and every claim met the threshold its policy sets - 2-of-3 for a reviewed profile, unanimous 3-of-3 for a long-tail creator. Per ADR 0005 this gate verifies the form of a citation and never its correspondence to the cited URL, so it is not evidence that any quote is real; a separate reproducibility checker addressed that. gate-evidencenot requiredThe 19 bundles it refused, run to confirm the withholding mechanically 1 FailExit 1 on all 19, every one of them long-tail, naming the claims that fell short of the unanimous 3-of-3 bar. The run honoured every refusal exactly: nothing from any of these bundles is in this pull request, so this gate run governs no published record. It is expensive rather than wrong - refusal is per bundle, so 145 claims the panel had individually accepted were discarded alongside the blocking ones, and that cost is filed as a follow-up rather than worked around. gate-datasetThe applied tranche, 634 records 0 PassSchema, referential integrity and the family-has-release rule all hold. An earlier tranche failed this gate when dropping a release orphaned its family; the family and the 4 sources whose sole referents were dropped went with it rather than landing as a data error. gate-scopeBranch diff against merge-base d395d2c8fd 0 PassOne dataset document changed, web/src/data/sources.json, with outOfClass empty. web/asset-budgets.json was not touched in any field. npm run validateweb/ 0 Pass139 of 139 test files and 3189 of 3189 tests passed with 0 Astro errors, matching the pre-edit baseline exactly. Every asset-budget test passes. ci-preflightnot requiredRepository root, checks selected from the branch diff 0 Passweb-ci, skills-ci and source-link-health-tests all pass locally. It does not cover the networked link-health sweep, the second Python interpreter, web-e2e or the runner itself, so it predicts CI rather than binding it. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
qwen3-8-flash-next-repository-last-checked-2026-09-11Editorial and entity boundariesThe current source record notes import the comparative phrase "delivers superior capabilities in coding and office tasks" into data; editorial forbids rank/winner language even when a source says it.alibaba-qwen-drive-1-0-4b-release-addProvenanceThe attached quotes support downloadability (`The model can be downloaded from Hugging Face or ModelScope`), weights, createdAt, and Apache licensing, but none of this claim's quoted evidence states `google-gemini-3-8-flash-release-addEditorial and entity boundariesGoogle profile says the canonical release name is the API model id and the marketing name is an alias; the proposal sets canonicalName/displayName to Gemini 3.8 Flash while treating gemini-3.8-flash amicrosoft-mai-transcribe-2-announcement-last-checked-2026-09-11Editorial and entity boundariesThe source title and claim statement import a winner/rank phrase, "fastest, most accurate and cheapest speech recognition model in the world"; editorial forbids ranking language in data or phrasing.microsoft-mai-image-2-6-announcement-last-checked-2026-09-11Editorial and entity boundariesThe source title and claim statement import the composite ranking phrase "quality-cost frontier"; editorial forbids frontier/rank language in data or phrasing.microsoft-mai-code-1-1-flash-announcement-source-addEditorial and entity boundariesThe proposed source title imports comparative ranking, "Better, faster, at a quarter of the cost", and the notes/quote carry quality-efficiency-cost comparisons; editorial forbids rank/composite winneopenai-gpt-image-2-5-sunburst-release-addEditorial and entity boundariesOpenAI release naming rule is dated snapshot as canonical release and dateless name as API alias; the proposal sets canonicalName/displayName to GPT-Image-2.5 Sunburst while placing gpt-image-2.5-sunbopenai-gpt-image-2-5-flare-release-addEditorial and entity boundariesOpenAI release naming rule is dated snapshot as canonical release and dateless name as API alias; the proposal sets canonicalName/displayName to GPT-Image-2.5 Flare while placing gpt-image-2.5-flare-2
Posted 190 edits
190 edits across 1 document, a net change of 0 records.
Dataset documents this run changed Document Before After What changed web/src/data/sources.json291 291 Re-dates lastCheckedDate on 189 sources across 25 creators against freshly fetched and hashed primary pages, and corrects one notes field. No record added, removed or re-identified, which is why posted.records is empty: nothing landed in releases, families or organizations this run. Each document links to the file as this run left it, not as it stands today.
Not posted 82 items
Rejected by the review panel
ai-singapore:osi-license-index-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. Reject: the cited contentHash sha256:94c72ae924ee69fa5e1fd3481013d287afdd941a8b84bdcff8e629e279923432 does not match any stored .raw page for this run; the only osi-license-index.raw present hashes to sha256:9304b108a09fbd97c7c25474726be0cbd5760dd72be2ff9ed5b1b9b615af9965, so theBlocked byprovenanceai-singapore:ai-singapore-sea-lion-v3-description-record-three-models1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. Reject: the only attached quote says SEA-LION v3 is "a collection of 3 models (and their variants)" but does not name Gemma 9B or Llama 70B, so it does not support the proposed description's specific named list.Blocked byprovenanceeditorialai-singapore:ai-singapore-gemma-sea-lion-v3-9b-release-add1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. Reject: the quotes support the name, 9B size, 8192 context, download availability, and "License: Gemma Community License", but no attached quote states lifecycle status "current", and the generic OSI-process quote does not state that Gemma Community License is not OSI-approved.Blocked byprovenanceeditorialai-singapore:ai-singapore-llama-sea-lion-v3-70b-release-add1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. Reject: the attached Llama quote states "70 billion parameters" and download availability, but no quote states the proposed 128000-token context window or lifecycle status "current", and the generic OSI-process quote does not state that the Llama 3.1 Community License is not OSI-Blocked byprovenanceeditorialai-singapore:ai-singapore-llama-sea-lion-v3-8b-sibling-ids-add-v3-peers2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. Reject: the quote says only that SEA-LION v3 is "a collection of 3 models (and their variants)"; it does not name the proposed sibling ids Gemma-SEA-LION-v3-9B and Llama-SEA-LION-v3-70B.Blocked byprovenanceai-singapore:ai-singapore-sea-lion-v4-5-family-add1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. Reject: the quotes support v4.5 as "foundational, agentic, and multimodal" and identify Qwen/Gemma sources, but none of the attached quotes states the proposed "coding" category.Blocked byprovenanceeditorialai21-labs:osi-approved-licenses-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The cited contentHash sha256:b4c6f6a756d1eba09084fd3db309aec635290e58fb4ddda677ed92a5e3837343 does not match the fetched osi-approved-licenses.raw bytes, which hash to sha256:087690e000fec800e0e5e6394a74f0391e65b60e409df4c4b8d55b65fcb571ff.Blocked byprovenanceai21-labs:ai21-labs-jamba-instruct-release-add2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quote says "Jamba-Instruct ... is in public preview on the AI21 Platform", so the proposed status unknown is not supported; the source states a preview lifecycle value.Blocked byprovenanceapple:apple-openelm-license-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quoted string "Disclaimer: IMPORTANT: This Apple Machine Learning Research Model is" is generic Apple model-license boilerplate; it does not name OpenELM, OpenELM-3B-Instruct, or any source-specific license title, so it does not prove the catalogued OpenELM license source stiBlocked byprovenanceapple:apple-fastvlm-license-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quoted string "Disclaimer: IMPORTANT: This Apple Machine Learning Research Model is" is generic Apple model-license boilerplate; it does not name FastVLM, FastVLM-7B, or any source-specific license title, so it does not prove the catalogued FastVLM license source still identiBlocked byprovenancecohere:cohere-command-a-plus-announcement-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quote is only the generic banner 'Skip to content AI for Empowerment: Your freedom. Your focus...' and does not identify the Command A+ announcement page, so it does not re-verify that catalogued source.Blocked byprovenancecohere:cohere-transcribe-announcement-source2/3 accepts against a threshold of 3. Dissenting rubric(s): editorial. Reject: the proposed source title and claim statement import "state-of-the-art" into ModelTree source data/phrasing; the editorial rubric forbids rank or winner language even when it appears in an announcement title.Blocked byeditorialcohere:cohere-north-small-translate-announcement-source2/3 accepts against a threshold of 3. Dissenting rubric(s): editorial. Reject: the proposed source title and claim statement import the rank term "leading" into ModelTree source data/phrasing; the editorial rubric forbids rank or winner language.Blocked byeditorialcohere:cohere-north-small-translate-model-card-source2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quote is tokenizer/chat-template text beginning with eos_token/pad_token/platform_instruction_override; it does not name North Small Translate or CohereLabs/North-Small-Translate-1.0, so it does not support adding that model-card source.Blocked byprovenancecohere:cohere-transcribe-03-2026-release2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quotes support ASR, audio-in/text-out, 2B parameters, and download availability, but no model quote states Apache 2.0 as the model licence or hosted/API/Model Vault access for accessType 'both'; the OSI page only proves Apache approval.Blocked byprovenancecohere:cohere-north-small-translate-1-0-release2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The model-card quote supports an open-weights research MoE with 25B active and 218B total parameters for translation, but no provided quote states the CC BY-NC 4.0 licence or the proposed 16K input/output context windows.Blocked byprovenancecohere:cohere-parse-v5-0-release1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. The quote supports Cohere Parse as a document-processing vision language model, but it does not state the proposed parse-v5.0 API alias/version, current hosted availability, or proprietary-hosted access type.Blocked byprovenanceeditorialdeepseek:deepseek-v4-1-family1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. The quoted source names the release 'DeepSeek-V4.1-Flash'; it does not state a separate 'DeepSeek-V4.1' family, and the hub-record quote provided does not include createdAt to support the proposed 2026-09-10 firstReleaseDate.Blocked byprovenanceeditorialdeepseek:deepseek-v4-1-flash-release1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. The model-card quote supports multimodal MoE, 552B backbone parameters, image/text input, text output, and one-million-token context, but the provided hub-record quote does not include createdAt, so the proposed 2026-09-10 releaseDate/dateBasis is not quoted.Blocked byprovenanceeditorialdeepseek:deepseek-v4-pro-0813-release1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. The model-card quote says DeepSeek-V4-Pro-0813 is the official release superseding the preview, but neither quote states the proposed 2026-08-13 repository-created release date; the hub-record quote shown does not include createdAt.Blocked byprovenanceeditorialdeepseek:deepseek-v4-flash-0731-release1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. The model-card quote says DeepSeek-V4-Flash-0731 is the official release superseding the preview, but neither quote states the proposed 2026-07-31 repository-created release date; the hub-record quote shown does not include createdAt.Blocked byprovenanceeditorialhugging-face:hugging-face-gguf-docs-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quoted string begins with the documentation navigation text "?? View all docs AWS Trainium & Inferentia..." and does not name GGUF or the GGUF docs page; it is boilerplate navigation, not evidence that the catalogued GGUF source still identifies itself.Blocked byprovenanceibm:ibm-granite-4-0-h-micro-release-add2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. For open-weight access the evidence has "open sourced under a standard Apache 2.0 license" and "Download language models", but no quote states weights are downloadable; the OSI Apache evidence hash also does not match the fetched raw bytes.Blocked byprovenanceibm:ibm-granite-4-0-micro-release-add2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. For open-weight access the evidence has "open sourced under a standard Apache 2.0 license" and "Download language models", but no quote states weights are downloadable; the OSI Apache evidence hash also does not match the fetched raw bytes.Blocked byprovenancekyutai:osi-license-index-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. Reject: the cited contentHash sha256:e68145cd8074237ae72422bf5eb1d6559f1c30896ab494cb82265edce4eb89f7 does not match any stored .raw page for this run; the only osi-license-index.raw present hashes to sha256:9304b108a09fbd97c7c25474726be0cbd5760dd72be2ff9ed5b1b9b615af9965, so theBlocked byprovenancelelapa-ai:osi-license-index-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. Reject: the cited contentHash sha256:f0205cc1dd07ecbca01b3343be9bbaa9957840aeb864a753a43c24718bb8d33f does not match any stored .raw page for this run; the only osi-license-index.raw present hashes to sha256:9304b108a09fbd97c7c25474726be0cbd5760dd72be2ff9ed5b1b9b615af9965, so theBlocked byprovenancemaritaca-ai:osi-license-index-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. Reject: the cited contentHash sha256:a2c93600a45439ef11e3d486926e007d79eb5862dde2fb6305113a5ea08dd4e7 does not match any stored .raw page for this run; the only osi-license-index.raw present hashes to sha256:9304b108a09fbd97c7c25474726be0cbd5760dd72be2ff9ed5b1b9b615af9965, so theBlocked byprovenanceminimax:minimax-model-license-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quoted string "Model Release Date: 15 January 2025" is present, but it does not name MiniMax, the MiniMax Model License, or the catalogued license source; the quote alone supports only a date line, not re-verification of that source.Blocked byprovenancemistral-ai:mistral-medium-3-5-128b-release1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. The quotes support the name, 128B size, open weights, and Modified MIT wording, but no provided quote states the proposed 256k context window, image input, public-preview lifecycle, or hosted/API access required by accessType 'both'.Blocked byprovenanceeditorialmistral-ai:mistral-voxtral-tts-4b-2603-release1/3 accepts against a threshold of 3. Dissenting rubric(s): provenance, editorial. The quotes state Voxtral TTS, 4B parameters, text-to-speech/voice-cloning details, open weights/BF16 weights, and CC BY-NC wording, but none states the proposed hosted API/Studio access needed for accessType 'both' or the exact release date.Blocked byprovenanceeditorialmoonshot-ai:osi-license-index-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The cited contentHash sha256:66f0bc037700f668da07e931e1a703c2347a8d32b396e2eeb36b95402aac0b71 does not match the fetched osi-license-index.raw bytes, which hash to sha256:9304b108a09fbd97c7c25474726be0cbd5760dd72be2ff9ed5b1b9b615af9965.Blocked byprovenancemoonshot-ai:moonshot-ai-kimi-k2-instruct-0905-model-card-source-add2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The cited contentHash sha256:12d67666d2a8b6790e738ab87925848d97b4628350bfdf88452106c734092278 does not match moonshot-ai-kimi-k2-instruct-0905-model-card.raw sha256:12d676ae66ec2b3dab35b76660626e94121b85b8868d98bf1df047b78b9f7ec4.Blocked byprovenancemoonshot-ai:hugging-face-kimi-k2-instruct-0905-hub-record-source-add2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The cited contentHash sha256:35b6f803c9e89f1175ad86f43b2779fed6468d9f40e5f749b8120bd4fb7e3f3e does not match hugging-face-kimi-k2-instruct-0905-hub-record.raw sha256:35b6f8eb5a864d7f8aca36dbe2cdfc16b4c69184e572a13105a1e3f6adacdd6f.Blocked byprovenancemoonshot-ai:moonshot-ai-kimi-k2-instruct-0905-release-add2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The load-bearing 0905 model-card, Hub-record, and OSI-index evidence hashes do not match the fetched raw files, so the quoted strings are not tied to the cited bytes for the release, date, access, or OSI-false fields.Blocked byprovenancenaver:naver-hyperclova-x-seed-announcement-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quoted sentence "Businesses and research institutions... can download these models" is anaphoric and does not name HyperCLOVA X SEED, NAVER, or the announcement title; read alone it does not identify the catalogued source.Blocked byprovenancesakana-ai:sakana-ai-evollm-jp-model-card-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quoted sentence "This model is provided for research and development purposes only..." is a generic model-card disclaimer and does not name EvoLLM-JP-v1-7B or SakanaAI, so it does not prove the catalogued model-card source still identifies itself.Blocked byprovenancesakana-ai:sakana-ai-evollm-jp-license-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quoted sentence "These license terms are an agreement between you and Microsoft Corporation..." is generic Microsoft license boilerplate and does not name EvoLLM-JP-v1-7B, SakanaAI, or the catalogued license source.Blocked byprovenancesarvam-ai:sarvam-ai-sarvam-m-blog-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The quoted sentence "Download the model from Hugging Face, try it on our playground, and build with our APIs" says only "the model" and does not name Sarvam-M or the announcement source; read alone it does not identify the catalogued source.Blocked byprovenancetencent:tencent-hunyuanimage-3-0-instruct-release2/3 accepts against a threshold of 3. Dissenting rubric(s): editorial. Reject: the release naming keeps `HunyuanImage-3.0-Instruct` and the Hub download separate, but the proposed `intendedUse` adds `chain-of-thought-style visual reasoning`; the attached quotes say only `Instruct (with reasoning)` / `Instruction reasoning` and image-to-image generatBlocked byeditorialupstage:osi-license-mit-last-checked-2026-09-112/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The cited contentHash sha256:64b9a6077ff220cdb35ad7f6f4fa3f91bdd95a72112de4d1d7233c8356e083bf does not match the fetched osi-license-mit.raw bytes, which hash to sha256:fccedc89f37399eefd47c15460bf41426cc3a3d296ea9ee3d5947aa055724716.Blocked byprovenancezhipu-ai:zhipu-ai-glm-4-5-release-add2/3 accepts against a threshold of 3. Dissenting rubric(s): provenance. The MIT OSI evidence is not tied to the fetched bytes: the cited sha256:64b9a6077ff220cdb35ad7f6f4fa3f91bdd95a72112de4d1d7233c8356e083bf does not match osi-license-mit.raw sha256:fccedc89f37399eefd47c15460bf41426cc3a3d296ea9ee3d5947aa055724716.Blocked byprovenance
Accepted by the panel, then dropped
alibaba-qwen-drive-1-0-4b-release-addAdding this release puts a twelfth id into the deliberately pinned registry in web/src/data/date-basis-policy.test.ts (dateBasis 'platform-repository-created'), and a sixteenth into the pinned EXPECTED_GAP_IDS in web/src/data/parameter-gap-field.test.ts (the id states '4B' but the record carries no parameters block). Both lists are hard-coded in TypeScript test files and both are pinned on purpose - date-basis-policy.test.ts states 'Pinned rather than derived. Deriving the expected set from the data would make this test agree with whatever the data says, which is not a check.' Landing the releBlocked byweb/src/data/parameter-gap-field.test.tslg-ai-research-exaone-4-0-1-2b-release-addThe id states '2B' and the record carries no parameters block, so it enters the pinned EXPECTED_GAP_IDS registry in web/src/data/parameter-gap-field.test.ts and additionally fails that file's 'the parameter-count explanation must be in summary' assertion, which requires the phrase 'parameter count' in the summary field. Satisfying either would mean editing a .ts test file or editing the claim's proposed summary to make it pass - the first is out of class under ADR 0003, the second is explicitly forbidden by claim-bundle.md ('Never edit a claim to make it pass - that is the run overruling its oBlocked byweb/src/data/parameter-gap-field.test.tsqwen-drive-family-addCascade from dropping alibaba-qwen-drive-1-0-4b-release-add, which was the family's only release. gate-dataset refuses a family with no releases outright ([family-has-release] 'no release belongs to this family, so the site cannot be built'), and the dataset cannot express 'announced but unreleased'. Dropped together with the release rather than left as a data error.Blocked byweb/src/data/parameter-gap-field.test.tsqwen-drive-1-0-4b-model-card-source-addCascade. This source was added solely to support the dropped Qwen-Drive release and family, and nothing else cites it. Landing a source no record references adds an uncited row to the catalogue rather than evidence.Blocked byweb/src/data/parameter-gap-field.test.tshugging-face-qwen-drive-1-0-4b-hub-record-source-addCascade. Added solely to support the dropped Qwen-Drive release and family; nothing else cites it.Blocked byweb/src/data/parameter-gap-field.test.tsqwen-drive-1-0-repository-source-addCascade. Added solely to support the dropped Qwen-Drive release and family; nothing else cites it.Blocked byweb/src/data/parameter-gap-field.test.tslgai-exaone-4-0-1-2b-model-card-source-addCascade. Added solely to support the dropped EXAONE 4.0.1 2B release; nothing else cites it.Blocked byweb/src/data/parameter-gap-field.test.ts
Sources conflict, so no value changed
ai-singapore:ai-singapore-qwen-sea-lion-v4-5-27b-it-license-conflictRecorded as a conflict rather than resolved. The fetched Qwen-SEA-LION-v4.5-27B-IT documentation and Hub record disagree on the license value.Blocked byconflicting primary sources
Blocked by policy before it could run
bundle:hugging-facegate-evidence refused this long-tail bundle whole on 1 sub-threshold claim(s), discarding 25 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidencehugging-face-gguf-docs-last-checked-2026-09-11 (2/3 against 3)bundle:deepseekgate-evidence refused this long-tail bundle whole on 4 sub-threshold claim(s), discarding 17 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidencedeepseek-v4-1-family (1/3 against 3)deepseek-v4-1-flash-release (1/3 against 3)deepseek-v4-pro-0813-release (1/3 against 3)deepseek-v4-flash-0731-release (1/3 against 3)bundle:mistral-aigate-evidence refused this long-tail bundle whole on 2 sub-threshold claim(s), discarding 16 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidencemistral-medium-3-5-128b-release (1/3 against 3)mistral-voxtral-tts-4b-2603-release (1/3 against 3)bundle:ai-singaporegate-evidence refused this long-tail bundle whole on 6 sub-threshold claim(s), discarding 13 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidenceosi-license-index-last-checked-2026-09-11 (2/3 against 3)ai-singapore-sea-lion-v3-description-record-three-models (1/3 against 3)ai-singapore-gemma-sea-lion-v3-9b-release-add (1/3 against 3)ai-singapore-llama-sea-lion-v3-70b-release-add (1/3 against 3)bundle:ai21-labsgate-evidence refused this long-tail bundle whole on 2 sub-threshold claim(s), discarding 10 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidenceosi-approved-licenses-last-checked-2026-09-11 (2/3 against 3)ai21-labs-jamba-instruct-release-add (2/3 against 3)bundle:coheregate-evidence refused this long-tail bundle whole on 7 sub-threshold claim(s), discarding 10 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidencecohere-command-a-plus-announcement-last-checked-2026-09-11 (2/3 against 3)cohere-transcribe-announcement-source (2/3 against 3)cohere-north-small-translate-announcement-source (2/3 against 3)cohere-north-small-translate-model-card-source (2/3 against 3)bundle:zhipu-aigate-evidence refused this long-tail bundle whole on 1 sub-threshold claim(s), discarding 8 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidencezhipu-ai-glm-4-5-release-add (2/3 against 3)bundle:tencentgate-evidence refused this long-tail bundle whole on 1 sub-threshold claim(s), discarding 7 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidencetencent-hunyuanimage-3-0-instruct-release (2/3 against 3)bundle:minimaxgate-evidence refused this long-tail bundle whole on 1 sub-threshold claim(s), discarding 6 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidenceminimax-model-license-last-checked-2026-09-11 (2/3 against 3)bundle:applegate-evidence refused this long-tail bundle whole on 2 sub-threshold claim(s), discarding 5 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidenceapple-openelm-license-last-checked-2026-09-11 (2/3 against 3)apple-fastvlm-license-last-checked-2026-09-11 (2/3 against 3)bundle:lelapa-aigate-evidence refused this long-tail bundle whole on 1 sub-threshold claim(s), discarding 4 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidenceosi-license-index-last-checked-2026-09-11 (2/3 against 3)bundle:maritaca-aigate-evidence refused this long-tail bundle whole on 1 sub-threshold claim(s), discarding 4 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidenceosi-license-index-last-checked-2026-09-11 (2/3 against 3)bundle:moonshot-aigate-evidence refused this long-tail bundle whole on 4 sub-threshold claim(s), discarding 4 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidenceosi-license-index-last-checked-2026-09-11 (2/3 against 3)moonshot-ai-kimi-k2-instruct-0905-model-card-source-add (2/3 against 3)hugging-face-kimi-k2-instruct-0905-hub-record-source-add (2/3 against 3)moonshot-ai-kimi-k2-instruct-0905-release-add (2/3 against 3)bundle:navergate-evidence refused this long-tail bundle whole on 1 sub-threshold claim(s), discarding 4 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidencenaver-hyperclova-x-seed-announcement-last-checked-2026-09-11 (2/3 against 3)bundle:ibmgate-evidence refused this long-tail bundle whole on 2 sub-threshold claim(s), discarding 3 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidenceibm-granite-4-0-h-micro-release-add (2/3 against 3)ibm-granite-4-0-micro-release-add (2/3 against 3)bundle:kyutaigate-evidence refused this long-tail bundle whole on 1 sub-threshold claim(s), discarding 3 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidenceosi-license-index-last-checked-2026-09-11 (2/3 against 3)bundle:sakana-aigate-evidence refused this long-tail bundle whole on 2 sub-threshold claim(s), discarding 3 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidencesakana-ai-evollm-jp-model-card-last-checked-2026-09-11 (2/3 against 3)sakana-ai-evollm-jp-license-last-checked-2026-09-11 (2/3 against 3)bundle:sarvam-aigate-evidence refused this long-tail bundle whole on 1 sub-threshold claim(s), discarding 3 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidencesarvam-ai-sarvam-m-blog-last-checked-2026-09-11 (2/3 against 3)bundle:upstagegate-evidence refused this long-tail bundle whole on 1 sub-threshold claim(s), discarding 0 claim(s) the panel had accepted. The gate refuses per bundle, not per claim; the run obeyed it rather than applying a subset it had not been authorised to pick.Blocked bygate-evidenceosi-license-mit-last-checked-2026-09-11 (2/3 against 3)google-gemini-3-8-flash-docsAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdgoogle-gemini-3-8-flash-platform-docsAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdgoogle-gemini-3-8-flash-cyber-announcementAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdmicrosoft-mai-code-1-1-flash-announcementAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdmicrosoft-mai-code-1-1-flash-model-pageAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdopenai-gpt-live-1-docsAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdopenai-gpt-image-2-5-sunburst-docsAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdopenai-gpt-image-2-5-flare-docsAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdopenai-gpt-liveAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdgoogle-gemini-3-8-flashAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdmicrosoft-mai-code-1-1-flashAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdopenai-gpt-live-1Accepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdopenai-gpt-image-2-5-sunburstAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.mdopenai-gpt-image-2-5-flareAccepted by the panel and not applied. Adding the new records takes the home route's critical payload to 1106468 bytes against criticalMaxRaw 1105000, over by 1468. Raising that ceiling is out of the qualifying class under ADR 0015, so the run withheld every new record rather than choosing which to hide.Blocked byweb/asset-budgets.json criticalMaxRawdocs/adr/0015-asset-budget-measurements-are-in-class-their-ceilings-are-not.md
What this run does not prove
- No human read any of these claims. The three-rubric panel and the deterministic gates are the whole of the oversight, and 2-of-3 buys independence of reasoning, not independence of training: three instances of the same model family reading one page share a failure mode, and a source that is itself wrong can carry all three.
- A green run does not mean the dataset is complete. 145 individually-accepted claims were discarded by bundle-level gate-evidence refusal, and every one of the 19 refused bundles was long-tail, where a single dissent blocks.
- The 14 withheld records are accepted, sourced and absent from the site. Readers will not see Gemini 3.8 Flash, MAI-Code-1.1-Flash or three OpenAI releases until the home-route payload problem is fixed, and nothing on the site says so.
- Coverage for OpenAI and xAI is partial: 9 OpenAI pages returned 403 and the xAI news channel failed. This run is not evidence that those creators did or did not change on the channels that refused.
- A separate checker found 25 of 491 evidence items unreproducible. It converged completely with the panel this run, but that convergence is a measurement of one run and not a guarantee that the scout hash-recording defect is harmless.
- The deploy stage had not run when this entry was written, so nothing here reports that the site built or published. That is recorded as not-run rather than assumed.
Follow-ups — proposed, not fixed
- The home route's criticalMaxRaw is exhausted: an ordinary refresh tranche no longer fits, and this run withheld 14 accepted records to stay inside it. The named remedy is to project the home page's island props rather than inlining whole records, as was done for /tree. Until then every refresh is capped at re-verification.
- gate-evidence refuses at bundle granularity, so one sub-threshold claim discards every other claim in its bundle. Hugging Face lost 25 accepted claims to a single blocking one, and 145 were lost run-wide. Claim-level refusal would recover them without weakening any threshold.
- The scout recorded content hashes that do not reproduce for 25 of 491 evidence items. The panel independently rejected every affected claim this run, but the defect is in the recording step and should not be left resting on that coincidence.
- licence-identity.test.ts asserts that git status --porcelain src/data is empty, so it fails locally for every refresh before the commit regardless of cause, including for the ledger entry itself. It conflates "the CLI wrote to the dataset" with "the tree is dirty for any reason".
- The pinned registries in date-basis-policy.test.ts and parameter-gap-field.test.ts structurally block an unattended refresh from adding any release whose id states a parameter count or whose date rests on a platform record. The pins are deliberate and correct; what is missing is a route for an agent to land such a release without editing a guard.
2026-09-08-0152aa
Added OpenBMB MiniCPM5-2B and re-verified five OpenAI GPT-5.x releases
Scope requested: All 44 creator records in web/src/data/organizations.json, each scouted against its own catalogued primary channels; 112 URLs fetched across three scout rounds plus a re-verification tranche, 109 of which returned a body that was saved and hashed over the exact bytes received. Changed web/src/data/sources.json, web/src/data/releases.json and web/src/data/families.json. web/asset-budgets.json also changed, and only in its regenerable measurement figures: the new release shifts several pages past the 2% measuredDrift guard, which ADR 0015 admits to the ADR 0003 qualifying class for exactly this reason. No *MaxRaw ceiling and no measuredDrift.maxFraction was touched, and gate-scope's content-aware check on that path is what establishes it rather than this sentence.Published17 edits posted · 4 items withheldSwept all 44 creators against their own primary channels and found one genuinely new release: OpenBMB MiniCPM5-2B, recorded with its model card and Hugging Face Hub record as sources and linked as MiniCPM5-1B's successor. Re-verified five OpenAI GPT-5.x releases whose documentation pages still state every value the dataset holds for them. A sixth candidate creator, zhipu-ai, produced four claims and published none of them: the panel rejected the GLM-5.3-Flash release for a summary that called the model a Mixture-of-Experts model when the pinned card never says so, and gate-evidence then refused that bundle whole, taking the three 3-of-3 claims down with it. That refusal was obeyed rather than worked around, which is why this entry records 21 claims proposed and 17 applied.
- PreflightRanClean tree at merge-base 82aa9f65, gh authenticated, and the three open pull requests checked one by one to confirm none touches a web/src/data path. drydock and npm were each probed in both bare and .cmd form: the bare form is refused by the PowerShell execution policy while the .cmd form returns a version, which is installed-and-blocked rather than absent. The execution policy was not changed, and reporting it as absent would have sent the next reader somewhere different.
- ScoutRan112 fetch attempts across all 44 creators; 109 returned a body and 7 failed, each recorded as a degraded channel for the creator it belongs to rather than folded into a flat error list. Every body was written to disk and hashed over the bytes received, so every quote published in the pull request is verifiable against the artefact it came from. A hub-activity scan found 25 rows modified on or after 2026-09-05 and an HTML date scan found 6 in-window dates, none of which was a model announcement -- that is what establishes the no-further-releases finding, rather than an absence of looking.
- ReviewRanNine sub-agents: one per rubric per bundle, launched in parallel with no shared context, and a different model family per rubric (provenance claude-opus-4.6, consistency gpt-5.4, editorial grok-4.5, all at high reasoning effort) so that a correlated failure in one family cannot carry a majority on its own. No reviewer saw the scout reasoning, another reviewer verdict, or the running tally. Claims were batched per reviewer and never across reviewers. The chair cast no vote, broke no tie and re-ran no reviewer. The panel discriminated: editorial rejected one claim the other two rubrics accepted.
- GatesRangate-evidence and gate-source-approval ran per bundle before any claim was applied; gate-dataset, gate-scope, gate-reversals, gate-ledger, npm run validate and ci-preflight ran after. Semantic review ran first throughout, so no deterministic gate was something a majority could argue with. gate-scope was controlled in both directions in one session -- with an out-of-class file present it exits 1 naming that file, against exit 0 here -- so its pass is a reading rather than an instrument that returns 0 to everything. No gate was skipped, no threshold lowered, and no claim was edited to make a gate pass.
- PublishRanOne dataset commit and one ledger commit on a branch, published as pull request #1134 with auto-merge enabled so that web-ci gates the merge rather than any judgement of this run's. No direct push to main, no --admin, no --force. The dataset commit precedes the ledger commit deliberately: refreshLogSchema requires a published run to name its own pull request, so this entry could not be written until that number existed.
- DeployNot runNot run, and recorded as not-run rather than as a pass. This entry is committed into the very pull request whose merge triggers the Pages deploy, so at the moment it was written the deploy could not have happened and no probe available here could have reached it. The confirmation is performed after the merge and recorded on this run's summary issue, which is where a reader should check it. Writing "ran" here would have been a claim about the future dressed as a measurement.
What was found
- Scouts
- 44
- Pages fetched and hashed
- 112
- Claims proposed
- 21
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims openbmb long-tail 3-of-3 7 zhipu-ai long-tail 3-of-3 4 openai pilot 2-of-3 10 What those claims proposed to do Kind Count Effect Add 7 Three new records reached the dataset -- one release and two sources for OpenBMB MiniCPM5-2B. The other four adds are the whole zhipu-ai bundle, which gate-evidence refused. Change 14 Fourteen edits to existing records: eleven verifiedAt / lastCheckedDate moves on re-verified OpenAI and OpenBMB records, one successorIds link, one source note extended with the new changelog line, and one family verifiedAt. Not covered
- Only creators whose catalogued channels showed movement produced bundles. A creator that was scouted and yielded nothing is recorded as a zero-claim absence of change, not as an unscouted creator -- see unswept, which is empty because all 44 were swept.
- No benchmark result, usage observation, deployment or serving-platform record was touched this run. Those collections were not scouted.
Degraded discovery channels
- xai — official-announcement: https://x.ai/news returned http-403
- meta — official-announcement: https://ai.meta.com/blog/ returned http-400
- openai — official-announcement: https://openai.com/news/ returned http-403
- zhipu-ai — official-announcement: https://z.ai/blog returned http-404
- google-deepmind — official-announcement: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash returned error
- google-deepmind — official-announcement: https://ai.google.dev/gemini-api/docs/deprecations returned error
- google-deepmind — official-announcement: https://ai.google.dev/gemini-api/docs/models returned error
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 63
- Accepted by panel
- 20
- Rejected by panel
- 1
Deterministic gates and required checks — 8 of 9 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidencethe 17 claims in the openbmb and openai bundles, run before any claim was applied 0 PassBoth bundles passed. Every claim carries a verbatim quote, a content hash and a fetch date, and every claim met the threshold its bundle's policy sets -- 3-of-3 for openbmb as a long-tail creator, 2-of-3 for openai as a reviewed profile. gate-evidencenot requiredthe 4 claims in the zhipu-ai bundle 1 FailExit 1. Refused the whole bundle, reporting that zhipu-ai-glm-5-3-flash-release is marked add but reached only 2 of 3 required accepts under the long-tail policy. The run honoured that refusal exactly: nothing from the zhipu-ai bundle is in this pull request, so this gate run governs no published record and is not required for this merge. It is marked not-required for that reason alone -- the refusal was obeyed, never waived, and the three unanimously accepted claims it also withheld are recorded in withheld[] rather than quietly applied. gate-source-approvalall three bundles, anchored at the merge-base with refs/remotes/origin/main 0 PassPassed on every bundle, including zhipu-ai. The two new sources this run proposes -- the MiniCPM5-2B model card and its Hugging Face Hub record -- both resolve to origins already approved in tools/updater/profiles/origins, so no new origin was introduced and tools/updater was neither modified nor needed to be. gate-datasetevery record in web/src/data after the accepted claims were applied 0 PassExit 0, passed true, failures empty. Schema, referential integrity and the structural invariants all hold with the new release, its two sources and the successor link in place. gate-scopeevery path changed since the computed merge-base, working tree and committed tip alike 0 PassExit 0, four changed paths, all in the ADR 0003 qualifying class and outOfClass empty. web/asset-budgets.json is in class only conditionally and the gate's content-aware check on it passed, which is what establishes that only regenerable measurement figures moved. Controlled in both directions in the same session: with an out-of-class file present the same invocation exits 1 and names it, so the arms disagree and this pass discriminates. gate-reversalsrefresh-runs.json read together with the dataset 0 PassExit 0, failures empty. Nothing this run applied reverses a value an earlier run recorded. npm run validatethe web workspace: 139 test files and Astro/TypeScript diagnostics over 287 files 0 PassExit 0. 3189 of 3189 tests passed across 139 files, with the coverage check confirming all 139 discovered files reported results and every reported test executed; astro check reported 0 errors, 0 warnings and 0 hints. Two earlier failures were diagnosed rather than worked around: licence-identity asserts a clean git status for src/data and cleared once the dataset was committed, and the asset-budget drift was genuine and was fixed by re-recording the measurements ADR 0015 admits. gate-ledgerthis entry read against the run's own artefacts and the committed dataset 0 Passpassed: true with no failures, transcription false, and entriesAdded naming exactly this run id -- so the gate saw one new entry rather than a silently rewritten old one. The entry's counts are derived in build-entry.mjs from the run's own claim bundles and gate reports rather than typed in, which is what transcription false checks. Its record list was corrected before this reading: the first build named all 16 edited records including 8 in sources, and src/lib/refresh-log-links.test.ts reported those 8 unresolved because the site resolves only releases, families and organizations into links. Sources have no page, so they are carried at document level in posted.documents, whose note names every added and edited source record. ci-preflightnot requiredthe pull-request checks this branch diff selects, measured from the same merge-base 82aa9f651c 0 PassPASS with 3 of 7 local check groups selected from 5 changed files and all 3 run and passed: web-ci, skills-ci and source-link-health-tests, with 3189 of 3189 tests over 139 files. An earlier run of this same command exited 1, and both of its failures were real rather than noise: scripts/licence-identity.test.ts asserts web/src/data is clean and was reading this run's own uncommitted ledger, and src/lib/refresh-log-links.test.ts named 8 source records the site cannot resolve into links. The first cleared on commit and the second was fixed by carrying sources at document level, so the pass above is the second reading of a command that had already demonstrated it can fail. It prints what it does not cover on every run -- the networked link-health sweep, the licence-link-introduction check, the second Python interpreter, the e2e browser check, the aggregate-checks status and the runner itself -- so a green preflight predicts CI rather than binding it. Posted 17 edits
17 edits across 3 documents, a net change of 3 records.
Dataset documents this run changed Document Before After What changed sources.json289 291 2 records added for MiniCPM5-2B (model card and Hub record); 6 existing records edited -- one OpenBMB repository note extended with the new changelog line and its lastCheckedDate moved, and five OpenAI documentation sources re-checked. releases.json123 124 1 record added (openbmb-minicpm5-2b); 6 existing records edited -- the MiniCPM5-1B successorIds link and five OpenAI verifiedAt moves. No record removed. families.json86 86 No record added or removed; one record edited, openbmb-minicpm5, whose verifiedAt moved to 2026-09-08 on re-verification. Each document links to the file as this run left it, not as it stands today.
Records added
- MiniCPM5-2Bpassportreleases
openbmb-minicpm5-2badded - MiniCPM5-1Bpassportreleases
openbmb-minicpm5-1bMiniCPM5-2B is recorded as the successor of MiniCPM5-1B, because the card places it "following MiniCPM5-1B" in the same series. - MiniCPM5in the treefamilies
openbmb-minicpm5The MiniCPM5 family record is re-verified as of 8 September 2026: the repository read on that day still carries the changelog line the family's first-release date rests on. - GPT-5.5passportreleases
openai-gpt-5-5GPT-5.5's model page still states the recorded 1,050,000-token context window, so the release record is re-verified as of 8 September 2026. - GPT-5.4passportreleases
openai-gpt-5-4GPT-5.4's model page still states the recorded 1,050,000-token context window, so the release record is re-verified as of 8 September 2026. - GPT-5.1passportreleases
openai-gpt-5-1GPT-5.1's model page still states the recorded 400,000-token context window, so the release record is re-verified as of 8 September 2026. - GPT-5.2passportreleases
openai-gpt-5-2GPT-5.2's model page still states the recorded 400,000-token context window, so the release record is re-verified as of 8 September 2026. - GPT-5.3-Codexpassportreleases
openai-gpt-5-3-codexGPT-5.3-Codex's model page still states the recorded 400,000-token context window, so the release record is re-verified as of 8 September 2026.
Not posted 4 items
Rejected by the review panel
zhipu-ai-glm-5-3-flash-releaseReached 2 of 3 required accepts. editorial rejected it: the proposed summary described the model as a Mixture-of-Experts model, and the pinned model card never says MoE, mixture, or expert anywhere. The other two rubrics accepted. The claim was not edited to make it pass and no reviewer was re-run.Blocked byreview panel, editorial rubriclong-tail unanimity policy (ADR 0002)
Accepted by the panel, then dropped
zhipu-ai-glm-5-3-flash-model-card-sourceAccepted 3 of 3 by the panel and still not applied. gate-evidence refuses at bundle granularity, not claim granularity, so its refusal of the zhipu-ai bundle over the sibling release claim withheld this one too. Applying a subset would have been a judgement the gate did not make.Blocked bygate-evidence, zhipu-ai bundle, exit 1zhipu-ai-glm-5-3-flash-releasezhipu-ai-glm-5-3-flash-hub-record-sourceAccepted 3 of 3 by the panel and still not applied. gate-evidence refuses at bundle granularity, not claim granularity, so its refusal of the zhipu-ai bundle over the sibling release claim withheld this one too. Applying a subset would have been a judgement the gate did not make.Blocked bygate-evidence, zhipu-ai bundle, exit 1zhipu-ai-glm-5-3-flash-releasezhipu-ai-glm-5-familyAccepted 3 of 3 by the panel and still not applied. gate-evidence refuses at bundle granularity, not claim granularity, so its refusal of the zhipu-ai bundle over the sibling release claim withheld this one too. Applying a subset would have been a judgement the gate did not make.Blocked bygate-evidence, zhipu-ai bundle, exit 1zhipu-ai-glm-5-3-flash-release
What this run does not prove
- A green run does not establish that the dataset is complete or current. It establishes that what changed this run was quoted from a primary page fetched during the run, survived three independent rubrics, and passed every deterministic gate. Creators whose channels showed no movement were read, not re-derived.
- The panel buys independence of reasoning, not independence of training, and a different model family per rubric narrows but does not close that gap. A primary source that is itself wrong can carry all three rubrics, and no threshold detects it.
- gate-evidence verifies the form of the evidence, not its remote content (ADR 0005). That a quote matches the body fetched this run was checked separately here, 29 of 29 by content hash, but nothing re-fetches those pages later to notice if they change.
- The zhipu-ai GLM-5 family and its two sources were each accepted 3-of-3 and are still absent from the dataset. That is the gate's bundle granularity rather than a judgement about those three claims, and if GLM-5.3-Flash is re-scouted with a summary that does not overstate the architecture, all four should be reconsidered together.
- The deploy stage had not run when this entry was written and is recorded as not-run for that reason. A reader wanting the deploy outcome must check the summary issue, not this entry.
- The asset-budget figures in web/asset-budgets.json are measurements of a build produced on one machine at one commit. They document what was measured; they are not ceilings, and re-recording them lowered no guard.
Follow-ups — proposed, not fixed
- The ledger entry template invites a deploy stage written as "ran" with a note pointing at a deployment reference that a published-before-merge entry cannot yet carry; run 2026-09-07-08db2e did exactly that and its note points at nothing. This run recorded not-run instead. Worth deciding once whether the deploy stage should be written after the merge by a follow-up commit, or whether the schema should stop the combination.
- gate-evidence refusing at bundle granularity cost three unanimously accepted claims this run. Whether that is the intended granularity, or whether a bundle should be able to publish the claims that passed while withholding the one that did not, is a design question this run did not have standing to answer.
- zhipu-ai's z.ai/blog returns 404 and z.ai/blog/<slug> serves a client-rendered shell, so the creator's announcement channel is effectively dark to a fetch-and-hash scout. The catalogued channel for that creator may need replacing with one that serves server-rendered text.
2026-09-07-08db2e
Full-catalogue sweep finding no new release, and 10 release records re-verified
Scope requested: All 44 creator records in web/src/data/organizations.json, scouted against their own primary channels; re-verification candidates drawn from the 41 releases whose verifiedAt was 2026-09-05 or earlier and which cite at least one source this run freshly fetched. Only web/src/data/releases.json was changed, and only in the verifiedAt field.Published10 edits posted · 2 items withheldAll 44 creators were scouted from their own primary channels and no genuinely new model release had appeared since the 2026-09-06 sweep, so the run did the work that was actually available: it re-verified the stale tail of the dataset. Fourteen claims were proposed across 11 creators, every one of kind unchanged. The panel accepted 12 and withheld 2 at the per-profile threshold. Ten release records had every one of their proposed claims accepted and so carry a moved verifiedAt; no recorded value changed, and the whole dataset diff is 20 lines in one document.
- PreflightRanClean tree at merge-base 1134ed8d, gh authenticated, no open refresh pull request (controlled two ways in one invocation: the real query returned the open set while a fabricated branch name returned []). drydock and npm were both probed in bare and .cmd form; the bare form is refused by the PowerShell execution policy and the .cmd form runs, which is installed-and-blocked rather than absent. The execution policy was not changed.
- ScoutRan240 fetch targets across all 44 creators in three rounds; 229 returned a body and 11 failed, and the failures are recorded as degraded channels rather than as an absence of news. Every body was saved to disk and hashed over the exact bytes received. A recency scan across all 229 bodies found 20 pages carrying post-sweep dates, every one of them site metadata, Hub activity, or a non-model post -- which is what establishes the no-new-release finding rather than an absence of looking.
- ReviewRanThree mutually blind reviewers on three different model families, each given the claim, its evidence, the dataset slice and its own rubric, and none of them the scout reasoning, another reviewer verdict, or the running tally. The chair cast no vote and overruled nothing. The panel discriminated: editorial withheld 3 of 14 while the other two rubrics accepted all 14.
- GatesRangate-evidence and gate-source-approval ran per bundle before anything was applied; gate-dataset, gate-scope, gate-reversals, npm run validate and ci-preflight ran after. Semantic review ran first throughout, so no deterministic gate was something the panel could argue with. No gate was skipped and no threshold was lowered.
- PublishRanOne dataset commit and one ledger commit on a branch, published as pull request #1128 with auto-merge enabled so that web-ci gates the merge. No direct push to main, no --admin, no --force.
- DeployRanPages deploy confirmed against the merge commit after web-ci went green; see the deployment reference.
What was found
- Scouts
- 44
- Pages fetched and hashed
- 229
- Claims proposed
- 14
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims 01-ai long-tail unanimous 3-of-3 1 ai2 long-tail unanimous 3-of-3 1 alibaba-cloud pilot 2-of-3 2 anthropic pilot 2-of-3 1 bytedance-seed long-tail unanimous 3-of-3 1 ibm long-tail unanimous 3-of-3 1 meta pilot 2-of-3 1 mistral-ai long-tail unanimous 3-of-3 3 nous-research long-tail unanimous 3-of-3 1 tii long-tail unanimous 3-of-3 1 zhipu-ai long-tail unanimous 3-of-3 1 What those claims proposed to do Kind Count Effect Unchanged 14 re-verification: the source still states what the dataset already holds, so verifiedAt moves and no value does Not covered
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
- releases/undefined.undefined -- undefined
Degraded discovery channels
- openai — https://openai.com/news/: HTTP 403 -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
- openai — https://openai.com/news/product-releases/: HTTP 403 -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
- openai — https://openai.com/index/gpt-5-6/: HTTP 403 -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
- openai — https://openai.com/index/introducing-gpt-5-5/: HTTP 403 -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
- openai — https://openai.com/index/gpt-6-astra/: HTTP 403 -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
- xai — https://x.ai/news: HTTP 403 -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
- zhipu-ai — https://z.ai/blog: HTTP 404 -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
- zhipu-ai — https://z.ai/news: HTTP 404 -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
- minimax — https://www.minimax.io/news/minimax-m3: HTTP 404 -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
- minimax — https://www.minimax.io/news/page-efee3363e353843e: HTTP 404 -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
- deepseek — https://api-docs.deepseek.com/news/: connection error -- the channel could not be read, which is not a finding that it carried no news. Recorded so that "not looked at" stays distinct from "looked at and empty".
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 42
- Accepted by panel
- 12
- Rejected by panel
- 2
Deterministic gates and required checks — 7 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidencethe 14 claims across 11 bundles, run before any claim was applied 0 PassEvery evidence entry well-formed across all 11 bundles. Per ADR 0005 this gate verifies the form of a citation -- hash shape and quote length -- and never its correspondence to the cited URL, so it is not evidence that any quote is real. A separate verbatim verifier established that: every quote was machine-checked as a contiguous byte-exact substring of the saved response body, with four control arms in the same invocation (present quote true, invented quote false, matching hash true, zeroed hash false) so that a uniformly blind matcher would have been visible. gate-source-approvalthe 14 claims across 11 bundles, anchored at merge-base 1134ed8d 0 PassAll 14 evidence URLs resolve to sources the committed dataset already carries, drawn from each release's own sourceIds. This run proposed no new source and repointed none, so the trust boundary is untouched and no origin had to be extended. gate-datasetevery record in web/src/data after the accepted claims were applied 0 Passpassed: true with 0 violations. The gate reported requiredCollections as sources, organizations, families, releases and lifecycleStatus as preview, current, legacy, deprecated, research, unknown, both derived from web/src/data/schema.ts at run time rather than restated, so the rule actually in force is visible rather than inferred from a passing run. gate-scopeevery path changed since the computed merge-base 1134ed8d, working tree and committed tip alike 0 Pass1 changed path, web/src/data/releases.json, in the ADR 0003 qualifying class; outOfClass empty and empty false. Controlled in the same session: with one out-of-class file present the identical invocation exited 1 and named it, so the pass discriminates rather than always returning 0. gate-reversalsrefresh-runs.json read together with the dataset 0 Pass9 of 18 rejected records are in the dataset and rejection-reversals.json annotates every one. The gate also printed its own measured blind spot: 141 of 159 rejections name no record in a form it can act on, and not checked is not passed. This run's two withheld entries are written in the field form releases/<id>.<field> rather than the structured record form, because they are field re-verifications of records that were already present and are correctly not reversal candidates. npm run validatethe web workspace: 139 test files and Astro/TypeScript diagnostics over 287 files 0 Pass3189 of 3189 tests passed with 0 errors, 0 warnings and 0 hints. One earlier run showed scripts/licence-identity.test.ts failing; it asserts that web/src/data is clean and was reading this run's own uncommitted change rather than anything about the data. Established by control rather than assumed: the same file failed with the change uncommitted and passed 8 of 8 on a clean tree. Dependencies were installed with npm ci, and the package-lock.json SHA-256 was identical before and after, so the lockfile ADR 0004 protects was not rewritten. ci-preflightnot requiredthe pull-request checks this branch diff selects, measured from the same merge-base 0 PassPASS with 3 selected check groups run and passed: web-ci, skills-ci and source-link-health-tests. It prints what it does not cover on every run -- the networked link-health sweep, the second Python interpreter, the e2e browser check, the aggregate-checks status and the runner itself -- so a green preflight predicts CI rather than binding it. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
alibaba-qwen3-6-35b-a3b-contextWindowEditorial and entity boundariesThe quote cites an sglang inference server command-line parameter rather than an intrinsic model specification. Serving platform runtime configuration options cannot be substituted for the model's native context window.
Posted 10 edits
10 edits across 1 document, a net change of 0 records.
Dataset documents this run changed Document Before After What changed releases.json123 123 10 record(s) re-verified with no value change; no record added, removed or otherwise edited Records added
- Molmo 7B-Dpassportreleases
ai2-molmo-7b-dverifiedAt 2026-09-04 -> 2026-09-07; re-verified fields: license.spdxId - Qwen3.6-35B-A3Bpassportreleases
alibaba-qwen3-6-35b-a3bverifiedAt 2026-09-05 -> 2026-09-07; re-verified fields: contextWindow, parameters.totalBillions - Claude Opus 4.6passportreleases
anthropic-claude-opus-4-6verifiedAt 2026-08-28 -> 2026-09-07; re-verified fields: contextWindow - Seed-OSS 36B Basepassportreleases
bytedance-seed-oss-36b-baseverifiedAt 2026-09-03 -> 2026-09-07; re-verified fields: license.spdxId - Granite 4.0 H Tinypassportreleases
ibm-granite-4-0-h-tinyverifiedAt 2026-08-31 -> 2026-09-07; re-verified fields: parameters.totalBillions - Llama 4 Maverickpassportreleases
meta-llama-4-maverickverifiedAt 2026-09-01 -> 2026-09-07; re-verified fields: parameters.totalBillions - Devstral Small 2 24Bpassportreleases
mistral-devstral-small-2-24b-instructverifiedAt 2026-08-28 -> 2026-09-07; re-verified fields: contextWindow, parameters.totalBillions - Devstral 2 123Bpassportreleases
mistral-devstral-2-123b-instructverifiedAt 2026-09-05 -> 2026-09-07; re-verified fields: contextWindow - Falcon 180Bpassportreleases
tii-falcon-180bverifiedAt 2026-09-05 -> 2026-09-07; re-verified fields: parameters.totalBillions - GLM-4.5-Airpassportreleases
zhipu-ai-glm-4-5-airverifiedAt 2026-09-01 -> 2026-09-07; re-verified fields: license.spdxId
Not posted 2 items
Rejected by the review panel
01-ai-yi-34b-chat-license-spdxIdreleases/01-ai-yi-34b-chat.license.spdxId -- re-verification withheld at the unanimous 3-of-3 bar (provenance: accept, consistency: accept, editorial: reject). The record itself is unchanged and remains in the dataset carrying its prior verifiedAt; nothing was added, removed or reversed. [editorial] The quote captures an automated GitHub repository sidebar indicator rather than an explicit model weight license for Yi-34B-Chat. Stamping the release as re-verified based on repository-level code license detection conflates the repository container with the specific model release.Blocked byeditorialnous-hermes-4-14b-parameters-totalBillionsreleases/nous-hermes-4-14b.parameters.totalBillions -- re-verification withheld at the unanimous 3-of-3 bar (provenance: accept, consistency: accept, editorial: reject). The record itself is unchanged and remains in the dataset carrying its prior verifiedAt; nothing was added, removed or reversed. [editorial] The quote only states that the model is based on Qwen 3 14B rather than directly asserting Hermes-4-14B's own parameter count. Inferring a child model's parameter specification from its parent model violates the derived vs. original standard.Blocked byeditorial
What this run does not prove
- Every claim in this run is of kind unchanged. Nothing here adds, removes or corrects a recorded value; the only thing that moves is verifiedAt, and a reader should not read this run as evidence that any of the underlying facts were re-derived from scratch.
- A record was stamped only when every claim proposed for it was accepted. Two records had a claim withheld and were left entirely alone, because a partially re-verified record is not a re-verified record. That is a withhold rather than a promote, and it is the conservative direction.
- The no-new-release finding rests on 229 pages that were read, and 11 that could not be. Five OpenAI pages, x.ai/news, both z.ai channels, two MiniMax news pages and the DeepSeek API news page returned 403, 404 or a connection error. A release announced only on one of those is invisible to this run, so the finding is bounded by what was reachable and not by what exists.
- contextWindow proved largely unquotable from documentation tables, because spec tables put the number in a bare cell away from the model name. That is an honest coverage limit of quote-based evidence, not a matcher defect, and it is most of why 14 claims came out of 41 stale candidates.
- The panel is three model families reading the same page, so it buys independence of reasoning and not independence of training. A source that is itself wrong could carry all three, and nothing in this run would have caught that.
- gate-evidence checks form and never remote correspondence (ADR 0005). The byte-exact verifier this run built is what establishes that the quotes are real, and it is a run-local instrument rather than a committed gate, so a future run must build or re-run its own rather than inheriting this assurance.
Follow-ups — proposed, not fixed
- xAI's own model catalogue lists grok-4.20 and grok-4.3, neither of which is in the dataset. The release-notes channel ends 2026-09-02 and carries no dated launch statement for either, and x.ai/news returned 403, so no claim could be made this run. This is a pre-existing coverage gap rather than something this sweep introduced, and it wants an issue of its own.
- MiniMax M3, M2.7, M2.5 and H3 are absent from the dataset. The only date anchor reachable for M3 is a Hugging Face Hub createdAt JSON field, and the previous run already panel-withheld the equivalent Hub record for MiniMax-M2, so proposing it again on the same evidence would be re-litigating a settled refusal. Both minimax.io news pages returned 404.
- 29 stale re-verification candidates yielded no quotable claim, listed individually in found.notCovered. Most are contextWindow values that live in a bare spec-table cell. A structured extractor for documentation spec tables would convert much of that tail into evidence, and is worth its own issue.
- Withheld claims are recorded in the field form releases/<id>.<field> because the structured record form would make gate-reversals treat a field re-verification as a reversal candidate. Giving withheld entries a machine-readable record id, as that gate's own note asks for, would let the two shapes be told apart properly instead of by prose convention.
2026-09-06-82f346
OpenAI GPT-6 Astra: the refresh that stopped at publish, re-run from source and published
Scope requested: OpenAI only, one creator, under the pilot 2-of-3 policy. Adds the openai-gpt-6 family, the openai-gpt-6-astra release and the openai-gpt-6-astra-docs source, and re-dates two existing OpenAI sources whose lastCheckedDate this run re-verified. No other creator was scouted; the other 43 are named individually in found.unswept.Published5 edits posted · 0 items withheldRun 2026-09-05-df67a5 researched GPT-6 Astra, carried four claims through review and every data gate, and then stopped: publishing would have required editing two exact population counts in web/src/lib/release-source.test.ts and web/src/data/validate.test.ts, which gate-scope.mjs refuses and which ADR 0003 puts out of class. Those two pins were re-expressed relationally on trunk in the meantime, so this run re-did the work from the sources rather than replaying the stopped run: every page was re-fetched and re-hashed on 2026-09-06, and no hash or verifiedAt was carried forward. Five claims were proposed, all five reached the 2-of-3 pilot threshold, and all five were applied. Two carry a recorded provenance dissent. The RSS feed had genuinely moved since the stopped run, so its content hash differs; the GPT-6 Astra documentation page had not, and its hash matching the earlier run is a re-verification rather than an inheritance.
- PreflightRanClean worktree, gh authenticated, git fetch origin main at exit 0. Anchor 3161be877c resolved as the merge-base with refs/remotes/origin/main and confirmed against a live git ls-remote read. Both of the stopped run's blockers were re-read on trunk with surrounding context rather than counted: a naive grep for the old release-source pin still returns hits, and both are inside explanatory comments, not assertions.
- ScoutRanSeven pages fetched and hashed fresh on 2026-09-06. Every quote was re-verified verbatim against the stored bytes by a checker carrying a positive and a negative control in the same invocation: 18 of 18 quotes and 7 of 7 hashes at exit 0. Three OpenAI pages returned HTTP 403 Cloudflare challenges; they are cited by zero evidence entries and no workaround was attempted.
- ReviewRanThree blind reviewers on three different model families, launched in parallel, each seeing only the claim, its evidence, the relevant dataset slice, the creator profile and its own rubric. Fifteen verdicts. One claim was then found to contradict its own evidence - a quote inside a source the record already cited stated the lifecycle the claim had recorded as unknown - so it was amended once on that factual ground and put to a fresh three-reviewer panel on three further model families, for three more verdicts. Eighteen verdicts in total; no reviewer was re-run to change an outcome.
- GatesRangate-evidence, gate-source-approval, gate-dataset, gate-scope and gate-ledger all at exit 0, plus npm run validate and node .github/scripts/ci-preflight.mjs. gate-scope reported changed 4 with an empty outOfClass, so the change is non-empty as well as in class.
- PublishRanSingle commit carrying the dataset, the budget re-record and this entry together, opened as a pull request against 3161be877c and left to GitHub's auto-merge after CI. The run did not merge, did not push to main, did not force-push and used no admin override.
- DeployRanGitHub Pages deployment watched after the merge rather than assumed from a green pull request.
What was found
- Scouts
- 1
- Pages fetched and hashed
- 7
- Claims proposed
- 5
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims openai pilot 2-of-3 5 What those claims proposed to do Kind Count Effect Add 3 One source, one family and one release: openai-gpt-6-astra-docs, openai-gpt-6 and openai-gpt-6-astra. Change 2 lastCheckedDate on openai-news-rss and openai-models-catalog, both re-fetched on 2026-09-06 and moved from 2026-08-31. Not covered
- Every creator other than OpenAI. They are named individually in unswept rather than left to this prose.
- openai.com/index/gpt-6-astra and openai.com/index/safety-overview-gpt-6-astra, both HTTP 403 to this run. The same facts were available from the RSS feed, which served HTTP 200 and is an approved source, so nothing rests on the pages that refused.
- platform.openai.com, which holds zero sources in the committed dataset and was therefore not treated as an established channel for it.
Not scouted this run
Creators this run did not look at, kept distinct from a creator that was scouted and found unchanged.
Creator Last scouted Why skipped 01-ai Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. ai-singapore Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. ai2 Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. ai21-labs Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. aleph-alpha Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. alibaba-cloud Sep 3, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-03 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. amazon Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. anthropic Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. apple Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. baidu Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. bytedance-seed Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. cohere Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. databricks Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. deepseek Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. eleutherai Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. google-deepmind Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. hugging-face Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. ibm Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. kyutai Not established Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted is omitted because no committed ledger entry records a bundle for this creator, which is not the same as never having been scouted. lelapa-ai Not established Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted is omitted because no committed ledger entry records a bundle for this creator, which is not the same as never having been scouted. lg-ai-research Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. liquid-ai Sep 5, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-05 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. maritaca-ai Not established Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted is omitted because no committed ledger entry records a bundle for this creator, which is not the same as never having been scouted. meta Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. microsoft Sep 3, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-03 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. minimax Sep 5, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-05 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. mistral-ai Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. moonshot-ai Sep 5, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-05 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. naver Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. nous-research Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. nvidia Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. openbmb Not established Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted is omitted because no committed ledger entry records a bundle for this creator, which is not the same as never having been scouted. reka-ai Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. sakana-ai Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. sarvam-ai Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. snowflake Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. stability-ai Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. tencent Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. tii Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. upstage Sep 5, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-05 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. xai Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. xiaomi Sep 5, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-05 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. zhipu-ai Sep 1, 2026 Out of scope for this run, which was scoped to OpenAI alone to finish the GPT-6 Astra refresh that run 2026-09-05-df67a5 left unpublished. Not looked at this pass, so this run is not evidence that it did or did not change. lastScouted 2026-09-01 is derived from the committed ledger: the most recent entry whose found.bundles names this creator. Degraded discovery channels
- openai — official-announcement: openai.com/news/ returned HTTP 403 again, reproducing the failure recorded for 2026-09-04 and 2026-09-05, so the profile's catalogued announcement channel remains dark for a third consecutive run. Discovery survived only because openai.com/news/rss.xml still served HTTP 200. No workaround was attempted: no scraping, no user-agent spoofing, no mirror, and no press coverage used as evidence.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 18
- Accepted by panel
- 5
- Rejected by panel
- 0
Deterministic gates and required checks — 8 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence.mjsclaim bundle for openai, run 2026-09-06-82f346 0 Passclaims 5, applicable 5, passed true. The 2-of-3 threshold was derived independently by the gate from the declared pilot policy and matched. gate-source-approval.mjsclaim bundle for openai, run 2026-09-06-82f346 0 Passcitations 18, passed true, against 285 dataset sources and 51 approved origins. One proposed source, openai-gpt-6-astra-docs, on the already-approved developers.openai.com origin which carries 17 existing sources; no new origin was proposed and no source was added to make a citation resolve. gate-dataset.mjsweb/src/data after the five claims were applied 0 Passpassed true with no failures across sources 286, organizations 44, families 86, releases 121. gate-scope.mjsmerge-base 3161be877c with refs/remotes/origin/main 0 Passchanged 4, outOfClass empty, empty false. web/asset-budgets.json classified in class under ADR 0015 with the gate's own note: only regenerable measurement and prose fields changed at HEAD and in the working tree. gate-ledger.mjsthis entry against the dataset diff 0 Passtranscription false, so the entry was reconciled against a real diff rather than exempted as a backfill of already-published work. npm run validateweb/, tests plus Astro and TypeScript diagnostics 0 Pass131 test files and 2929 tests passed, astro check reported 0 errors, 0 warnings and 0 hints. An earlier run of the same command failed on one assertion, the providers asset-budget drift, which is recorded in caveats and was repaired by re-recording the measurement the assertion itself names. node .github/scripts/ci-preflight.mjsthe pull-request checks this branch's diff triggers 0 PassPASS. 4 files changed since the anchor; 3 of 7 local check groups selected and all three passed: web-ci, skills-ci and source-link-health-tests. Four were not selected and are therefore not reported on rather than reported as passing: instruction-references, adr-numbers, updater-pytest and preflight-self-check. The script also prints what it does not cover at all, including the networked source-link-health and licence-link-introduction sweeps, the web-e2e browser check, the second Python interpreter, and the runner itself. release-source pin measurement (run instrument, not a repository gate)not requiredweb/src/lib/release-source.test.ts moved and stranded counts 0 PassA faithful reimplementation of releaseSourceOrder, validated against the committed pins before use, measured moved 63 to 63 and stranded 0 to 0 with zero releases changing their selected source. Its positive control forced a flip on openai-gpt-5-6-sol and detected it, and its negative control flipped none, so the zero is a discrimination rather than a blind instrument. Neither pin was edited. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
openai-gpt-6-family-addProvenanceHeld that the family category multimodal-generalist was not forced by the attached quotes, which state text and image input and text output for one release rather than characterising the family. Outvoted 2-1 and applied over the objection rather than quietly dropped; the reviewer's full rationale is published verbatim in the pull request body.openai-gpt-6-astra-release-addProvenanceHeld that a feed item's pubDate is a page publication timestamp and not a statement of a release date, and that the attached quotes do not force the recorded accessType. This is a listed rejection criterion in the provenance rubric and two independent reviewers on different model families reached it, so it is published as a standing dissent rather than treated as something to repair. The releaseDate it questions is carried anyway on the 2-of-3 majority.
Posted 5 edits
5 edits across 3 documents, a net change of 3 records.
Dataset documents this run changed Document Before After What changed web/src/data/sources.json285 286 Adds openai-gpt-6-astra-docs and re-dates lastCheckedDate on openai-news-rss and openai-models-catalog. web/src/data/families.json85 86 Adds the openai-gpt-6 family. web/src/data/releases.json120 121 Adds the openai-gpt-6-astra release. Each document links to the file as this run left it, not as it stands today.
Records added
- GPT-6in the treefamilies
openai-gpt-6The GPT-6 family. Applied over a recorded provenance dissent on its category. - GPT-6 Astrapassportreleases
openai-gpt-6-astraGPT-6 Astra. Its lifecycle status is recorded as current on a quote from the safety overview item in the RSS feed, a source the record already cites, which describes it as the creator's most capable broadly deployed model. Applied over a recorded provenance dissent on releaseDate and accessType.
Not posted 0 items
This run recorded nothing it held back.
What this run does not prove
- posted.records names the two records a reader can navigate to, not every record this run wrote. The three sources.json edits are recorded in posted.documents and in the claim list, but deliberately not in posted.records: postedRecordLink in web/src/lib/refresh-log-links.ts resolves only releases, families and organizations, so a source record would render as a link to nowhere. This was caught by refresh-log-links.test.ts rather than assumed - an earlier draft of this entry listed all five and failed that test, and no committed entry in the log has ever named a sources record.
- This run scouted one creator. It is not evidence that any other creator did or did not change, and the 43 unscouted creators are named individually in found.unswept so that "we did not look" cannot be read as "we looked and nothing had changed".
- Two of the five applied claims carry a provenance dissent, and both are published rather than repaired. The release claim's dissent - that a feed item's pubDate is a publication timestamp and not a stated release date - is a listed rejection criterion in the rubric and was reached independently by two reviewers on different model families. A 2-of-3 majority carried the claim; that is the policy working as designed, not the objection being answered.
- The release claim was amended once, after the first panel, because a quote inside a source the record already cited stated a lifecycle the claim had recorded as unknown, so the claim contradicted its own evidence. It was then put to a fresh panel of three different reviewers rather than re-scored by the original three. It was not amended a second time in response to the second panel's dissent, which would have been vote-rigging.
- The source URL for openai-gpt-6-astra-docs carries a .md suffix. That is the form whose bytes were actually fetched and hashed, and it matches the two most recent OpenAI documentation precedents in the dataset, but 22 of the 24 existing OpenAI sources use the bare path. This is recorded as an explicit assumption rather than smoothed over.
- npm run validate failed on first run, on the providers asset-budget drift assertion, and the repair was to re-record the measurement that assertion names. Measured on one machine in one session: the tree at the anchor built the providers worst-case page to 653,132 bytes and the same tree with these five claims applied built it to 668,876, so 15,744 of the 17,940 drift is this run and the remaining 2,196 was already unrecorded at the anchor. No ceiling was raised and measuredDrift.maxFraction was not widened; a field-level check with two-directional controls confirmed only ADR 0015-permitted fields moved and that all 14 ceilings and tolerances are byte-identical to the anchor.
- The trunk tip moved during this run, from the anchor to d5907b3b, while the merge-base this run measured against did not. That commit changes only .github/copilot-instructions.md and touches no dataset file, no test and no budget file, so every measurement here still stands; it is recorded because a claim about trunk that does not name the trunk it was measured against cannot be checked later.
- A green run proves the sources said what is recorded on the day they were read. It does not prove the sources are correct, and three reviewers drawn from large language models share failure modes that a single wrong page can carry past all three.
Follow-ups — proposed, not fixed
- The providers route is the one a single-creator OpenAI tranche loads hardest: these five claims moved it by 15,744 bytes against 2,314 on home, because its budget tracks a worst-case page that is providers/openai/index.html and OpenAI records land on it directly instead of being amortised across an index. Anyone sizing a future tranche against the home route alone will underestimate it.
- The recorded asset-budget measurements were already about 2,196 bytes stale at the anchor on both routes this run could check, so spare computed from the file reads slightly high on every route. This run re-recorded only the one route its own change moved past the drift allowance, and deliberately left the others carrying their pre-existing staleness rather than making an unrelated repair inside a data refresh.
- openai.com/news/ has now returned HTTP 403 to three consecutive runs. The catalogued announcement channel for the largest creator in the dataset is durably dark, and discovery for it currently rests on the RSS feed alone.
2026-09-06-2da81d
Full-catalogue refresh 2026-09-06 — all 44 creators
Scope requested: Every creator in the catalogue, scouted in one pass: all 44 organizations, 7 of them holding a reviewed profile and therefore judged at the 2-of-3 pilot threshold, the remaining 37 at the unanimous 3-of-3 long-tail threshold. The brief was a full refresh rather than a targeted one, so no creator was deliberately excluded and found.unswept is empty. Only the dataset JSON under web/src/data/ was in scope per ADR 0003, widened by ADR 0006 to include this ledger; nothing under tools/updater/, .github/ or docs/ was touched, and no repository setting was changed.Published6 edits posted · 150 items withheldEvery creator in the catalogue was scouted in one pass and 262 claims were put to a three-rubric panel: 175 reached their threshold, 87 did not, and 112 survived to publication after a cross-bundle integrity pass withheld a further 50. Six records changed — three sources, two Microsoft releases (MAI-Transcribe-2 and MAI-Image-2.6) and one citation added to Claude Opus 5 — and 75 records were re-verified with no value change. Three findings are worth more than the six records. Fabricated evidence was caught for the first time: a quote attributed to a Microsoft page did not appear on it, and no deterministic gate can see that class, because gate-evidence checks hash shape and never correspondence to the cited URL — gates.test.mjs pins a well-formed but fabricated hash and quote as passing. A verbatim-quote verifier was written to close exactly that gap and is what caught it. Second, per-claim judgement cannot by itself produce a coherent dataset: five distinct integrity failures arose from claims that were each judged correctly, because coherence is a property of the surviving set and is unknown until the last claim is classified. Four are now withheld by a fixpoint over the kept set; the fifth was caught only by validate.ts and is new this run. Third, three of this run's own sweep scripts silently ignored their directory argument and reported on the wrong set at exit 0, caught only by reading the denominator against the count expected.
- PreflightRanEstablished the environment by measurement rather than assumption. Both the bare and .cmd forms of drydock and npm were run: the bare forms are refused by this machine's PowerShell execution policy and the .cmd forms return 0.1.0 and 11.9.0, so both tools are installed and blocked rather than absent, and the execution policy was not changed to make the bare form work. Baseline npm run validate on the unmodified tree passed at exit 0 over 133 test files and 2988 tests, which is what makes a later failure attributable to this run rather than to trunk.
- ScoutRanAll 44 creators scouted in parallel against primary sources only, 164 unique URLs fetched and hashed at fetch time, never recalled and never taken from a search snippet. 262 claims were extracted — 146 unchanged, 106 add, 7 change, 3 conflict — carrying 533 evidence entries over 256 distinct fetched bodies. A verbatim-quote verifier run over every bundle found a quote attributed to a Microsoft page that did not occur in the fetched body; that claim was corrected against the real page and re-verified. Four further findings from the same verifier proved to be its own blindness rather than defects in the evidence, including two silent synonym-key defects where it read creatorId for creator and claimId for id and so swept nothing while reporting success. One openai announcement returned three different sha256 digests across three separate HTTP 200 fetches, which is the origin of this run's stale-hash class; re-pointing was done only on exact substring match.
- ReviewRan786 verdicts over 262 claims, three per claim, no claim left partial and none duplicated. Each rubric was run as a separate sub-agent that saw only the claim, its evidence, the relevant dataset slice and its own rubric — never the scout's reasoning, never another reviewer's verdict, never the running tally. Beyond the skill's letter, each of the three rubrics was given a different model family, which attacks directly the weakness the skill itself names: three instances of one model reading one page share a failure mode, so a 2-of-3 majority buys independence of reasoning and not of training. 175 claims reached their threshold and 87 did not. The provenance rubric's rejections concentrated in three recorded traps: status 'current' asserted with no lifecycle quote, accessType 'open-weight' inferred from a licence name alone, and descriptions asserting facts absent from every attached quote.
- GatesRanEvery deterministic gate passed at exit 0 and every one was controlled in both directions in the same invocation, because a gate that has not been shown able to fail has measured nothing. gate-dataset was run against a pristine copy and against the same copy with one enum member corrupted, returning 0 and 1; gate-scope was run with and without an out-of-class file present, returning 0 and 1; gate-evidence and gate-source-approval were each run against the gated set and against an absent bundle, returning 0 and 2. The bundle-pairing checker initially reported every bundle coherent over a zero denominator, which is not a pass, and was only shown to discriminate after a deliberately malformed bundle carrying an uncited source add was fabricated and correctly refused.
- PublishRanSix mutations applied in dependency order — sources before releases, so references resolve as they are created — with 0 refusals, followed by 75 re-verifications. The diff touches four documents and nothing outside web/src/data/. npm run validate passed at exit 0 over 133 test files and 2988 tests after application; its one failure during this run is recorded as a caveat and was cleared by withholding two date moves rather than by editing the dataset by hand. No bypass was used at any point: no --force, no --admin, no skipped gate, no lowered threshold, and no direct push to main.
- DeployRanMerge and deploy left to GitHub. Auto-merge was enabled on the pull request so that web-ci gates the merge rather than any judgement of this run's, and the Pages deployment was confirmed afterwards against the full 40-character merge commit rather than against a branch name, with a two-armed control whose arms were required to disagree before either reading was believed.
What was found
- Scouts
- 10
- Pages fetched and hashed
- 164
- Claims proposed
- 262
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims 01-ai long-tail 3-of-3 5 ai-singapore long-tail 3-of-3 3 ai2 long-tail 3-of-3 12 ai21-labs long-tail 3-of-3 6 aleph-alpha long-tail 3-of-3 6 alibaba-cloud pilot 2-of-3 8 amazon pilot 2-of-3 4 anthropic pilot 2-of-3 6 apple long-tail 3-of-3 6 baidu long-tail 3-of-3 6 bytedance-seed long-tail 3-of-3 4 cohere long-tail 3-of-3 6 databricks long-tail 3-of-3 5 deepseek long-tail 3-of-3 6 eleutherai long-tail 3-of-3 12 google-deepmind pilot 2-of-3 6 hugging-face long-tail 3-of-3 8 ibm long-tail 3-of-3 8 kyutai long-tail 3-of-3 3 lelapa-ai long-tail 3-of-3 3 lg-ai-research long-tail 3-of-3 6 liquid-ai long-tail 3-of-3 6 maritaca-ai long-tail 3-of-3 3 meta pilot 2-of-3 5 microsoft pilot 2-of-3 9 minimax long-tail 3-of-3 7 mistral-ai long-tail 3-of-3 9 moonshot-ai long-tail 3-of-3 7 naver long-tail 3-of-3 5 nous-research long-tail 3-of-3 6 nvidia long-tail 3-of-3 7 openai pilot 2-of-3 7 openbmb long-tail 3-of-3 6 reka-ai long-tail 3-of-3 6 sakana-ai long-tail 3-of-3 3 sarvam-ai long-tail 3-of-3 4 snowflake long-tail 3-of-3 4 stability-ai long-tail 3-of-3 7 tencent long-tail 3-of-3 4 tii long-tail 3-of-3 10 upstage long-tail 3-of-3 3 xai long-tail 3-of-3 6 xiaomi long-tail 3-of-3 3 zhipu-ai long-tail 3-of-3 6 What those claims proposed to do Kind Count Effect Unchanged 146 re-verification: the source still states what the dataset already holds Add 106 a record proposed for a collection that did not hold it Change 7 a field whose recorded value the sources contradict Conflict 3 sources disagree, so the disagreement itself is the finding What was evaluated
- Reviewers
- 3
- Verdicts cast
- 786
- Accepted by panel
- 175
- Rejected by panel
- 87
Deterministic gates and required checks — 6 of 6 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-scopeevery path changed since the computed merge-base a51e196d, working tree and committed tip alike 0 Pass4 dataset documents changed, 0 paths outside the ADR 0003 class. Controlled: with an out-of-class file present the same invocation exited 1 and named it, so the pass discriminates. gate-datasetall 631 records in web/src/data after the claims were applied 0 PassAll gates passed over 631 records. Controlled: the identical invocation against a copy with one lifecycle status set to a non-member exited 1 and named the field, so both arms exercised the checking logic rather than the argument parser. gate-evidencethe 112 gated claims across 38 surviving bundles 0 PassEvery evidence entry well-formed. This gate verifies form and never correspondence to the cited URL, per ADR 0005, so it is not evidence that any quote is real; the separate verbatim verifier is what establishes that. gate-source-approvalevery source cited by a surviving claim, against the approved origin profiles 0 Pass0 unapproved origins in the surviving set. Reached exit 0 only after a fourth integrity rule withheld claims whose evidence cited a source that a withheld claim would have added; before that rule it exited 1 over 7 creators. verify-quotesnot requiredall 450 distinct evidence entries over 256 fetched bodies, a superset of the 155 in the gated set 0 Pass0 failures, 0 stale hashes. Not a repository gate but a check written for this run, because no deterministic gate can see a fabricated quote and gates.test.mjs pins that as accepted behaviour. Coverage of the gated subset was proved rather than assumed: all 155 gated entries occur among the 450 verified, with a discriminating control. npm run validatethe full web test suite and Astro/TypeScript diagnostics against the applied dataset 0 Pass133 test files, 2988 tests, 0 errors. It failed at exit 1 on the first application over five model-fit freshness violations, which is the finding recorded in the caveats; it was cleared by withholding the two offending date moves, never by editing a fit statement's own verification date. Applied over a recorded dissent
These met their threshold and were applied. The objection stands on the record and was not overruled.
alibaba-cloud-creator-attribution-unchangedCross-source consistencyThe unchanged claim targets organizations/alibaba-cloud field summary, but that record has no summary field; its stored field is description, whose longer conflict-aware text is not the claimed currentValue. The card therefore does not re-verify a value the current record holds.meta-llama-4-scout-release-date-unchangedEditorial and entity boundariesThe quote is "Llama 4 Version Effective Date: April 5, 2025". That names the Llama 4 family, not Llama 4 Scout. Recording a family version-effective date as meta-llama-4-scout's releaseDate conflates family with release.microsoft-fara-1-5-27b-release-date-unchangedProvenanceNeither evidence quote mentions any date. Quote 1 ('Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers.') and quote 2 ('- 262K context. Long enough for multi-screenshot trajectories with full action history.') describe the model's kind and context, but the field being re-verified is releaseDate '2026-05-21' and no quote states 21 May 2026 or any release date. The card's 'Release date | 21 May 2026' line exists on the source page but was not brought into the evidence, so the claim as it stands is unquoted.microsoft-mai-transcribe-2-release-addProvenanceThe proposedValue's intendedUse — 'Real-world transcription workloads spanning clinical note-taking, legal documentation, accessibility, and closed captioning' — enumerates four specific application domains that appear in NONE of the four evidence quotes. The quotes cover only availability on Foundry/MAI Playground/Open Router, the 'most capable transcription model yet' framing, the September 3, 2026 date, and the 'Turn noisy audio into precise, domain-specific transcripts' models-page line. The 'clinical note-taking … legal documentation … accessibility … closed captioning' sentence exists on the announcement page but was not brought into evidence, so the intendedUse field is unquoted 'unknown filled in'.microsoft-mai-transcribe-2-source-addEditorial and entity boundariesproposedValue.title is 'MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world'. That ranking claim would be written into the dataset. A verbatim superlative in a proposedValue field is still a superlative.openai-gpt-6-astra-cite-announcementCross-source consistencyAppending openai-gpt-6-astra-announcement to openai-gpt-6-astra.sourceIds would make the release cite a fetched page while its unchanged summary explicitly says that announcement page returned HTTP 403 to this run and was not read.openai-gpt-6-astra-release-date-unchangedCross-source consistencyAlthough 2026-09-03 matches openai-gpt-6-astra.releaseDate, this card says the launch page reconfirmed it while the current release summary says that page returned HTTP 403 to this run and was not read; the batch does not update that contradictory provenance text.openai-gpt-6-astra-summary-correct-403-clauseCross-source consistencyreleases.json exactly contains the claimed old HTTP 403 clause, and the proposal changes only that clause consistently with the surrounding RSS-based 2026-09-03 date. However, the replacement asserts that openai-gpt-6-astra-announcement returned HTTP 200 and has no printed date while openai-gpt-6-astra.sourceIds remains only [openai-gpt-6-astra-docs, openai-news-rss]; because the announcement source addition was rejected and is missing from sources.json, the corrected summary would make an uncited page claim and is not internally coherent.
Posted 6 edits
6 edits across 4 documents, a net change of 5 records.
Dataset documents this run changed Document Before After What changed families.json86 86 8 record(s) re-verified with no value change organizations.json44 44 4 record(s) re-verified with no value change releases.json121 123 2 record(s) added; 1 field(s) changed; 63 record(s) re-verified with no value change sources.json286 289 3 record(s) added; ids: anthropic-opus-5-docs, microsoft-mai-transcribe-2-announcement, microsoft-mai-image-2-6-announcement Records added
- MAI-Transcribe-2passportreleases
microsoft-mai-transcribe-2add from microsoft::microsoft-mai-transcribe-2-release-add — MAI-Transcribe-2 is a new Microsoft AI speech recognition release published on 3 September 2026 that accepts audio and produces text. - MAI-Image-2.6passportreleases
microsoft-mai-image-2-6add from microsoft::microsoft-mai-image-2-6-release-add — MAI-Image-2.6 is a new Microsoft AI image release published on 4 September 2026 that accepts text or photo prompts and produces images, currently available in Public Preview through Microsoft Foundry. - Claude Opus 5passportreleases
anthropic-claude-opus-5change to sourceIds from anthropic::anthropic-claude-opus-5-cite-docs — Cite Anthropic's Claude Opus 5 model documentation page on the anthropic-claude-opus-5 release.
Not posted 150 items
Rejected by the review panel
01-ai::01-ai-yi-1-5-34b-chat-status-unchangedreleases 01-ai-yi-1-5-34b-chat — Yi-1.5-34B-Chat is still listed on the 01-ai Hugging Face org page as one of its 34B text-generation models, with no successor announcement replacing it on the Yi repository's News section. [withheld: rejected-by-panel, unapproved-source, short-quote; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelunapproved-sourceshort-quote01-ai::01-ai-yi-1-5-34b-chat-provenance-unchangedfamilies 01-ai-yi-1-5 — The Yi-1.5 family description is still supported by the Yi-1.5-34B-Chat model card, which restates the 500B-token corpus, 3M fine-tuning samples, and apache-2.0 licence. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panel01-ai::01-ai-yi-34b-chat-status-unchangedreleases 01-ai-yi-34b-chat — Yi-34B-Chat is still listed on 01-ai's Hugging Face org page and the Yi repository's News section still records its 2023-11-23 open-source release without a superseding statement. [withheld: rejected-by-panel, unapproved-source, short-quote; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelunapproved-sourceshort-quoteai2::ai2-add-family-tulu-3families ai2-tulu-3 — Add families record "ai2-tulu-3". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelai2::ai2-add-release-olmo-2-13breleases ai2-olmo-2-13b [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quoteai2::ai2-add-release-olmo-2-32breleases ai2-olmo-2-32b [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quoteai2::ai2-add-release-tulu-3-8breleases ai2-tulu-3-8b — Add releases record "ai2-tulu-3-8b": canonical name "Llama-3.1-Tulu-3-8B", organization "ai2", family "ai2-tulu-3". [withheld: rejected-by-panel; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelai2::ai2-add-release-molmo-72breleases ai2-molmo-72b [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quoteai2::ai2-unchanged-molmo-family-firstreleasedatefamilies ai2-molmo [withheld: rejected-by-panel, no-statement]Blocked byrejected-by-panelno-statementai2::ai2-unchanged-olmo-2-family-statusfamilies ai2-olmo-2 — Re-verified: families record "ai2-olmo-2" field "status" still holds "current". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelai21-labs::ai21-jamba-reasoning-3b-card-source-addsources ai21-jamba-reasoning-3b-card — The ai21labs/AI21-Jamba-Reasoning-3B Hugging Face model card is a primary AI21-published source and belongs in sources so a new Jamba Reasoning 3B release record can cite it. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelai21-labs::ai21-jamba-reasoning-3b-release-addreleases ai21-labs-jamba-reasoning-3b — AI21 Labs released Jamba Reasoning 3B, a 3B-parameter hybrid Transformer-Mamba reasoning model under Apache 2.0 with a 256K context window, published on Hugging Face on 2025-10-05. [withheld: rejected-by-panel, short-quote; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelshort-quotealeph-alpha::aleph-tfree-hat-7b-base-card-source-addsources aleph-tfree-hat-7b-base-card — The Aleph-Alpha/tfree-hat-pretrained-7b-base Hugging Face model card is a primary Aleph-Alpha-published source and is needed to back the new TFree-HAT-Pretrained-7B-Base release record and family record. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelaleph-alpha::aleph-alpha-tfree-hat-family-addfamilies aleph-alpha-tfree-hat — Aleph Alpha's tokenizer-free Hierarchical Autoregressive Transformer models constitute a distinct family from Pharia-1; the org page describes it as a separate family based on the HAT architecture paper, with the pretrained 7B base checkpoint published in July 2025. [withheld: rejected-by-panel, unapproved-source; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelunapproved-sourcealeph-alpha::aleph-alpha-tfree-hat-7b-base-release-addreleases aleph-alpha-tfree-hat-pretrained-7b-base — Aleph Alpha published TFree-HAT-Pretrained-7B-Base, a ~7B-parameter tokenizer-free HAT foundation model, on Hugging Face on 2025-07-31 under the Open Aleph License (weights downloadable, licence not OSI-approved), pre-trained in English and German with a long-context adapted checkpoint of 32,900 words. [withheld: rejected-by-panel, short-quote; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelshort-quotealibaba-cloud::alibaba-qwen3-6-27b-release-addreleases alibaba-qwen3-6-27b — Qwen3.6-27B is a new open-weight Alibaba Cloud dense release published on 22 April 2026 with 27B parameters, a 262,144-token native context window, and vision-language input. [withheld: rejected-by-panel; 1 of 3 accepts against a 2 threshold]Blocked byrejected-by-panelalibaba-cloud::alibaba-qwen3-5-27b-release-addreleases alibaba-qwen3-5-27b — Qwen3.5-27B is a new open-weight Alibaba Cloud dense release published on 24 February 2026 with 27B parameters, a 262,144-token native context window, and vision-language input. [withheld: rejected-by-panel; 0 of 3 accepts against a 2 threshold]Blocked byrejected-by-panelalibaba-cloud::alibaba-qwen3-6-27b-source-addsources qwen3-6-27b-model-card — Add the Qwen3.6-27B Hugging Face model card as a model-card source; it is cited by the Qwen3.6-27B release add. [withheld: rejected-by-panel; 1 of 3 accepts against a 2 threshold]Blocked byrejected-by-panelalibaba-cloud::alibaba-qwen3-5-27b-source-addsources qwen3-5-27b-model-card — Add the Qwen3.5-27B Hugging Face model card as a model-card source; it is cited by the Qwen3.5-27B release add. [withheld: rejected-by-panel; 1 of 3 accepts against a 2 threshold]Blocked byrejected-by-panelamazon::amazon-organization-provenance-unchangedorganizations amazon — The Amazon Nova portfolio positioning (foundation models plus Nova Forge / Nova Act, built on internal Amazon AI, served on Bedrock) is re-confirmed on the current AWS Nova overview page. [withheld: rejected-by-panel; 1 of 3 accepts against a 2 threshold]Blocked byrejected-by-panelapple::apple-fastvlm-0-5b-release-addreleases apple-fastvlm-0-5b — Apple published a FastVLM 0.5B open-weight vision language model on Hugging Face under the apple-amlr research licence, alongside the FastVLM 7B already recorded. [withheld: rejected-by-panel, short-quote; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelshort-quotebaidu::baidu-ernie-4-5-family-status-unchangedfamilies baidu-ernie-4-5 — The ERNIE 4.5 family status remains 'current': Baidu continues to publish new ERNIE 4.5 variants on Hugging Face, with ERNIE-4.5-21B-A3B-Thinking created on 2025-09-08. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelbaidu::hugging-face-ernie-4-5-21b-thinking-hub-record-source-addsources hugging-face-ernie-4-5-21b-thinking-hub-record — Add the Hugging Face hub JSON record for baidu/ERNIE-4.5-21B-A3B-Thinking as a source; its createdAt timestamp anchors the release date under dateBasis platform-repository-created. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelbaidu::baidu-ernie-4-5-21b-thinking-release-addreleases baidu-ernie-4-5-21b-thinking — Add release ERNIE-4.5-21B-A3B-Thinking, Baidu's 21B/3B open-weight post-trained MoE within the ERNIE 4.5 family with a 131,072-token context length, whose HF model card was created on 2025-09-08. [withheld: rejected-by-panel, short-quote; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelshort-quotecohere::cohere-add-family-transcribefamilies cohere-transcribe — Add the Cohere Transcribe family, the creator's dedicated audio-in, text-out automatic speech recognition line with weights released under Apache 2.0 on Hugging Face and hosted access on Cohere's Audio Transcriptions endpoint. [withheld: rejected-by-panel; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelcohere::cohere-add-release-transcribe-03-2026releases cohere-transcribe-03-2026 — Add cohere-transcribe-03-2026: a 2B parameter conformer-based encoder-decoder speech recognition model whose weights are released on Hugging Face under Apache 2.0 and which is also live on Cohere's hosted Audio Transcriptions endpoint. [withheld: rejected-by-panel; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-paneldeepseek::deepseek-release-add-v4-flash-vision-expreleases deepseek-v4-flash-vision-exp [withheld: rejected-by-panel, no-statement]Blocked byrejected-by-panelno-statementdeepseek::deepseek-v3-2-license-unchangedreleases deepseek-v3-2 — Re-verified: releases record "deepseek-v3-2" field "license" still holds { name: "MIT License", spdxId: "MIT", url: "https://huggingface.co/deepseek-ai/DeepSeek-V3.2/blob/main/LICENSE", weightsDownloadable: true, osiApproved: true }. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-paneleleutherai::add-source-mesh-transformer-jax-repositorysources eleutherai-mesh-transformer-jax-repository — The Mesh Transformer JAX GitHub repository, which the GPT-J-6B model card names as the training codebase, is a citable source for the GPT-J-6B release's licence and downloadable weights. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-paneleleutherai::add-family-gpt-neoxfamilies eleutherai-gpt-neox — Add a GPT-NeoX family record for EleutherAI, whose sole publicly identified branded release is GPT-NeoX-20B. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelgoogle-deepmind::google-gemini-3-8-flash-docs-source-addsources google-gemini-3-8-flash-docs — Add the Google AI for Developers docs page for gemini-3.8-flash as a new Google source. [withheld: rejected-by-panel; 1 of 3 accepts against a 2 threshold]Blocked byrejected-by-panelgoogle-deepmind::google-gemini-3-8-flash-release-addreleases google-gemini-3-8-flash — Add Gemini 3.8 Flash as a new Google DeepMind release in the Gemini 3 family, dated September 2, 2026, with a 1,048,576-token input context and a 65,536-token output limit, served through the gemini-3.8-flash API alias. [withheld: rejected-by-panel; 0 of 3 accepts against a 2 threshold]Blocked byrejected-by-panelgoogle-deepmind::google-gemini-3-1-flash-lite-status-unchangedreleases google-gemini-3-1-flash-lite — Gemini 3.1 Flash-Lite remains an actively documented generally-available model on ai.google.dev, still described in the present tense with no deprecation banner on its docs page. [withheld: rejected-by-panel; 1 of 3 accepts against a 2 threshold]Blocked byrejected-by-panelhugging-face::hf-add-family-smollm2families hugging-face-smollm2 — Add families record "hugging-face-smollm2". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelhugging-face::hf-add-family-smolvlmfamilies hugging-face-smolvlm — Add families record "hugging-face-smolvlm". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelhugging-face::hf-add-release-smollm2-1-7b-instructreleases hugging-face-smollm2-1-7b-instruct [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quotehugging-face::hf-add-release-smolvlm-instructreleases hugging-face-smolvlm-instruct — Add releases record "hugging-face-smolvlm-instruct": canonical name "SmolVLM-Instruct", organization "hugging-face", family "hugging-face-smolvlm". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelibm::ibm-granite-4-2-8b-release-addreleases ibm-granite-4-2-8b — Add Granite-4.2-8B as a distinct 8B dense-decoder Granite 4.2 reasoning release from IBM, dated August 25, 2026 with a native 128K context and Apache 2.0 licence. [withheld: rejected-by-panel; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelibm::ibm-granite-4-2-3b-release-addreleases ibm-granite-4-2-3b — Add Granite-4.2-3B as a distinct 3B dense-decoder Granite 4.2 reasoning release from IBM, dated August 25, 2026 with a native 128K context and Apache 2.0 licence. [withheld: rejected-by-panel; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panellg-ai-research::lg-ai-research-exaone-4-0-1-2b-add-releasereleases lg-ai-research-exaone-4-0-1-2b — LG AI Research released EXAONE 4.0 1.2B on 2025-07-15 as the small-size member of the EXAONE 4.0 series alongside the already-recorded 32B; it is a text-only model with a 65,536-token context length and 1.07B non-embedding parameters, distributed under the EXAONE AI Model License Agreement 1.2 - NC. [withheld: rejected-by-panel; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelliquid-ai::liquid-lfm25-2-6b-card-source-addsources liquid-lfm25-2-6b-card — The LiquidAI/LFM2.5-2.6B Hugging Face model card is a primary Liquid-AI-published source and belongs in sources so a new LFM2.5-2.6B release record and LFM2.5 family record can cite it. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelliquid-ai::liquid-lfm2-5-family-addfamilies liquid-lfm2-5 — Liquid AI introduced the LFM2.5 family as a follow-on to LFM2 with a 128K context window and agentic post-training; five post-trained checkpoints are published in the LFM2.5 collection on the Liquid AI Hugging Face org page. [withheld: rejected-by-panel, unapproved-source; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelunapproved-sourceliquid-ai::liquid-lfm2-5-2-6b-release-addreleases liquid-lfm2-5-2-6b — Liquid AI published LFM2.5-2.6B, a ~2.6B-parameter LFM2.5 post-trained hybrid model with a 128K context window, on Hugging Face on 2026-07-28 under the LFM Open License v1.0 (bespoke, not OSI-approved) with downloadable weights. [withheld: rejected-by-panel, short-quote; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelshort-quotemicrosoft::microsoft-mai-voice-2-release-addreleases microsoft-mai-voice-2 — MAI-Voice-2 is a new Microsoft AI text-to-speech release published on 2 June 2026 that accepts text and reference audio and produces audio. [withheld: rejected-by-panel, short-quote; 1 of 3 accepts against a 2 threshold]Blocked byrejected-by-panelshort-quoteminimax::hugging-face-minimax-m2-hub-record-source-addsources hugging-face-minimax-m2-hub-record — Add the Hugging Face hub JSON record for MiniMaxAI/MiniMax-M2 as a source; its createdAt timestamp anchors the release date and its cardData carries the license_name. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelminimax::minimax-m2-release-addreleases minimax-m2 — Add release MiniMax-M2, MiniMax's 230B-total / 10B-active open-weight MoE for coding and agentic workflows, whose HF model card was created on 2025-10-22. [withheld: rejected-by-panel; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelmistral-ai::mistral-add-family-medium-3-5families mistral-medium-3-5 — Add the Mistral Medium 3.5 family, the creator's dense flagship line that unifies instruction-following, reasoning and coding into a single model, released under a Modified MIT License. [withheld: rejected-by-panel; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelmistral-ai::mistral-add-release-medium-3-5-128breleases mistral-medium-3-5-128b — Add Mistral Medium 3.5 128B: a dense 128B multimodal model with a 256k context window, released 22 May 2026 in public preview under a Modified MIT License on Hugging Face and via the creator's hosted API. [withheld: rejected-by-panel; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelmistral-ai::mistral-change-devstral-2-123b-status-deprecatedreleases mistral-devstral-2-123b-instruct — Devstral 2 (API name devstral-2512) is listed by Mistral's own model docs under "Deprecated & retired models" with deprecation 5/22/2026 and retirement 7/31/2026, both of which are in the past as of 2026-09-06. [withheld: rejected-by-panel, unapproved-source; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelunapproved-sourcemistral-ai::mistral-change-devstral-small-2-24b-status-deprecatedreleases mistral-devstral-small-2-24b-instruct — Devstral Small 2 (API name labs-devstral-small-2512) is listed by Mistral's own model docs under "Deprecated & retired models" with deprecation 2/27/2026 and retirement 3/31/2026, both of which are in the past as of 2026-09-06. [withheld: rejected-by-panel, unapproved-source; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelunapproved-sourcemoonshot-ai::moonshot-family-add-kimi-k3families moonshot-ai-kimi-k3 [withheld: rejected-by-panel, no-statement]Blocked byrejected-by-panelno-statementmoonshot-ai::moonshot-release-add-kimi-k3releases moonshot-ai-kimi-k3 — Add releases record "moonshot-ai-kimi-k3": canonical name "Kimi-K3", organization "moonshot-ai", family "moonshot-ai-kimi-k3". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelmoonshot-ai::moonshot-kimi-k2-instruct-context-unchangedreleases moonshot-ai-kimi-k2-instruct [withheld: rejected-by-panel, no-statement]Blocked byrejected-by-panelno-statementmoonshot-ai::moonshot-kimi-k2-instruct-license-unchangedreleases moonshot-ai-kimi-k2-instruct — Re-verified: releases record "moonshot-ai-kimi-k2-instruct" field "license" still holds { name: "Modified MIT License", url: "https://huggingface.co/moonshotai/Kimi-K2-Instruct/blob/main/LICENSE", weightsDownloadable: true, osiApproved: false }. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelnaver::naver-hyperclova-x-seed-text-instruct-0-5b-add-releasereleases naver-hyperclova-x-seed-text-instruct-0-5b — NAVER Cloud released HyperCLOVAX-SEED-Text-Instruct-0.5B as the smallest of three lightweight HyperCLOVA X SEED text models on 2025-04-24; it is a Korean-focused text-to-text model with 0.57B total parameters, distributed under the HyperCLOVA X SEED Model License Agreement, and the same announcement already recorded for the 1.5B sibling covers it. [withheld: rejected-by-panel, unapproved-source; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelunapproved-sourcenaver::naver-hyperclova-x-seed-0-5b-add-licensesources naver-hyperclova-x-seed-0-5b-license — Add the copy of the HyperCLOVA X SEED Model License Agreement that ships in the 0.5B repository as a primary source; the new 0.5B release cites it as the licence-terms artefact and it names NAVER Corp. as the intellectual-property holder. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelnous-research::nous-add-release-hermes-4-70breleases nous-hermes-4-70b [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quotenous-research::nous-add-release-hermes-4-405breleases nous-hermes-4-405b [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quotenous-research::nous-unchanged-hermes-4-14b-accesstypereleases nous-hermes-4-14b [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quotenous-research::nous-unchanged-hermes-4-family-statusfamilies nous-hermes-4 — Re-verified: families record "nous-hermes-4" field "status" still holds "current". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelnvidia::nvidia-cosmos-1-0-diffusion-14b-text2world-release-addreleases nvidia-cosmos-1-0-diffusion-14b-text2world — Add Cosmos-1.0-Diffusion-14B-Text2World as a distinct 14B open-weight video-generating release in the Cosmos 1.0 family, released on 6 January 2025 under the NVIDIA Open Model License. [withheld: rejected-by-panel; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelopenai::openai-add-gpt-6-astra-announcement-sourcesources openai-gpt-6-astra-announcement — Add source openai-gpt-6-astra-announcement pointing at https://openai.com/index/gpt-6-astra, OpenAI's launch announcement for GPT-6 Astra. [withheld: rejected-by-panel; 0 of 3 accepts against a 2 threshold]Blocked byrejected-by-panelreka-ai::reka-edge-2603-card-source-addsources reka-edge-2603-card — The RekaAI/reka-edge-2603 Hugging Face model card is a primary Reka-published source and belongs in sources so a new Reka Edge 2603 release record and a Reka Edge family record can cite it. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelreka-ai::reka-edge-family-addfamilies reka-edge — Reka AI introduced Reka Edge, an efficient 7B multimodal vision-language model line published on Hugging Face in March 2026 and listed alongside Reka Flash on the RekaAI org page. [withheld: rejected-by-panel, unapproved-source, short-quote; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelunapproved-sourceshort-quotereka-ai::reka-edge-2603-release-addreleases reka-edge-2603 — Reka AI published Reka Edge 2603, a ~7B multimodal vision-language model, on Hugging Face on 2026-03-11 under a bespoke "reka-edge-2603-license" (open weights, commercial use permitted for organizations under $1M USD annual revenue; not OSI-approved). [withheld: rejected-by-panel; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelsakana-ai::sakana-evollm-jp-family-first-release-date-still-2024-03families sakana-ai-evollm-jp — The EvoLLM-JP family's first release date remains March 2024, as re-stated in Sakana AI's own retrospective blog. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelsarvam-ai::sarvam-m-release-status-still-currentreleases sarvam-ai-sarvam-m-v1 — Sarvam-M remains Sarvam AI's currently-promoted flagship model on its own model card: the card describes it in present-tense terms as a multilingual hybrid-reasoning language model built on Mistral-Small, with no supersession or legacy notice. [withheld: rejected-by-panel; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelsnowflake::snowflake-arctic-instruct-release-date-unchangedreleases snowflake-arctic-instruct — The Snowflake Arctic Instruct model card continues to state Model Release Date April 24, 2024, matching the launch date recorded in the Snowflake-Labs README changelog. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelsnowflake::snowflake-arctic-instruct-license-unchangedreleases snowflake-arctic-instruct — The Snowflake Arctic Instruct card continues to publish under Apache-2.0, and the Snowflake-Labs GitHub repo carries the same Apache-2.0 license header; OSI publishes the Apache License, Version 2.0 as an OSI-approved licence. [withheld: rejected-by-panel, short-quote; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelshort-quotesnowflake::snowflake-arctic-instruct-modality-unchangedreleases snowflake-arctic-instruct — The Snowflake Arctic Instruct card continues to state Arctic is a text-only model — input text only and output text and code only — so the release stays classified as text output. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelstability-ai::stability-sd-3-5-medium-release-addreleases ? — Stability AI has released Stable Diffusion 3.5 Medium (2.5B parameters, MMDiT-X architecture) as part of the SD 3.5 open release; the launch announcement records its release on October 29, 2024 as a companion to the Large variant already in the dataset. [withheld: rejected-by-panel, no-targetId; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelno-targetIdstability-ai::stability-sd-3-5-large-turbo-release-addreleases ? — Stability AI has released Stable Diffusion 3.5 Large Turbo, a distilled variant of SD 3.5 Large that generates images in 4 steps, alongside the Large model on October 22, 2024, under the same Community License. [withheld: rejected-by-panel, no-targetId; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelno-targetIdtencent::tencent-hunyuanimage-3-0-family-status-changefamilies tencent-hunyuanimage-3-0 — Move HunyuanImage-3.0 family status from 'unknown' to 'current'. The creator's own News list on the HunyuanImage-3.0 README now records two further checkpoint releases in the same family on 2026-01-26 (HunyuanImage-3.0-Instruct-Distil and HunyuanImage-3.0-Instruct) after the initial 2025-09-28 open-source release, and a 2025-10-30 vLLM Acceleration release, which together assert the family is in active development. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-paneltii::tii-add-family-falcon-3families tii-falcon-3 — Add families record "tii-falcon-3". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-paneltii::tii-add-family-falcon-mambafamilies tii-falcon-mamba — Add families record "tii-falcon-mamba". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-paneltii::tii-add-release-falcon-3-10b-basereleases tii-falcon-3-10b-base [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quotetii::tii-add-release-falcon-mamba-7breleases tii-falcon-mamba-7b [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quotetii::tii-add-release-falcon-h1-7b-instructreleases tii-falcon-h1-7b-instruct [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quotetii::tii-unchanged-falcon-180b-accesstypereleases tii-falcon-180b — Re-verified: releases record "tii-falcon-180b" field "accessType" still holds "open-weight". [withheld: rejected-by-panel; 1 of 3 accepts against a 3 threshold]Blocked byrejected-by-paneltii::tii-unchanged-falcon-h1-family-statusfamilies tii-falcon-h1 — Re-verified: families record "tii-falcon-h1" field "status" still holds "current". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelupstage::upstage-solar-pro-preview-instruct-context-window-unchangedreleases upstage-solar-pro-preview-instruct — Solar Pro Preview Instruct still states a maximum context length of 4K on its Hugging Face model card; under the run brief's Upstage rule 4K is recorded as 4096, matching the current dataset value. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelxai::xai-grok-4-5-context-unchangedreleases xai-grok-4-5 [withheld: rejected-by-panel, no-statement]Blocked byrejected-by-panelno-statementxai::xai-grok-4-6-status-unchangedreleases xai-grok-4-6 — Re-verified: releases record "xai-grok-4-6" field "status" still holds "current". [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelxai::xai-family-add-grok-buildfamilies xai-grok-build [withheld: rejected-by-panel, no-statement]Blocked byrejected-by-panelno-statementxai::xai-release-add-grok-build-0-1releases xai-grok-build-0-1 [withheld: rejected-by-panel, no-statement, short-quote]Blocked byrejected-by-panelno-statementshort-quotezhipu-ai::hugging-face-glm-4-5-hub-record-source-addsources hugging-face-glm-4-5-hub-record — Add the Hugging Face hub JSON record for zai-org/GLM-4.5 as a source; the record's createdAt timestamp anchors the release date under dateBasis platform-repository-created. [withheld: rejected-by-panel; 2 of 3 accepts against a 3 threshold]Blocked byrejected-by-panelzhipu-ai::zhipu-ai-glm-4-5-release-addreleases zhipu-ai-glm-4-5 — Add release GLM-4.5, the full-size 355B/32B open-weight MIT-licensed sibling of GLM-4.5-Air, whose HF model card was created on 2025-07-20. [withheld: rejected-by-panel; 0 of 3 accepts against a 3 threshold]Blocked byrejected-by-panel
Accepted by the panel, then dropped
ai2::ai2-src-olmo-2-13b-cardsources ai2-olmo-2-13b-model-card — Add sources record "ai2-olmo-2-13b-model-card": "allenai/OLMo-2-1124-13B model card" ("model-card") published by "ai2" at https://huggingface.co/allenai/OLMo-2-1124-13B. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceai2::ai2-src-olmo-2-32b-cardsources ai2-olmo-2-32b-model-card — Add sources record "ai2-olmo-2-32b-model-card": "allenai/OLMo-2-0325-32B model card" ("model-card") published by "ai2" at https://huggingface.co/allenai/OLMo-2-0325-32B. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceai2::ai2-src-tulu-3-8b-cardsources ai2-tulu-3-8b-model-card — Add sources record "ai2-tulu-3-8b-model-card": "allenai/Llama-3.1-Tulu-3-8B model card" ("model-card") published by "ai2" at https://huggingface.co/allenai/Llama-3.1-Tulu-3-8B. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceai2::ai2-src-molmo-72b-cardsources ai2-molmo-72b-model-card — Add sources record "ai2-molmo-72b-model-card": "allenai/Molmo-72B-0924 model card" ("model-card") published by "ai2" at https://huggingface.co/allenai/Molmo-72B-0924. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceapple::apple-fastvlm-0-5b-model-card-source-addsources apple-fastvlm-0-5b-card — The apple/FastVLM-0.5B Hugging Face page is a primary Apple-published model card and belongs in sources so a new FastVLM-0.5B release record can cite it. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceapple::apple-fastvlm-7b-siblingids-add-0-5breleases apple-fastvlm-7b — The existing FastVLM 7B record should list FastVLM 0.5B as a sibling, since Apple's own Hugging Face card and evaluations table present both as members of the same family launch. [withheld: dangling-reference; 3 of 3 accepts against a 3 threshold]Blocked bydangling-referencebaidu::baidu-ernie-4-5-21b-thinking-source-addsources baidu-ernie-4-5-21b-thinking-model-card — Add the ERNIE-4.5-21B-A3B-Thinking model card on Hugging Face as a Baidu source; the release is not yet in the dataset. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcecohere::cohere-add-source-transcribe-model-cardsources cohere-transcribe-model-card — Add the Cohere Transcribe (cohere-transcribe-03-2026) Hugging Face model card as a model-card source. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcedeepseek::deepseek-source-add-vision-exp-cardsources deepseek-v4-flash-vision-exp-model-card — Add sources record "deepseek-v4-flash-vision-exp-model-card": "DeepSeek-V4-Flash-Vision-Exp model card" ("model-card") published by "deepseek" at https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcedeepseek::deepseek-source-add-vision-exp-hub-recordsources hugging-face-deepseek-v4-flash-vision-exp-hub-record — Add sources record "hugging-face-deepseek-v4-flash-vision-exp-hub-record": "DeepSeek-V4-Flash-Vision-Exp Hugging Face hub record" ("platform-record") published by "hugging-face" at https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceeleutherai::add-family-gpt-jfamilies eleutherai-gpt-j — Add a GPT-J family record for EleutherAI, whose sole publicly identified member is GPT-J-6B. [withheld: dangling-citation; 3 of 3 accepts against a 3 threshold]Blocked bydangling-citationeleutherai::add-release-gpt-j-6breleases eleutherai-gpt-j-6b — Add EleutherAI's GPT-J-6B as a 6.053-billion-parameter, Apache-2.0-licensed, open-weight text model whose Hugging Face repository was created on 2022-03-02. [withheld: dangling-citation; 3 of 3 accepts against a 3 threshold]Blocked bydangling-citationeleutherai::add-release-gpt-neox-20breleases eleutherai-gpt-neox-20b — Add EleutherAI's GPT-NeoX-20B as a 20.555-billion-parameter, Apache-2.0-licensed, open-weight text model with a 2,048-token sequence length whose Hugging Face repository was created on 2022-04-07. [withheld: dangling-reference; 3 of 3 accepts against a 3 threshold]Blocked bydangling-referencehugging-face::hf-src-smollm2-1-7b-instruct-cardsources hugging-face-smollm2-1-7b-instruct-model-card — Add sources record "hugging-face-smollm2-1-7b-instruct-model-card": "HuggingFaceTB/SmolLM2-1.7B-Instruct model card" ("model-card") published by "hugging-face" at https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B-Instruct. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcehugging-face::hf-src-smolvlm-instruct-cardsources hugging-face-smolvlm-instruct-model-card — Add sources record "hugging-face-smolvlm-instruct-model-card": "HuggingFaceTB/SmolVLM-Instruct model card" ("model-card") published by "hugging-face" at https://huggingface.co/HuggingFaceTB/SmolVLM-Instruct. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceibm::ibm-granite-4-2-8b-source-addsources ibm-granite-4-2-8b-model-card — Add the Hugging Face model card for granite-4.2-8b as a new IBM source. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceibm::ibm-granite-4-2-3b-source-addsources ibm-granite-4-2-3b-model-card — Add the Hugging Face model card for granite-4.2-3b as a new IBM source. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcemicrosoft::microsoft-mai-voice-2-source-addsources microsoft-mai-voice-2-announcement — Add the MAI-Voice-2 announcement page on microsoft.ai as an official-announcement source; it is cited by the MAI-Voice-2 release add. [withheld: orphan-source; 3 of 3 accepts against a 2 threshold]Blocked byorphan-sourceminimax::minimax-m2-family-addfamilies minimax-m2 — Add family MiniMax-M2, a distinct open-weight MoE model line from MiniMax-M1, whose HF model card was created on 2025-10-22. [withheld: dangling-citation; 3 of 3 accepts against a 3 threshold]Blocked bydangling-citationminimax::minimax-m2-license-conflictreleases minimax-m2 — MiniMax-M2 model card body labels the licence 'MIT' while the Hugging Face hub record's cardData labels it 'modified-mit' (license_name) with license_link pointing at github.com/MiniMax-AI/MiniMax-M2/blob/main/LICENSE. A human should retrieve the LICENSE file itself to determine whether the licence is vanilla MIT (OSI-approved) or a modification that must not be recorded as OSI-approved MIT. [withheld: evidence-source-unapproved; 3 of 3 accepts against a 3 threshold]Blocked byevidence-source-unapprovedmistral-ai::mistral-add-source-medium-3-5-announcementsources mistral-medium-3-5-announcement — Add the Mistral Medium 3.5 announcement page as an official-announcement source on the creator's own domain. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcemistral-ai::mistral-add-source-medium-3-5-model-cardsources mistral-medium-3-5-model-card — Add the Mistral Medium 3.5 128B Hugging Face model card as a model-card source. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcemoonshot-ai::moonshot-source-add-kimi-k3-cardsources moonshot-ai-kimi-k3-model-card — Add sources record "moonshot-ai-kimi-k3-model-card": "Kimi-K3 model card" ("model-card") published by "moonshot-ai" at https://huggingface.co/moonshotai/Kimi-K3. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcemoonshot-ai::moonshot-source-add-kimi-k3-hub-recordsources hugging-face-kimi-k3-hub-record — Add sources record "hugging-face-kimi-k3-hub-record": "Kimi-K3 Hugging Face hub record" ("platform-record") published by "hugging-face" at https://huggingface.co/api/models/moonshotai/Kimi-K3. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcenaver::naver-hyperclova-x-seed-0-5b-add-model-cardsources naver-hyperclova-x-seed-0-5b-model-card — Add the Hugging Face model card for HyperCLOVAX-SEED-Text-Instruct-0.5B as a primary source; it is required to evidence the new 0.5B release's parameter count, modality and licence claims. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcenaver::hugging-face-hyperclova-x-seed-0-5b-add-hub-recordsources hugging-face-hyperclova-x-seed-0-5b-hub-record — Add the Hugging Face registry JSON record for HyperCLOVAX-SEED-Text-Instruct-0.5B as a primary source; it fixes the repository createdAt (2025-04-22) that anchors the platform-observable release timing and reports total safetensors parameters. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcenous-research::nous-src-hermes-4-70b-cardsources nous-hermes-4-70b-model-card — Add sources record "nous-hermes-4-70b-model-card": "NousResearch/Hermes-4-70B model card" ("model-card") published by "nous-research" at https://huggingface.co/NousResearch/Hermes-4-70B. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcenous-research::nous-src-hermes-4-405b-cardsources nous-hermes-4-405b-model-card — Add sources record "nous-hermes-4-405b-model-card": "NousResearch/Hermes-4-405B model card" ("model-card") published by "nous-research" at https://huggingface.co/NousResearch/Hermes-4-405B. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcenvidia::nvidia-cosmos-diffusion-14b-text2world-addsources nvidia-cosmos-diffusion-14b-text2world-model-card — Add the Hugging Face model card for Cosmos-1.0-Diffusion-14B-Text2World as a new NVIDIA source. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceopenai::openai-gpt-6-astra-cite-announcementreleases openai-gpt-6-astra — Cite the OpenAI launch announcement for GPT-6 Astra on the openai-gpt-6-astra release, now that the announcement page is retrievable. [withheld: dangling-citation; 2 of 3 accepts against a 2 threshold]Blocked bydangling-citationopenai::openai-gpt-6-astra-release-date-unchangedreleases openai-gpt-6-astra — The OpenAI GPT-6 Astra release date remains 2026-09-03, as reconfirmed by the OpenAI news RSS pubDate and the launch announcement page. [withheld: evidence-source-unapproved; 2 of 3 accepts against a 2 threshold]Blocked byevidence-source-unapprovedopenai::openai-gpt-6-astra-summary-correct-403-clausereleases openai-gpt-6-astra — The openai-gpt-6-astra release summary's clause describing the announcement page as HTTP 403 and unread is out of date: the announcement page was retrievable at HTTP 200 to this run and carries no printed calendar date of its own, so the 2026-09-03 date still comes from the news feed pubDate rather than the page. [withheld: evidence-source-unapproved; 2 of 3 accepts against a 2 threshold]Blocked byevidence-source-unapprovedstability-ai::stability-sd-3-5-medium-source-addsources stability-ai-sd-3-5-medium-model-card — The Stable Diffusion 3.5 Medium model card on Hugging Face is a primary source for the SD 3.5 Medium release and needs its own source id. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcestability-ai::stability-sd-3-5-large-turbo-source-addsources stability-ai-sd-3-5-large-turbo-model-card — The Stable Diffusion 3.5 Large Turbo model card on Hugging Face is a primary source for the SD 3.5 Large Turbo release and needs its own source id. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcetii::tii-src-falcon-3-10b-cardsources tii-falcon-3-10b-model-card — Add sources record "tii-falcon-3-10b-model-card": "tiiuae/Falcon3-10B-Base model card" ("model-card") published by "tii" at https://huggingface.co/tiiuae/Falcon3-10B-Base. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcetii::tii-src-falcon-mamba-7b-cardsources tii-falcon-mamba-7b-model-card — Add sources record "tii-falcon-mamba-7b-model-card": "tiiuae/falcon-mamba-7b model card" ("model-card") published by "tii" at https://huggingface.co/tiiuae/falcon-mamba-7b. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcetii::tii-src-falcon-h1-7b-instruct-cardsources tii-falcon-h1-7b-instruct-model-card — Add sources record "tii-falcon-h1-7b-instruct-model-card": "tiiuae/Falcon-H1-7B-Instruct model card" ("model-card") published by "tii" at https://huggingface.co/tiiuae/Falcon-H1-7B-Instruct. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcexai::xai-source-add-pricingsources xai-pricing — Add sources record "xai-pricing": "Pricing" ("official-docs") published by "xai" at https://docs.x.ai/developers/pricing. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcezhipu-ai::zhipu-ai-glm-4-5-source-addsources zhipu-ai-glm-4-5-base-model-card — Add the GLM-4.5 (base) model card on Hugging Face as a Zhipu AI source; the existing zhipu-ai-glm-4-5-model-card points at the Air variant. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceai2::ai2-unchanged-molmo-7b-d-accesstypereleases ai2-molmo-7b-d — Re-verified: releases record "ai2-molmo-7b-d" field "accessType" still holds "open-weight". [withheld: evidence-source-unapproved; 3 of 3 accepts against a 3 threshold]Blocked byevidence-source-unapprovedeleutherai::add-source-gpt-j-6b-model-cardsources eleutherai-gpt-j-6b-model-card — The GPT-J-6B model card on Hugging Face is a citable source for the GPT-J-6B release and the GPT-J family. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceeleutherai::add-source-gpt-neox-20b-model-cardsources eleutherai-gpt-neox-20b-model-card — The GPT-NeoX-20B model card on Hugging Face is a citable source for the GPT-NeoX-20B release and the GPT-NeoX family. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceeleutherai::add-source-gpt-neox-library-repositorysources eleutherai-gpt-neox-library-repository — The GPT-NeoX library GitHub repository is a citable source for the GPT-NeoX family and the GPT-NeoX-20B release. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceeleutherai::add-source-gpt-neox-20b-papersources eleutherai-gpt-neox-20b-paper — The GPT-NeoX-20B arXiv paper by EleutherAI authors is a citable source for the GPT-NeoX-20B release. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourceminimax::minimax-m2-source-addsources minimax-m2-model-card — Add the MiniMax-M2 model card on Hugging Face as a MiniMax source; the MiniMax-M2 model line is not yet in the dataset. [withheld: orphan-source; 3 of 3 accepts against a 3 threshold]Blocked byorphan-sourcexai::xai-grok-4-6-context-unchangedreleases xai-grok-4-6 — Re-verified: releases record "xai-grok-4-6" field "contextWindow" still holds 500000. [withheld: evidence-source-unapproved; 3 of 3 accepts against a 3 threshold]Blocked byevidence-source-unapprovedzhipu-ai::zhipu-ai-glm-4-5-air-parameters-unchangedreleases zhipu-ai-glm-4-5-air — GLM-4.5-Air still has 106B total parameters and 12B active parameters, as stated on the GLM-4.5 model card. [withheld: evidence-source-unapproved; 3 of 3 accepts against a 3 threshold]Blocked byevidence-source-unapprovedzhipu-ai::zhipu-ai-glm-4-5-air-license-unchangedreleases zhipu-ai-glm-4-5-air — GLM-4.5-Air is still released under the MIT open-source license, as stated on the GLM-4.5 model card, and MIT remains OSI-approved on the OSI's own MIT page. [withheld: evidence-source-unapproved; 3 of 3 accepts against a 3 threshold]Blocked byevidence-source-unapproved
Verification date deliberately held back
anthropic::anthropic-claude-haiku-4-5-context-window-unchangedreleases anthropic-claude-haiku-4-5 — Claude Haiku 4.5's context window remains 200,000 tokens, per the Sonnet 5 documentation page's model comparison table. [withheld: freshness-dependant-stale; 3 of 3 accepts against a 2 threshold]Blocked byfreshness-dependant-stalemeta::meta-llama-4-scout-release-date-unchangedreleases meta-llama-4-scout — Llama 4 Scout continues to state a version effective date of April 5, 2025 on its Hugging Face model card. [withheld: freshness-dependant-stale; 2 of 3 accepts against a 2 threshold]Blocked byfreshness-dependant-stale
Source refused by the approval gate
ai-singapore::sea-lion-v3-8b-parameters-still-8-billionreleases ai-singapore-llama-sea-lion-v3-8b — Llama-SEA-LION-v3-8B still measures at ~8B parameters: the Hugging Face Hub record's safetensors total is 8,030,261,248, which rounds to the recorded 8. [withheld: unapproved-source; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-sourceai-singapore::sea-lion-v3-family-first-release-date-still-2024-12families ai-singapore-sea-lion-v3 — The SEA-LION v3 family's first-release month remains 2024-12: AI Singapore's SEA-LION v3 documentation still states the version was released in Dec 2024, and the Hugging Face Hub record shows the 8B repository was created on 2024-12-11. [withheld: unapproved-source; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-sourceai21-labs::ai21-jamba-v0-1-status-conflict-newer-versionsreleases ai21-labs-jamba-v0-1 — The Jamba v0.1 model card names two later versions (Jamba-1.5-Mini and Jamba-1.5-Large) as replacements, while the earlier announcement page and the file itself remain live and are still linked from AI21's own resources. The two readings do not agree on whether v0.1 should still carry "current" status; both are recorded here as a finding rather than resolved. [withheld: unapproved-source, short-quote; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-sourceshort-quoteai21-labs::ai21-jamba-v0-1-context-window-unchangedreleases ai21-labs-jamba-v0-1 — The Jamba v0.1 model card still states a 256K context length, matching the 262,144 already recorded. [withheld: unapproved-source, short-quote; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-sourceshort-quoteai21-labs::ai21-jamba-v0-1-access-type-unchangedreleases ai21-labs-jamba-v0-1 — Jamba v0.1 remains available as downloadable weights on Hugging Face under Apache 2.0, matching the existing open-weight access type. [withheld: unapproved-source, short-quote; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-sourceshort-quoteai21-labs::ai21-labs-jamba-family-status-unchangedfamilies ai21-labs-jamba — The Jamba family status remains current, corroborated by AI21's October 2025 release of Jamba Reasoning 3B under the same family name. [withheld: unapproved-source; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-sourcehugging-face::hf-unchanged-smollm3-family-statusfamilies hugging-face-smollm3 [withheld: no-statement, unapproved-source]Blocked byno-statementunapproved-sourcelg-ai-research::lg-ai-research-exaone-4-0-1-2b-add-sourcesources lgai-exaone-4-0-1-2b-model-card — Add the Hugging Face model card for EXAONE-4.0-1.2B as a primary source; it is required to evidence the new 1.2B release record's context length, parameter count, licence and modality claims. [withheld: short-quote, unapproved-source; 3 of 3 accepts against a 3 threshold]Blocked byshort-quoteunapproved-sourceliquid-ai::liquid-lfm2-family-status-unchangedfamilies liquid-lfm2 — The LFM2 family status remains current; Liquid AI's HF org page continues to list LFM2 variants alongside the newer LFM2.5 series and the LFM2-1.2B card is still active. [withheld: unapproved-source, short-quote; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-sourceshort-quotereka-ai::reka-flash-family-status-unchangedfamilies reka-flash — The Reka Flash family status remains current; the RekaAI HF org page still lists Reka Flash 3.1 (updated Jul 10, 2025) and its 3.5-bit quantized companion alongside the newer Reka Edge 2603. [withheld: unapproved-source, short-quote; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-sourceshort-quotexiaomi::mimo-7b-family-first-release-date-still-2025-05-30families xiaomi-mimo-7b — The MiMo-7B family's first release date remains 2025-05-30 under the platform-repository-created basis: the Hugging Face Hub record for XiaomiMiMo/MiMo-7B-RL-0530 still shows createdAt "2025-05-30T01:19:37.000Z". [withheld: unapproved-source; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-sourcexiaomi::mimo-7b-rl-0530-access-type-still-open-weightreleases xiaomi-mimo-7b-rl-0530 — MiMo-7B-RL-0530's accessType remains open-weight: Xiaomi's own MiMo repository README states that the series is open-sourced with downloadable checkpoints for the RL model, and the Hugging Face Hub record still marks the repository as ungated with a safetensors weight enumeration. [withheld: unapproved-source; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-sourcexiaomi::mimo-7b-rl-0530-canonical-name-still-matches-hubreleases xiaomi-mimo-7b-rl-0530 — The recorded canonical name "MiMo-7B-RL-0530" still matches the Hugging Face repository id and card title verbatim. [withheld: unapproved-source; 3 of 3 accepts against a 3 threshold]Blocked byunapproved-source
What this run does not prove
- Fabricated evidence is invisible to every deterministic gate, by design. gate-evidence.mjs checks contentHash shape and, in its own words, never correspondence to the cited url; ADR 0005 records that limit and gates.test.mjs pins as expected behaviour that a well-formed but fabricated hash and quote pass. This run caught one such fabrication — a quote attributed to a Microsoft page that did not occur in the fetched body — and it was caught only by a verbatim-quote verifier written for this run and not by any gate. Any run that relies on the gates alone is not checking whether its quotes are real.
- Per-claim judgement cannot on its own produce a coherent dataset. Five distinct integrity failures this run share one shape: the panel accepted claim A and rejected claim B, each correctly, and the surviving pair was incoherent — an added source nothing cites, a citation to a source no longer added, a reference to a record no longer created, evidence citing a source a withheld claim would have added, and a re-verification that strands guidance resting on it. Coherence is a property of the kept set and is unknown until the last claim is classified, so bundles were held unwritten until classification finished and a fixpoint over four rules then withheld 50 claims across 3 passes. The rules only ever withhold and never promote.
- The fifth of those failures was caught by validate.ts and by nothing else. Applying an unchanged claim moves a record's verifiedAt to the run date, and validate.ts line 433 refuses a model fit statement whose own verifiedAt is older than any fact it rests on — where a fit fact resolves to the release's or family's own verifiedAt. Re-verifying a release therefore strands the guidance built on it. Five violations across two releases, meta-llama-4-scout and anthropic-claude-haiku-4-5, were cleared by withholding those two date moves and recording them as verification-held. Bumping the fit statements instead would have asserted that their guidance was re-verified, which no reviewer did.
- Three of this run's own sweep scripts accepted a directory argument and silently ignored it, measuring the run root instead of the set named. Every one failed at exit 0 with plausible output, and the tell each time was the denominator — 44 bundles reported where the gated set holds 38. Two were genuinely measuring the wrong set; the third was harmless only because its superset contains the subject, which was then proved rather than assumed. Separately, the bundle-pairing checker takes exactly one bundle and silently ignores the rest, so passing it 38 files reported on the first alone: a zero denominator is not a pass.
- MINIMUM_QUOTE_LENGTH is satisfiable by padding. The 24-character floor was met on some HuggingFace model cards by widening a quote into adjacent sidebar metadata. Such quotes are verbatim and contiguous, so they are honest, but the added span carries no support for the claim. Scouting preferred prose on the second pass; 2 metadata-spanning quotes remain in the surviving set and are disclosed here rather than smoothed over.
- 21 claims were withheld because they cite a source origin that is not approved — HuggingFace organization pages and two Hub API endpoints — of which 6 were lost to that reason alone. They are recorded rather than published, and whether those origins should be approved is a question for a profile change and not for this run to decide.
- One openai announcement page returned three different sha256 digests across three separate HTTP 200 fetches, so a content hash on an unstable page dates a fetch rather than pinning a document. Stale hashes were re-pointed only where the recorded quote still occurred as an exact substring of the newly fetched body; where it did not, the claim was withheld rather than re-hashed.
Follow-ups — proposed, not fixed
- Fold the verbatim-quote verifier into the repository's own gates. It caught a real fabrication this run and no shipped gate can see that class; leaving it as a per-run script means the next run is only as safe as whoever remembers to write it again.
- The freshness dependency between a record and the guidance resting on it is enforced only by validate.ts, at the very end. Teaching gate-dataset or a gate of its own to refuse a verifiedAt move that strands a dependant would catch it where the other four integrity rules are caught, rather than as a late test failure.
- 35 claims were withheld as orphan-source: a source add that no surviving claim cites. Many are model cards whose paired release add was rejected on other grounds, so pairing them properly is a scouting improvement rather than a policy question, and several would land next run.
- Consider whether HuggingFace organization pages and the two Hub API endpoints should be approved origins. 21 claims turn on that question this run, 6 of them on nothing else.
2026-09-03-660be8
Pilot depth 2026-09-03 — Microsoft Fara 1.5 and Alibaba Qwen3.8-Flash-Next
Scope requested: All creators, as instructed. In practice the seven creators carrying a reviewed profile under tools/updater/profiles were probed and two produced claims, so this is a depth pass over the pilot set and not a sweep of the catalogue; the 37 organisations with no reviewed profile were not scouted and are recorded as not covered.Published11 edits posted · 7 items withheldA depth pass over the seven reviewed pilot profiles that published two of them. Eleven claims were proposed across two bundles, every one of kind add, carrying 35 evidence entries; each quote was machine-checked as a contiguous verbatim substring of the exact stored bytes its contentHash names, with a positive and a negative span control in the same invocation. The three-rubric panel accepted all eleven unanimously at 3-of-3, which is above the 2-of-3 the pilot policy requires, and rejected none. The run stopped once before it published: on its first pass npm run validate reddened on web/asset-budgets.json, whose recorded measuredRaw figures had drifted past the 2% tolerance, and the only fix — re-recording them — lies outside the ADR 0003 qualifying class. Rather than edit that file or shrink the tranche to fit a stale figure, the run stopped and filed the conflict as issue #833. It resumed only after trunk PR #830 re-recorded those figures against its own merge tree for unrelated reasons; current trunk was then merged in and everything re-measured against the merge, which is the state CI builds. No budget figure, ceiling, threshold or tolerance was changed by this run.
- PreflightRanConfirmed the dataset round-trips byte-exactly through JSON.stringify(v, null, 2) + newline before touching it — no CRLF, no BOM, 2-space indent — and captured insertion anchors from the committed order rather than assuming one. Checked that every proposed id was absent from the dataset and that both target families microsoft-fara and qwen3-8 already existed, so no release could orphan. The first id-presence probe was a PowerShell git grep whose pattern was mangled by the shell: it reported every id absent including a known-present positive control, which is the only reason the fault was visible. It was replaced with a Node reader carrying its own positive and negative controls.
- ScoutRanStored 31 page bodies and cut every quote from those exact bytes. Eleven claims across two bundles, 35 evidence entries, all retrieval fetch and no search snippets. Several origins refused or could not be read and are recorded under withheld rather than worked around: openai.com returned 403, ai.meta.com/blog 400, the Gemini API model docs failed outright, and microsoft.ai/models/mai-code-1-1-flash returned 404. qwen.ai is reachable but is not in the approved origin catalogue, so nothing was cited from it; raw.githubusercontent.com is likewise unapproved, which forced the Qwen README to be re-fetched from its github.com blob page instead. derivedFromIds was left empty on both Fara records with the base model named only in prose, matching the committed microsoft-fara-1-5-27b precedent rather than inventing a lineage edge.
- ReviewRanThree reviewers, one per rubric, run independently: none saw the scout reasoning, another reviewer verdict, the threshold, or the running tally. 33 verdicts over 11 claims, all accept, so no claim was carried over a dissent and none was rejected. No reviewer was re-run. One harness casualty is worth recording: the provenance reviewer deleted check-bundle-pairing.mjs from the run directory as scratch, and it had to be rewritten with a live negative control before the bundles could be re-checked — a sub-agent tidying its own working directory can remove the harness that judges it.
- GatesRanRun in order and re-run in full after each trunk merge, because a verdict measured against the branch alone is not a verdict about what CI builds, and a verdict measured against a superseded trunk is not either. Trunk was merged twice: once to pick up the re-recorded budgets, and again when it moved on to carry ADR 0013 and a changed release schema while this entry was being written. The second merge shifted the anchor, which correctly invalidated the record counts written here and made gate-ledger exit 1 until they were re-measured — the gate catching a stale figure is the gate working. gate-evidence and gate-source-approval ran before any dataset file was touched, and gate-source-approval re-derived the anchor itself rather than being handed one. gate-ledger initially exited 1 against this branch for a different reason, naming both the missing entry and the run id declared in the commit subject: it is not part of npm run validate and does not run in CI, so nothing downstream would have caught it, and being fully in class is precisely what set unattended true and made the record mandatory. This entry is that failure being repaired rather than argued with. Two invocations exited 2 during the run — gate-evidence on a wrong flag and verify-quotes with no arguments — and both were read as refusals and re-run, never recorded as results.
- PublishNot runDeliberately not run. The dock boundary ends at a reviewable commit: it does not open a pull request, does not push, does not merge, does not rebase and does not gate its own work. The dataset change and this entry are committed to the branch and handed to the publishing step, which fills in the pull-request reference below once that pull request exists to be named. Landing remains GitHub’s once CI is green.
What was found
- Scouts
- 7
- Pages fetched and hashed
- 31
- Claims proposed
- 11
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims microsoft pilot 2-of-3 6 alibaba-cloud pilot 2-of-3 5 What those claims proposed to do Kind Count Effect Add 11 All eleven were accepted unanimously and all eleven were applied: three releases and eight sources. Nothing was dropped after acceptance, because both target families already held releases and no record could be orphaned. Not covered
- Only the seven creators carrying a reviewed profile under tools/updater/profiles were probed, and only microsoft and alibaba-cloud yielded claims. The other 37 organisations in the dataset were not scouted at all this run, so this is a depth pass and its silence about them is absence of evidence rather than evidence of absence.
- The four MAI models named on microsoft.ai, plus MAI-Image-2.6, were confirmed absent from releases.json but were not scouted into claims. Their absence was measured with a positive and a negative control, so it is a finding rather than an assumption.
- Link health was not swept and the second Python interpreter was not exercised; ci-preflight names both as outside what it covers, along with the web-e2e browser check and GitHub’s own production of the aggregate-checks status.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 33
- Accepted by panel
- 11
- Rejected by panel
- 0
Deterministic gates and required checks — 9 of 9 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidencethe microsoft bundle, before any dataset file was touched 0 Pass6 claims admissible under the pilot policy, all 6 applying to the dataset. The policy was derived from the reviewed-profile set rather than read from the bundle, and the derived pilot matched the declared pilot, so 2-of-3 is the threshold actually applied — which the panel cleared at 3-of-3 anyway. gate-evidencethe alibaba-cloud bundle, before any dataset file was touched 0 Pass5 claims admissible under the pilot policy, all 5 applying to the dataset. gate-source-approvalthe microsoft bundle against the approved origin set at the branch’s merge-base with refs/remotes/origin/main 0 Pass20 citations rest on approved sources: 1 inherited from the dataset at the merge base and 4 proposed on already-trusted origins. No origin was widened, and qwen.ai was left unapproved rather than added to make a citation fit. gate-source-approvalthe alibaba-cloud bundle against the approved origin set at the branch’s merge-base with refs/remotes/origin/main 0 Pass15 citations rest on approved sources: 1 inherited from the dataset at the merge base and 4 proposed on already-trusted origins. No origin was widened. gate-datasetweb/src/data after the eleven claims and this entry were applied 0 PassAll gates passed over the whole dataset. Both new Fara releases join a family that already holds a release and the Qwen release joins qwen3-8, so the empty-family refusal does not bite, and each release status agrees with its family. gate-scopethe branch against its computed merge-base with refs/remotes/origin/main 0 PassIn class. releases.json and sources.json changed, plus this ledger entry, which ADR 0006 admits to the class. web/asset-budgets.json is deliberately absent from the diff: trunk re-recorded it in PR #830 and that version became the merge base, so its figures were inherited rather than edited. No test, workflow, skill or document was touched. gate-ledgerthe branch against its computed merge-base, after this entry was added 0 PassExit 1 before this entry existed, naming two findings: a change confined to the qualifying class may auto-merge unattended and so must record itself (ADR 0006, the #419 failure), and the commit subject declared run 2026-09-03-660be8 with no entry reaching the ledger. Both are repaired by this entry, whose posted.documents figures the gate re-counts from the anchor and the working tree rather than trusting what is written here. The gate is not in npm run validate and does not run in CI, so its absence from the first verification set was invisible from inside the run. npm run validateweb/, the test suite plus Astro and TypeScript diagnostics 0 PassGreen on the merge commit, including tests/build/asset-budgets.test.ts, which is what had stopped the first pass. A green baseline was measured on the unmodified tree first, so any failure would have been attributable to this change. ci-preflightthe repository root, selecting the pull-request checks this diff triggers 0 PassSelected web-ci, skills-ci and source-link-health-tests, all passing. Run from the root because npm run validate reads only web/ and would not exercise the checks a diff outside it triggers; this diff is dataset-only, so the selection is narrow. Posted 11 edits
11 edits across 2 documents, a net change of 11 records.
Dataset documents this run changed Document Before After What changed releases.json114 117 Three releases appended after their family siblings: the two Fara 1.5 records after microsoft-fara-1-5-27b, and the Qwen record after alibaba-qwen3-8-27b. sources.json269 277 Eight sources, each cited by a record claim in the same bundle: microsoft-fara-1-5-4b-model-card, hugging-face-fara-1-5-4b-hub-record, microsoft-fara-1-5-9b-model-card, hugging-face-fara-1-5-9b-hub-record, qwen3-8-flash-next-model-card, hugging-face-qwen3-8-flash-next-hub-record, qwen3-8-flash-next-repository, qwen3-8-flash-next-license. They are counted here rather than listed under records, because a source resolves to no page on this site and would render as a link to nowhere. Each document links to the file as this run left it, not as it stands today.
Records added
- Fara1.5-4Bpassportreleases
microsoft-fara-1-5-4bMIT with osiApproved true against the OSI licence index, 262144 context, creator-stated releaseDate 2026-05-21, status unknown. derivedFromIds left empty with the base named only in prose, matching the committed 27B record. - Fara1.5-9Bpassportreleases
microsoft-fara-1-5-9bSame licence, context, date basis and status as the 4B, sourced separately from its own model card and Hub API record rather than inherited from its sibling. - Qwen3.8-Flash-Nextpassportreleases
alibaba-qwen3-8-flash-nextCreator-stated 2026-08-26 from the GitHub README, 262144 native context, Qwen Community License 1.0 with osiApproved false citing the OSI licence index, status preview. Only parameters.activeBillions 6 is recorded, because the card states no total.
Not posted 7 items
Source refused by the approval gate
qwen-ai-origin-citationsqwen.ai is reachable and carries creator-stated material, but it is not in the approved origin catalogue, so nothing was cited from it. The alternative — widening the catalogue mid-run to admit a source the run wanted — is the move the approval gate exists to prevent, so the origin was left unapproved and the facts were taken from approved origins instead.Blocked bygate-source-approvalraw-githubusercontent-fetchesraw.githubusercontent.com is not approved while github.com is, so the Qwen README could not be cited from its raw URL and was re-fetched from the github.com blob page. The quote is verbatim against the blob body actually stored, not against the raw file.Blocked bygate-source-approvalunreachable-creator-originsopenai.com/news and /index returned 403, ai.meta.com/blog returned 400, and ai.google.dev/gemini-api/docs/models failed to fetch. No claim was made about any of the three creators on the strength of recalled knowledge or a search snippet, so those creators contributed nothing this run.
Blocked by policy before it could run
asset-budget-re-recordOn the first pass the run needed web/asset-budgets.json re-recorded to keep web-ci green, because trunk drift plus this tranche pushed /tree and /compare past the 2% measuredDrift tolerance. That file is outside the ADR 0003 qualifying class, so the run stopped and filed issue #833 rather than edit it, and it declined to shrink the tranche to fit a figure known to be stale — sizing researched data by stale documentation is the failure #813 exists to end. The need disappeared when trunk PR #830 re-recorded those figures for its own reasons; the structural conflict did not, and #833 remains open.Blocked byweb/asset-budgets.jsonADR 0003
Out of the run’s reach
microsoft-mai-family-releasesFour MAI models named on microsoft.ai, plus MAI-Image-2.6, hold no release record in releases.json. Their absence was measured rather than assumed, with a known-present positive control and a fabricated negative control in the same lookup. They were not scouted into claims this run: microsoft.ai/models/mai-code-1-1-flash returned 404, and the run did not widen its scope to chase the rest once the Fara bundle was complete.microsoft-fara-7b-releaseFara-7B and Fara-7B-onnx already carry sources in the dataset but no release records. The run confirmed this while placing the Fara 1.5 records and left it alone: adding them is a separate researched claim, not a side effect of this tranche.microsoft-vibevoice-asr-streamingVibeVoice-ASR-Streaming appears in no dataset document at all. Observed during the Microsoft sweep and not scouted, for the same reason as the MAI models.
What this run does not prove
- Coverage is seven reviewed pilot creators, not the catalogue. Nothing here says anything about the 37 organisations that were not probed, and a green run is not a statement that the rest of the dataset is current.
- The panel reviewed the claims, not the world. Unanimous acceptance means three rubrics found each claim properly sourced against the bytes fetched on 2026-09-03; it does not mean the creator has not changed the page since, and it does not verify facts no source asserts.
- Both Fara records leave derivedFromIds empty although the cards name a base model in prose. That under-claims the lineage rather than over-claiming it, and it matches the committed 27B record, but a reader should not read the empty edge as evidence that no lineage exists.
- The measuredDrift figures this run passed against were re-recorded by PR #830, not by this run. Its margin is real but inherited: on the tree this run gated, /tree sat 5113 bytes and /compare 7110 bytes inside a 2% tolerance, so the next tranche of comparable size meets the same wall #833 describes unless something re-records again first. Trunk growth eats that margin without anyone deciding to spend it.
- gate-ledger is absent from npm run validate and from CI. This run reached a full green verification set without it and was wrong to; the entry exists because a reviewer ran the gate the run had not.
Follow-ups — proposed, not fixed
- The four MAI models and MAI-Image-2.6 carry no release records, and microsoft.ai/models/mai-code-1-1-flash returns 404 — the model pages may have moved, which is worth establishing before scouting them.
- Fara-7B and Fara-7B-onnx have sources but no release records.
- VibeVoice-ASR-Streaming is absent from every dataset document.
- qwen.ai is reachable, creator-operated and unapproved; whether it belongs in the approved origin catalogue is a decision for a human, not something a run should settle for itself mid-tranche.
- raw.githubusercontent.com is unapproved while github.com is approved, which makes README evidence needlessly awkward to cite and is worth a deliberate decision either way.
- scripts/run-check.test.ts has a load-dependent failure: "runs --root . and --root=. for real" fails under full-suite contention between two real astro check subprocesses and passes 15 of 15 in isolation. It was green in CI at the same SHA and is not attributable to this run.
2026-09-01-c3f81a
Long-tail depth 2026-09-01 — Cohere, TII, NVIDIA
Scope requested: Three creators already in the catalogue, each holding exactly one family — Cohere, TII and NVIDIA — targeted at one added family plus at least one release each. Deliberately excluded by the coordinating brief and untouched here: ai2-molmo and the other family ids held by in-flight issue #740, the four held by #751, and Stability AI, which is reserved to #754. No id proposed by this run collides with any of those.Published4 edits posted · 10 items withheldResearched a second model family for three creators that each hold exactly one — Cohere, TII and NVIDIA — and published one of the three. Fourteen claims were proposed across three bundles, all of kind add, carrying 44 evidence entries; every quote was machine-checked as a contiguous verbatim substring of the exact bytes its contentHash names, with positive and negative controls on the harness, and all 44 passed. The panel accepted eleven claims unanimously and blocked three, and the three blocks are what decided the run. The provenance rubric rejected the Falcon 3 release because its inputModalities and outputModalities were mapped to text from no quote at all: the card states no modality anywhere, so there was nothing to transcribe and the remedy of attaching a quote does not exist for that page. The editorial rubric rejected the Cohere model card and the Cohere release on entity attribution, holding that tools/updater/profiles/origins/cohere.json records huggingface.co/CohereLabs in deferred_origins as not approved, that resolving whether that org is Cohere's own voice is a human catalogue decision, and that this run did not take it. Because releaseSchema requires both modality fields and accessType with no unknown member, neither rejection could be repaired by dropping an optional field. Each rejection removed the creator's only release, which would have left an empty family that validateDataset refuses, so Cohere and TII were withheld whole — seven further claims that the panel had accepted unanimously were dropped after acceptance rather than applied and orphaned, per the scout contract's pairing rule. NVIDIA cleared unanimously on all four claims and is the run's entire output: the Nemotron Nano 2 family and its Nemotron Nano 9B v2 release, with two supporting sources. Both records carry status unknown, which under ADR 0008 records that the card states no lifecycle state and is a sourced value rather than a fallback; the family and its only release agree at unknown, so the run adds no family whose status outruns its releases. The release's accessType open-weight rests on the ADR 0009 composite — a Hugging Face Hub API record reporting the repository ungated and enabled with its safetensors shards enumerated — and not on the licence name, which states nothing about whether weights can be downloaded.
- PreflightRanRe-measured the base from the committed data rather than quoting the brief. The worktree started at trunk but trunk moved during planning, so with zero commits of its own and no divergence the branch was fast-forwarded to refs/remotes/origin/main at b435c69056; that fast-forward is recorded as an assumption in the run summary. At that anchor the dataset held 71 families, 100 releases and 239 sources, schema.ts carried unknown in lifecycleStatus, and none of cohere-command-r, tii-falcon-3 or nvidia-nemotron-nano-2 existed. The reviewed-profile set at tools/updater/profiles holds seven creators and none of these three, so gate-evidence derives long-tail for all three bundles and the declared policy matches the derived one.
- ScoutRanStored 18 page bodies and cut every quote out of those exact bytes between two literal markers, so a non-verbatim quote throws at build time rather than being emitted. Fourteen claims across three bundles, 44 evidence entries, all retrieval fetch and no search snippets. Two fields confirmed optional in schema.ts were dropped rather than sourced: maximumOutput on the Cohere release and parameters.totalBillions on both the Falcon 3 and Nemotron releases, where the size appears only in the model name. The Cohere family date came from a Hub createdAt with dateBasis platform-repository-created, because firstReleaseDate is required and Cohere's own docs date only the August 2024 update rather than the line.
- ReviewRanThree reviewers, one per rubric, run in parallel and independently: none saw the scout's reasoning, another reviewer's verdict, or any statement of the threshold or the running tally. 42 verdicts over 14 claims. Each was pointed at web/src/data/schema.ts as the definition of the controlled vocabulary, because .github/skills/modeltree-review/SKILL.md has drifted from ADR 0008 and still states that no field has an unknown member; the drift is recorded as a follow-up rather than edited here, since SKILL.md is outside the qualifying class gate-scope enforces. No reviewer that rejected was re-run. The editorial reviewer was asked once to re-transcribe rationales lost to output truncation, with its votes stated as final and closed to revision; it returned wording only.
- GatesRangate-evidence and gate-source-approval were run before any dataset file was touched. gate-evidence exited 0 on the published NVIDIA bundle and, run over the full Cohere and TII bundles, exited 1 naming exactly the three claims the panel had blocked — an independent confirmation of the withholding rather than the run's own word for it. gate-dataset and gate-scope both exited 0 after the change was applied. gate-ledger is the one gate this branch cannot pass by itself and the reason is structural, not a defect in the data: a published entry must name its pull request, and a dock does not open one.
- PublishNot runDeliberately not run. The dock's boundary ends at a reviewable commit: it does not open a pull request, does not merge, does not rebase and does not gate its own work. The dataset change is committed to the branch and this entry is handed to the publishing step to commit second, once the pull request exists to be named — the same two-commit shape run 2026-09-01-b41087 used. Landing remains GitHub's once CI is green.
What was found
- Scouts
- 3
- Pages fetched and hashed
- 18
- Claims proposed
- 14
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims cohere long-tail 3-of-3 5 tii long-tail 3-of-3 5 nvidia long-tail 3-of-3 4 What those claims proposed to do Kind Count Effect Add 14 Four applied to the dataset; ten withheld, three of them blocked by the panel and seven dropped after unanimous acceptance to avoid orphaning them. Not covered
- No second Falcon 3 or Command R variant was scouted after the first was blocked. Re-scouting the same creator for a page that scores better is vote-rigging however it is framed, and the Falcon 3 block in particular is a property of the card rather than of the variant chosen.
- The build.nvidia.com API Catalog channel the Nemotron card names was not read, so the release records accessType open-weight rather than both. That under-claims rather than over-claims.
- Link health was not swept and the second Python interpreter was not exercised; ci-preflight names both as outside what it covers.
What was evaluated
- Reviewers
- 3
- Verdicts cast
- 42
- Accepted by panel
- 11
- Rejected by panel
- 3
Deterministic gates and required checks — 6 of 7 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidencethe published NVIDIA bundle at .modeltree-refresh/runs/2026-09-01-c3f81a/published/nvidia.claims.json 0 Pass4 claims admissible under the long-tail policy, all 4 applying to the dataset. The policy was derived from the reviewed-profile set rather than taken from the bundle, and the derived long-tail matched the declared long-tail, so the unanimous 3-of-3 threshold is the one actually applied. gate-evidencenot requiredthe full Cohere and TII bundles, run to confirm the withholding mechanically 1 FailExit 1 on both, naming cohere-source-command-r-08-2024-card-add, cohere-release-command-r-08-2024-add and tii-release-falcon-3-7b-instruct-add as marked add but reaching only 2 of 3 required accepts. This is the expected and wanted result: it is the gate independently identifying the same three claims the panel blocked, and none of them is in the commit. Recorded as not required because these bundles are not what the run published. gate-source-approvalthe published NVIDIA bundle against the approved origin set at merge-base b435c69056 0 Pass13 citations rest on approved sources: 1 inherited from the dataset at the merge base and 2 proposed on already-trusted origins. No origin was widened and none needed to be. gate-datasetweb/src/data after the four claims were applied 0 PassAll gates passed over 545 records. In particular the new family carries a release, so the empty-family refusal that withheld Cohere and TII does not bite here, and the family's status unknown agrees with its only release. gate-scopethe branch against its computed merge-base with refs/remotes/origin/main 0 PassOnly families.json, releases.json and sources.json changed, all three inside the ADR 0003 qualifying class. No test, workflow, skill or document was touched: the SKILL.md drift found during review was left alone precisely because editing it would have taken the change out of class. npm run validateweb/, tests plus Astro and TypeScript diagnostics 0 PassGreen, with a green baseline measured on the unmodified base first so that any failure would have been attributable to this change. No pinned-enumeration test needed editing: the dataset-count assertions are all lower bounds. ci-preflightthe repository root, selecting the pull-request checks this branch's diff triggers 0 PassRun from the repository root because npm run validate reads only web/ and would not have exercised the checks a diff outside it triggers. This diff is dataset-only, so the selection is narrow. Posted 4 edits
4 edits across 3 documents, a net change of 4 records.
Dataset documents this run changed Document Before After What changed sources.json239 241 Two sources added: the Nemotron Nano 9B v2 model card and its Hugging Face Hub API record. families.json71 72 One family added: nvidia-nemotron-nano-2. releases.json100 101 One release added: nvidia-nemotron-nano-9b-v2. Records added
- Nemotron Nano 2in the treefamilies
nvidia-nemotron-nano-2New NVIDIA family, kept distinct from the earlier Nemotron-4 line. firstReleaseDate 2025-08-18 is the card's own stated release date and carries no dateBasis; status unknown. - Nemotron Nano 9B v2passportreleases
nvidia-nemotron-nano-9b-v2First release under that family. status unknown, accessType open-weight from the ADR 0009 platform composite, contextWindow 128000, modalities quoted from the card's Input Type(s) and Output Type(s) lines.
Not posted 10 items
Rejected by the review panel
tii-release-falcon-3-7b-instruct-addreleases record tii-falcon-3-7b-instruct for tii. [provenance] inputModalities ['text'] and outputModalities ['text'] are unsupported: the card contains no modality statement of any kind and no modality quote is attached, so mapping text/text is inference from the word LLM rather than a recording step. Every other field on the record — releaseDate 2024-12, status unknown, contextWindow 32000, the licence block with osiApproved false, and accessType open-weight on the ADR 0009 composite — was found properly sourced, which is what makes the single defect decisive: releaseSchema requires both modality fields and neither is optional, so the record cannot be written without them.Blocked byrubric:provenancecohere-source-command-r-08-2024-card-addsources record cohere-command-r-08-2024-model-card. [editorial] The claim attributes the card to publisherId cohere as Cohere's own voice, but tools/updater/profiles/origins/cohere.json records huggingface.co/CohereLabs in deferred_origins as not approved, reasoning that nothing there establishes it is Cohere's own org rather than a name anyone could take, and that resolving it is an organization mapping for a human. The run's own notes conceded that decision was not taken. The provenance rubric reached the opposite conclusion on the strength of an approved docs.cohere.com page naming the CohereLabs org, so this is a genuine split rather than a weak claim, and under the unanimous long-tail bar the split blocks it.Blocked byrubric:editorialprofile:cohere deferred_originscohere-release-command-r-08-2024-addreleases record cohere-command-r-08-2024 for cohere. [editorial] Its creator-authored facts — accessType both with the open-weight half, the CC-BY-NC licence block, parameters.totalBillions 32 and the modality lines — rest solely on the model card whose attribution the same rubric rejected, so accessType both over-attributes open weights to Cohere when only the docs' Live-on-API row is solidly Cohere's voice. It separately found variant 'Standard' to be a tier no source states. The unsourced maximumOutput 4000 the previous run carried was already dropped rather than sourced.Blocked byrubric:editorialclaim:cohere-source-command-r-08-2024-card-add
Accepted by the panel, then dropped
cohere-family-command-r-addfamilies record cohere-command-r. Accepted 3-of-3 on its own merits: firstReleaseDate 2024-03-11 with dateBasis platform-repository-created, status unknown, and sources that do not include the disputed card. Dropped because its only proposed release was blocked, which would leave a family with zero releases — the condition validateDataset refuses with 'family has no releases'. Publishing the family alone would have been incoherent, so family and release stand or fall together.Blocked byclaim:cohere-release-command-r-08-2024-addcohere-source-command-r-docs-addsources record cohere-command-r-docs, accepted 3-of-3 on an approved docs.cohere.com origin. Dropped because the only records citing it were the withheld Cohere family and release; applying it alone would leave an uncited source in the dataset.Blocked byclaim:cohere-family-command-r-addcohere-source-command-r-v01-hub-record-addsources record hugging-face-cohere-command-r-v01-hub-record, accepted 3-of-3 as a platform-authored ADR 0009 record. Dropped for the same orphaning reason: only the withheld Cohere family cited it.Blocked byclaim:cohere-family-command-r-addtii-family-falcon-3-addfamilies record tii-falcon-3. Accepted 3-of-3, with firstReleaseDate 2024-12 taken from the card's creator-stated Model Release Date: December 2024 and status unknown. Dropped because its only proposed release was blocked on modalities, leaving an empty family. This is the closest miss of the run: the family claim itself has no known defect.Blocked byclaim:tii-release-falcon-3-7b-instruct-addtii-source-falcon3-7b-instruct-card-addsources record tii-falcon3-7b-instruct-model-card, accepted 3-of-3. Dropped because both records citing it were withheld.Blocked byclaim:tii-family-falcon-3-addtii-source-falcon3-announcement-addsources record tii-falcon3-announcement, accepted 3-of-3 as an official-announcement in TII's first-person voice. Dropped for orphaning.Blocked byclaim:tii-family-falcon-3-addtii-source-falcon3-7b-instruct-hub-record-addsources record hugging-face-falcon3-7b-instruct-hub-record, accepted 3-of-3 as a platform-authored ADR 0009 record. Dropped for orphaning.Blocked byclaim:tii-release-falcon-3-7b-instruct-add
What this run does not prove
- One creator of three landed. The issue's exactly-one-family count moves by one, not by three, and this run does not by itself satisfy #689.
- status unknown is honest but it is not informative. Three of the dataset's families now sit at unknown plus this one, and a reader looking for whether Nemotron Nano 2 is current will not find that here, because NVIDIA's card does not say.
- accessType open-weight on the Nemotron release rests on a platform record reporting an ungated repository with safetensors shards, not on a statement by NVIDIA that weights are downloadable. That is the ADR 0009 composite and the panel accepted it unanimously, but it is a weaker basis than a creator sentence saying so, and it can go stale if the repository is later gated.
- The Cohere block is an unresolved human decision, not a finding that Cohere's card is wrong. Two rubrics reached opposite conclusions on whether huggingface.co/CohereLabs is Cohere's voice, and until a human updates that catalogue entry the question stays open and the evidence stays unused.
- The panel is three instances of one model family reading the same pages, so it buys independence of reasoning and not independence of training. A page that is itself wrong could carry all three.
- gate-ledger cannot pass on this branch as the dock leaves it, because this entry is not committed and cannot be until the pull request it must name exists.
Follow-ups — proposed, not fixed
- tools/updater/profiles/origins/cohere.json defers huggingface.co/CohereLabs as not approved on the ground that nothing establishes it as Cohere's own org. This run found first-party evidence bearing on exactly that: docs.cohere.com, an approved Cohere origin, links to a CohereLabs collection and describes it as where Cohere publishes open-weight models on Hugging Face. A human should decide whether that settles the mapping; the Cohere family and release are fully scouted and re-runnable the moment it does.
- .github/skills/modeltree-review/SKILL.md states that none of the controlled-vocabulary fields has an unknown member and lists five members of lifecycleStatus. ADR 0008 added a sixth, unknown, and web/src/data/schema.ts carries it. The rubric text and the schema now disagree on a point that decides claims, and reviewers in this run had to be pointed at the schema to resolve it. Editing SKILL.md was deliberately not done here because it is outside the ADR 0003 qualifying class and would have taken the dataset change out of class.
- The Falcon3 model card states no modality anywhere, which blocked an otherwise clean release. Whether TII states modalities on another Falcon 3 surface is unknown; a future run would need a different page, not a different reading of this one.
2026-09-01-b41087
Long-tail sweep: all 40 creators re-verified
Scope requested: All 40 creators in organizations.json. Seven carry a reviewed profile in tools/updater/profiles and were judged at 2-of-3 — alibaba-cloud, amazon, anthropic, google-deepmind, meta, microsoft, openai — and the other 33 were judged at the unanimous 3-of-3 long-tail bar. Predominantly re-verification: asking whether recorded facts still hold, and advancing the dates that say when they were last checked. Discovery was attempted only where a creator index page made it cheap, which surfaced exactly one unrecorded release. The model-fit guidance layer was deliberately left alone: all seven statements were re-verified on 2026-08-31, so the five records they pin needed nothing this run.Published179 edits posted · 5 items withheldAn agent-run, source-backed refresh under ADR 0003, and the breadth run the previous entry asked for. Run 2026-08-31-ae0342 traded breadth for depth, scouted six pilot creators and recorded the untouched long tail as the largest gap in its own entry; this run scouted every one of the 40 creators in organizations.json against 233 pages fetched and hashed during the run, 7 at the 2-of-3 pilot threshold and 33 at the unanimous 3-of-3 long-tail threshold. 202 claims were proposed and three rubrics voted blind on every one, casting 606 verdicts. 199 met their threshold, 197 were applied, and 5 were withheld with their rationales published. Not one claim was carried over a dissent, which is not a sign the rubrics went easy but a fact about this run being almost entirely re-verification, where there is little to disagree about: all four rejections landed on the three claims that attempted something harder than confirming a date. Run 2026-08-30-c0b6e9 also recorded zero dissents, so this is not a first. 179 field edits landed across three documents — 89 source lastCheckedDate values and 90 record verifiedAt values — and no record was added or removed. The one genuinely new release found this run, Gemini 2.5 Flash-Lite, was refused by the panel and is not in the dataset. No human reviewed any claim, verdict or edit.
- PreflightRanTree clean, gh authenticated, and no open pull request from a previous refresh when the run started. Run id 2026-09-01-b41087, with the run directory under .modeltree-refresh/ confirmed git-ignored before anything was written to it. Both npm and drydock were probed in both shim forms before either was relied on or ruled out: the bare names resolve to PowerShell shims this machine’s execution policy refuses, while npm.cmd reports 11.9.0 and drydock.cmd reports 0.1.0. Both are installed and blocked in bare form rather than absent, and the execution policy was not changed to work around it — that would be a machine-wide security change made to satisfy a probe. Dependencies were installed with npm ci after npm run validate failed on a missing vitest; package-lock.json is unmodified.
- ScoutRan240 URLs were attempted and 233 returned 200 and were hashed; the 7 failures are recorded under notCovered rather than retried until they looked like successes. Every quote was cut as a byte-exact slice of a page body saved during this run and asserted as a substring of those bytes by the builder, which throws rather than emit an unverifiable pair — so no claim here rests on recollection or on a search snippet. Discovery findings were genuinely mixed and both directions are recorded: OpenAI’s and Anthropic’s model catalogues list nothing the dataset does not already carry, which is a real negative result rather than an unchecked one, while Google’s listed Gemini 2.5 Flash-Lite, which the dataset lacks.
- ReviewRanThree rubrics — provenance, consistency, editorial — voted independently on all 202 claims across 9 reviewer invocations, each in its own context and each given only its own rubric text, the claim, the evidence and the relevant dataset slice. No reviewer saw another rubric’s verdicts, the running tally, the threshold, or the scout’s reasoning. 606 verdicts were cast, every one with a rationale, and all are published on the pull request. Four rejections were returned and every one was left to stand: no claim was edited and re-reviewed to chase a verdict, no reviewer was re-run, and the coordinating agent cast no vote and overruled none — including the one rejection whose stated premise is demonstrably wrong, which is recorded in the caveats and filed as a follow-up instead.
- GatesRangate-evidence and gate-source-approval ran across all 40 gated bundles before anything was applied, since gate-source-approval anchors on the committed dataset and running it afterwards would be asking the run’s own writes whether the run’s own sources are trustworthy. gate-dataset, npm run validate, gate-scope, gate-ledger and ci-preflight.mjs ran afterwards. origin/main moved during the review stage, so the branch was re-anchored from d9c5c403 onto 329a719f and both pre-apply gates were re-run against the new anchor rather than carried over. No gate was skipped, forced, or re-run to obtain a different answer, no threshold was lowered, and nothing was pushed to main.
- PublishRanOpened as pull request #752 carrying the full evidence trail — every claim, every quote with the SHA-256 of the bytes it was cut from, and all 606 rationales. The trail is roughly 260 KB against GitHub’s 65 536-character body limit, so the body carries the summary, the gate results, the source-approval block and all five withheld claims in full, and the per-claim detail follows as five comments on the same pull request rather than being summarised away. Merged by GitHub via --auto --squash once web-ci went green; the agent did not merge and did not use --admin. The data and this entry are two commits on the branch and one commit on main, so a bad run is still one revert.
- DeployNot runNot yet run when this entry was written, and recorded that way rather than predicted. The entry necessarily ships in the commit whose merge triggers pages.yml, so it cannot report the outcome of its own deploy without guessing. The run waits for the deploy after the merge, checks it against the merge SHA, and reverts by pull request if it failed; the result is reported in the run’s summary issue, which is where a reader should look to close this line.
What was found
- Scouts
- 40
- Pages fetched and hashed
- 233
- Claims proposed
- 202
Claims per creator bundle, with the review threshold its profile set Creator Policy Threshold Claims 01-ai long-tail 3-of-3 6 ai-singapore long-tail 3-of-3 4 ai2 long-tail 3-of-3 5 ai21-labs long-tail 3-of-3 4 aleph-alpha long-tail 3-of-3 4 alibaba-cloud pilot 2-of-3 6 amazon pilot 2-of-3 5 anthropic pilot 2-of-3 6 apple long-tail 3-of-3 4 baidu long-tail 3-of-3 5 bytedance-seed long-tail 3-of-3 4 cohere long-tail 3-of-3 5 databricks long-tail 3-of-3 4 deepseek long-tail 3-of-3 6 eleutherai long-tail 3-of-3 6 google-deepmind pilot 2-of-3 9 hugging-face long-tail 3-of-3 5 ibm long-tail 3-of-3 6 lg-ai-research long-tail 3-of-3 4 liquid-ai long-tail 3-of-3 4 meta pilot 2-of-3 6 microsoft pilot 2-of-3 6 minimax long-tail 3-of-3 6 mistral-ai long-tail 3-of-3 6 moonshot-ai long-tail 3-of-3 3 naver long-tail 3-of-3 5 nous-research long-tail 3-of-3 5 nvidia long-tail 3-of-3 5 openai pilot 2-of-3 6 reka-ai long-tail 3-of-3 5 sakana-ai long-tail 3-of-3 4 sarvam-ai long-tail 3-of-3 5 snowflake long-tail 3-of-3 5 stability-ai long-tail 3-of-3 5 tencent long-tail 3-of-3 4 tii long-tail 3-of-3 4 upstage long-tail 3-of-3 4 xai long-tail 3-of-3 6 xiaomi long-tail 3-of-3 5 zhipu-ai long-tail 3-of-3 5 What those claims proposed to do Kind Count Effect Change 107 A source page re-read this run, moving its lastCheckedDate to 2026-09-01. All 107 were accepted and applied, moving 89 distinct source records — 18 of the claims re-assert a shared licence page that several creators cite (osi-approved-licenses, osi-license-mit), and a shared source moves once no matter how many bundles confirm it. Unchanged 92 A recorded fact re-read from its primary source and found to still hold. 90 met their threshold and applied, advancing 50 release verifiedAt dates and 40 family verifiedAt dates to 2026-09-01. Two were refused by the panel and withheld, so those two records keep their previous dates. Add 3 One new release, Gemini 2.5 Flash-Lite, with the two source records it would have cited. None reached the dataset: the release was refused 1-of-2 by the panel, and the two sources — accepted on their own merits — fell with it, because a source no record cites is dead provenance that validateDataset refuses to load. Nothing was added or removed this run. Not covered
- ai.meta.com returned HTTP 400 to every request from this runner, across six URLs and both fetch paths tried. That is a status rather than a network failure, so it is a fact about how that host answers this client and not evidence the pages are gone. Four recorded sources live there — meta-llama-4-announcement, meta-muse-spark-announcement, meta-muse-spark-1-1-announcement and meta-muse-image-video-announcement — and none carries a verification date from this run. Meta itself was still covered, at 6 claims, through huggingface.co and github.com, which the dataset already trusts.
- abc.xyz returned HTTP 403. The one source on it, alphabet-about, was last checked 2026-08-15 and is now the oldest unverified source in the dataset. It is a corporate about-page backing a publisher record rather than a model fact, so nothing about a model rests on it, but it will keep aging until either the host answers this client or a human re-reads it.
- This run is re-verification first. It asks whether recorded facts still hold; it does not systematically ask what each creator has shipped that the dataset has never heard of. Discovery was attempted only where a creator index page made it cheap. So a green result here says the dataset is not stale; it says very little about whether it is complete.
- The known breadth gaps are unchanged and were not worked this run: OpenAI’s realtime, transcribe and tts models, the roughly twenty Cohere Command models against the one recorded, Microsoft’s MAI-Code-1.1-Flash, MAI-Image-2.5, MAI-Voice-2 and MAI-Transcribe-1.5, and Amazon’s Nova Act, Nova Forge and Nova Multimodal Embeddings. Each is breadth work with its own sourcing burden.
- A retirement-date conflict was found and deliberately not resolved. The Gemini API deprecations page says no shutdown date is announced for the 2.5 family, while the Cloud platform page gives a retirement date of 20 October 2026 for the same models. No dataset field holds either figure, so there was nothing to claim and nothing to withhold; it is reported here so the conflict is on the record rather than discovered again next run.
What was evaluated
- Reviewers
- 9
- Verdicts cast
- 606
- Accepted by panel
- 199
- Rejected by panel
- 3
Deterministic gates and required checks — 8 of 8 checks passed. Exit 0 is a pass; exit 2 means the gate could not run and is never treated as one. Check Scope Exit Result gate-evidence.mjs40 claim bundles 0 Pass197 claims admissible with complete three-rubric panels across all 40 gated bundles. Every bundle declared its policy and the gate independently re-derived the threshold from tools/updater/profiles, so no bundle’s declared policy went unchecked — which matters more than usual this run, since 33 of the 40 bundles claimed the stricter bar. gate-source-approval.mjs40 claim bundles 0 PassZero proposed sources across all 40 bundles: every cited source already existed in the dataset at anchor 329a719f, on one of the 46 approved origins derived there from 225 dataset sources and 18 catalogues across 19 profile files. The run therefore approved no source of its own. The two sources it did propose had already been withdrawn by the pairing cascade before this gate saw the gated bundles. Run before any edit was applied, and re-run after the re-anchor. gate-dataset.mjsweb/src/data 0 PassZero failures across the documents raw.ts composes, after the 179 edits were applied. npm run validateweb/ 0 PassTests green and astro check reported 0 errors and 0 warnings across 246 files. It failed on first invocation because this worktree had no node_modules — a missing vitest, not a data fault — and passed after npm ci. That first failure is recorded rather than quietly overwritten, because "the checker could not run" and "the data is good" are different findings and only one of them is a pass. gate-scope.mjsbranch vs merge-base 0 Passchanged: 3, empty: false, outOfClass: []. Anchored on 329a719f, the merge base the gate computed rather than one supplied to it — requestedBase was null. The three are sources.json, releases.json and families.json; this ledger is the fourth path in the qualifying class and is covered by gate-ledger instead. gate-ledger.mjsthis entry vs the diff it describes 0 Passtranscription: false — this entry was written by the run it describes and ships in the same pull request as the data, so its record counts were reconciled against the actual diff rather than taken on trust. That is the gap issue #419 recorded on three consecutive runs. ci-preflight.mjsnot requiredselected from the branch diff 0 PassSelected and ran the three check groups this diff triggers — web-ci, skills-ci and source-link-health-tests — locally from the repository root before the branch was pushed. It reports its own blind spots on every run, including the networked link-health sweep and the second Python interpreter. web-cithe pull request — PassGitHub performed the merge via --auto on this check going green; the agent did not merge. Posted 179 edits
179 edits across 3 documents, a net change of 0 records.
Dataset documents this run changed Document Before After What changed sources.json225 225 89 lastCheckedDate values moved to 2026-09-01, one per source page successfully re-read this run. No source was added or removed: gate-source-approval recorded zero proposed sources across all 40 bundles, so every date moved on a record that already existed. 107 accepted claims produced these 89 edits, because 18 of them re-assert one of two shared licence pages that several creators cite independently. releases.json96 96 50 verifiedAt dates moved to 2026-09-01, each backed by a release-level fact re-read from its primary source and found unchanged. Two releases that were scouted did not move because the panel refused their claims — hugging-face-smollm3-3b and ibm-granite-4-0-h-small, both listed under withheld. No release was added; the one new release found this run was refused. families.json67 67 40 verifiedAt dates moved to 2026-09-01 — one family per creator, each re-read and found unchanged. No family was added or removed and no family status value changed. Not posted 5 items
Rejected by the review panel
releases/google-gemini-2-5-flash-lite (proposed, not added)The only genuinely new release this run found, and it was refused 1-of-2 under the pilot bar. Provenance: the record sets accessType to proprietary-hosted, but none of the five cited quotes says anything about how the model is released or whether weights are downloadable, so that field had no quoted fact behind it. Editorial: the record puts the marketing name in canonicalName and the API model id in apiAliases, which it read as inverting the creator profile’s naming rule. Consistency accepted it. The release is not in the dataset and is left for a future run to propose properly rather than being trimmed until it passed. Note the caveat about the editorial rationale: its premise about existing Gemini records does not hold, and it was still not overruled.Blocked byprovenance rubriceditorial rubricreleases/hugging-face-smollm3-3b.verifiedAtReached 2 of the 3 accepts a long-tail creator requires. Provenance objected that the quote is a YaRN rope-scaling config comment showing the arithmetic 2 x 65536 = 131072, which does not by itself state the context window; the card does say so in prose nearby, but that sentence was not the quote. A stricter bar than the pilot creators face, applied as written. The record keeps its previous verifiedAt.Blocked byprovenance rubricreleases/ibm-granite-4-0-h-small.verifiedAtReached 2 of 3. Editorial objected not to the claim but to the record it would have re-verified: the existing summary opens "The largest Granite 4.0 model", which ranks models within a family, and this project forbids ranking prose. That is a pre-existing defect in committed data rather than anything this run proposed, and the rubric was right to refuse to bless it. Fixing the prose is a separate change and is filed as a follow-up; the record keeps its previous verifiedAt until then.Blocked byeditorial rubric
Accepted by the panel, then dropped
sources/google-gemini-2-5-flash-lite-docs (proposed, not added)Accepted unanimously on its own merits and still withdrawn, because the only record that would have cited it was refused. validateDataset treats a source no record references as dead provenance and refuses to load the dataset, so a source-add and the record citing it stand or fall together. Withdrawn by the run rather than by a reviewer.Blocked bypairing with the withheld release recordsources/google-gemini-2-5-flash-lite-platform-docs (proposed, not added)Accepted unanimously and withdrawn for the same reason as its sibling: the release that would have cited it was refused, and an uncited source is dead provenance.Blocked bypairing with the withheld release record
What this run does not prove
- No human reviewed any claim, verdict or edit in this run before it merged and deployed. That is what ADR 0003 authorises and what it costs.
- The three rubrics are three instances of the same model family reading the same page, so they share a failure mode: a source that is itself wrong can carry all three. The panel defends against unevidenced inference, not against a primary source that is confidently mistaken.
- This run is broad and shallow by design. It asks whether recorded facts still hold, not whether each creator has shipped something the dataset has never heard of. Only one genuinely new release surfaced, and it was withheld — so a green run here is weak evidence about completeness and strong evidence only about staleness.
- pagesFetched counts the 233 pages that returned 200. Seven further URLs were attempted and failed and are excluded from that figure rather than counted as reads.
- The editorial rubric's objection to the Gemini 2.5 Flash-Lite record asserted that the record inverts a naming convention 'every other Gemini entry' follows. That premise is factually wrong — google-gemini-2-5-flash already carries the marketing name in canonicalName. The claim was still withheld, because the chair does not vote and does not overrule a rubric; the error is recorded here and filed as a follow-up rather than corrected into a pass.
- reviewers and verdictsCast are derived rather than quoted: 3 rubrics x 3 batches = 9 reviewer invocations, and 202 claims x 3 rubrics = 606 verdicts, which reconciles with the nine per-rubric verdict files.
- The deploy had not run when this entry was written, and the entry says so rather than predicting it. A reader checking whether the site actually rebuilt should read the run’s summary issue, not this stage note.
- Zero claims carried over a dissent is a cleaner number than it looks. 197 of the 202 claims were re-verification of facts already in the dataset, where the three rubrics have little room to disagree; all four rejections landed on the three claims that were doing something harder. Read it as "this run attempted little that was contentious", not as "the panel found nothing to argue about".
- This entry was corrected in place shortly after it published, and the corrections are listed here rather than made silently. As first written it recorded issue #419 as open when it had closed on 2026-08-31; it called a zero-dissent panel "the cleanest panel result the ledger records" when run 2026-08-30-c0b6e9 had also recorded zero; and it twice miscounted the four rejections as landing on four and on five claims when they landed on three. All three were errors in this entry's own prose, not in the data the run published, and the dataset edits were unaffected. They are worth noting because of where they occurred: the ledger entry is the one document in the qualifying class that no reviewer rubric reads, so nothing in the panel or the gates was ever going to catch them. The gates checked this entry's record counts against the diff, which is exactly what they promise and no more.
Follow-ups — proposed, not fixed
- The editorial rubric refused the Gemini 2.5 Flash-Lite record on the ground that it inverts a naming convention every other Gemini entry follows. The premise is wrong: google-gemini-2-5-flash already carries the marketing name in canonicalName and the API id in apiAliases, exactly as the refused record did. The verdict was left standing because a chair that overrules a rubric it disagrees with is not running a panel, but the underlying question is real and unresolved — either the profile naming rule or the committed Gemini records are wrong, and until one is fixed this release will keep being refused for a reason that does not hold.
- ibm-granite-4-0-h-small’s summary opens with "The largest Granite 4.0 model", a size ranking inside committed data that the editorial rubric refused to re-verify. It will keep failing review every run until the prose is fixed, and it is worth checking whether other records carry similar superlatives, since nothing scans for them.
- ai.meta.com answers this runner with HTTP 400 on every URL and every fetch path tried, which is a different failure from the unreachable ai.google.dev of the previous run and may be client fingerprinting rather than a block. Four recorded sources are behind it. If it does not clear on its own, those sources need either a working fetch approach or an alternative approved origin, or they will keep aging out.
- The review panel was run as 9 invocations over 3 batches rather than one per creator, which is what made 202 claims affordable. It is worth deciding deliberately whether batching weakens the independence guarantee: a reviewer judging 76 claims in one context can see patterns across them that a per-creator reviewer cannot, which cuts both ways — more consistency, but also more room for one framing to carry a whole batch.
- This run advanced 89 source dates and 90 record dates without a single new fact reaching the dataset. That is a healthy outcome for a re-verification run, but two consecutive runs like it would mean the dataset is being kept fresh rather than kept complete. The next run should lead with discovery, and the ledger should probably distinguish the two modes explicitly rather than leaving a reader to infer it from claimsByKind.