Guidance is conditional, never a verdict
Every statement is filed as one of three kinds — good fit when, trade-off, or avoid when — and each carries the condition it applies under. There is no fourth, unconditional kind, no composite score, and no comparison against other models.
What the wording check does, and what it cannot do
Validation refuses a fixed list of vocabulary in ModelTree’s own editorial text: superlatives, best-in-class and go-to framing, beats-everything phrasing, numeric rankings, universal quantifiers, and composite-score wording. That is a vocabulary filter, not a judgement about meaning. It catches the usual ways a verdict gets written down, but a comparative claim phrased around those words would pass it, and it errs toward rejecting borderline wording that an author can simply rephrase. The check that actually holds is provenance, below: a statement may cite only the sources the facts beneath it already cite, so a comparison cannot pull in a source no recorded fact carries. That constrains where evidence comes from, not what a sentence means. Neither check reads a statement and decides whether its facts support it — that judgement stays with you, which is why every statement lists the facts and sources it rests on.
The rubric is disclosed, not weighted
A statement names the dimensions it was derived from, and each dimension must be answered by a fact of a kind that can answer it. Dimensions are never added up or weighed against each other; they exist so a reader can see which questions were asked.
- Context window — How much input and output length does the documentation state?
- Documented limits — What limits and intended uses does the documentation state outright?
- Modality coverage — Which inputs and outputs are documented?
- Access and licensing — How can it be obtained, and what does its licence permit?
- Lifecycle stability — What lifecycle stage do the vendor records place it in?
- Cost structure — What rates are recorded, in which unit and currency?
- Measured benchmark evidence — What did a recorded evaluation measure, under which setup?
- Usage evidence — What source-qualified usage has been observed, over which population?
The evidence threshold
A statement must rest on at least one structured fact already recorded here — a release or family field, a lifecycle event, an evaluation result, a usage observation, or a pricing record — and it must be a fact about this release. It may cite only sources those facts already cite, so guidance cannot pull in a source no recorded fact carries, and it cannot be dated earlier than the evidence beneath it.
Conflicts are shown, not resolved
When two statements about this release contradict each other on the same dimension, both are published and linked to one another. ModelTree does not decide between them, and the underlying disagreement between sources is left visible.
What this cannot tell you
Recorded facts are mostly documentation, and documentation states what an interface accepts rather than how well a model behaves. Where a dimension has no qualifying evidence, it is listed as a gap instead of being filled by inference. The date shown on a statement is the verification date of the newest fact beneath it, not a record that an editor re-read the reasoning; once that evidence is more than 180 days old the statement is marked stale. Nothing here is a recommendation tailored to a reader, and nothing here ranks models.