Stable Video Diffusion (img2vid-XT)
Stability AI's image-to-video latent diffusion model, generating 25 frames of video from a single conditioning image. Released with downloadable weights, distinct from the text-to-image Stable Diffusion family.
What it is
Stability AI's image-to-video latent diffusion model, generating 25 frames of video from a single conditioning image. Released with downloadable weights, distinct from the text-to-image Stable Diffusion family.
When to use it. Generating short videos from a still image; the model is conditioned on an image and cannot be controlled through text. The cited sources disagree on the scope of use: the launch announcement restricts it to research ("not intended for real-world or commercial applications at this stage"), while the model card states use is intended "for both non-commercial and commercial usage" under the Community License.
Model DNA
9 identity dimensions, each read from a single field of this release's record and shown in the same order on every model. Nothing here is a score, a rating, or a ranking, and a dimension the record does not carry says so rather than being left out.
- CreatorStability AI
- FamilyStable Video Diffusion
- Generation1.0
- Tierimg2vid-XT
- InputImage
- OutputVideo
- SpecializationVideo
- AccessOpen-weight
- WeightsDownloadable
What each segment means, and which field it comes from
- Creator
- Stability AIThe organization recorded as having built this release. A creator is not the same entity as a platform that serves the model or a product that ships it.Read from
organizationIdon this release record.How ModelTree keeps creator, model, product and platform separate - Family
- Stable Video DiffusionThe model line this release belongs to, as its creator names it. A family groups releases; it is not itself a release.Read from
familyIdon this release record.How ModelTree keeps creator, model, product and platform separate - Generation
- 1.0The version string the creator published for this release, recorded as written. ModelTree does not renumber, normalise, or order versions.Read from
versionon this release record. - Tier
- img2vid-XTThe variant name the creator gave this release within its family. It is a name, not a rank: what a creator means by it is recorded only where the creator has stated it.Read from
varianton this release record.Why ModelTree publishes no universal ranking - Input
- ImageThe kinds of input this release is documented as accepting.Read from
inputModalitieson this release record. - Output
- VideoThe kinds of output this release is documented as producing.Read from
outputModalitieson this release record. - Specialization
- VideoThe documented focus of this release. Categories are labels, never summed or weighted into a score.Read from
categorieson this release record. - Access
- Open-weightHow this release can be reached, as its creator documents it.Read from
accessTypeon this release record.How ModelTree defines each access type - Weights
- DownloadableWhether the licence record held for this release documents its weights as downloadable. Downloadable weights and an OSI-approved licence are separate claims, so this is not a statement that the release is open source.Read from
licenseon this release record.What “open weight” means
- Creator
- Stability AI
- Family
- Stable Video Diffusion
- Version
- 1.0
- Variant
- img2vid-XT
- Released
- Nov 21, 2023
- Lifecycle status
- Unknown
- Access
- Open-weight
- Record slug
- svd-img2vid-xt
Canonical record
- Canonical name
- stable-video-diffusion-img2vid-xt
- Canonical page
/ModelTree/models/svd-img2vid-xt/- Lifecycle
- Unknown. The creator’s own page states no lifecycle or availability term at all — common for a bare model card that names architecture, weights and licence but not whether the vendor still offers the model. It is the faithful value for “the source does not say”, not a claim that the model is unavailable; the record stays complete and is withheld from no filter.
Also known as
- stable-video-diffusion-img2vid-xt
No API identifier is recorded for this release.
Documented limits
- Input modalities
- image
- Output modalities
- video
- Context window
- Not applicable — this release takes no text input, so there is no token budget to state
- Maximum output
- Not applicable — this release emits no text, so an output limit in tokens counts nothing
- Parameters
- Not recorded
- Video
How you can get it
Open-weight. Model weights can be downloaded. This alone says nothing about the licence: the schema records downloadable weights and OSI-approval as two separate booleans, so open-weight does not imply open-source.
How ModelTree defines access and licensing
Licence
- SPDX identifier
- Not recorded
- Downloadable weights
- Weights are documented as downloadable.
- OSI-approved
- The licence is not recorded as OSI-approved, so this release is not described as open source.
What this passport does not record
These sections are absent from this page because no reviewed record exists for them. They are listed rather than dropped silently, because a missing section and a section ModelTree has decided nothing about look identical otherwise.
- Where it fits
- No predecessor, successor, sibling, or derivation is recorded for this release, and its creator has published no statement of what its variant name is for. ModelTree records a lineage link only where a source states it, and does not infer one from names or release dates.
- Where it is served
- No deployment record ties this release to a serving platform. ModelTree has not yet reviewed platform availability for this record; absence is not a claim that the model is unavailable.
- What it costs
- No published price is recorded for this release. A price is held only against a reviewed deployment on a named platform, with its currency, unit, and effective date; absence is not a claim that the model is free or unpriced.
- What has changed
- No dated release event is recorded for this release beyond its release date. Announcement, availability, and deprecation events are held as separate sourced records, and none has been reviewed for this one.
Who reports using it
No source-qualified usage evidence is recorded for Stable Video Diffusion (img2vid-XT).
ModelTree records a usage figure only when a citable source states the metric, the population it measured, the time window, and the method behind it. Absence here means no qualifying observation has been verified, not that the model is unused.
How ModelTree qualifies usage evidence
Incompatible populations are never merged
Weekly users of one assistant, downloads from one model hub, and routed tokens on one aggregator count different populations. ModelTree groups observations only when the metric, the unit, and the measured population match exactly. Nothing is converted, normalized, weighted, or ranked, and there is no composite popularity score.
Sources are qualified, not scored
Every observation names the exact sources behind it and states whether it is a creator self-report, a platform operator report, an independent measurement, a developer survey, or a community signal. Creator self-reports are kept in their own labelled list because the creator has an interest in the figure; they are still shown, never hidden.
A synthesis needs two independent publishers
A cross-source statement may only be made when at least two non-creator observations from at least two different publishers measure the same metric over the same population. A single-source observation is still published as an observation, but it cannot produce a cross-source statement.
Missing and conflicting evidence stays visible
When no qualifying source exists, this section says so rather than estimating. When two qualifying sources disagree, both readings are shown and labelled as conflicting; neither is dropped and no winner is declared. A figure that has not been re-checked within 180 days is marked stale.
When it fits, and when it does not
Everything in this section is ModelTree editorial synthesis: a reading of facts recorded elsewhere in this dataset, each traced to the sources beneath it. It is conditional guidance for a stated situation. It does not declare this model preferable to another, and no overall verdict is produced here or anywhere else in ModelTree.
No conditional-fit guidance is recorded for Stable Video Diffusion (img2vid-XT).
A statement is published only when it can be derived from specific recorded facts about this release and cite the sources those facts already carry. Absence here means no such statement has been verified, not that the model suits every situation or none.
How ModelTree derives conditional fit
Guidance is conditional, never a verdict
Every statement is filed as one of three kinds — good fit when, trade-off, or avoid when — and each carries the condition it applies under. There is no fourth, unconditional kind, no composite score, and no comparison against other models.
What the wording check does, and what it cannot do
Validation refuses a fixed list of vocabulary in ModelTree’s own editorial text: superlatives, best-in-class and go-to framing, beats-everything phrasing, numeric rankings, universal quantifiers, and composite-score wording. That is a vocabulary filter, not a judgement about meaning. It catches the usual ways a verdict gets written down, but a comparative claim phrased around those words would pass it, and it errs toward rejecting borderline wording that an author can simply rephrase. The check that actually holds is provenance, below: a statement may cite only the sources the facts beneath it already cite, so a comparison cannot pull in a source no recorded fact carries. That constrains where evidence comes from, not what a sentence means. Neither check reads a statement and decides whether its facts support it — that judgement stays with you, which is why every statement lists the facts and sources it rests on.
The rubric is disclosed, not weighted
A statement names the dimensions it was derived from, and each dimension must be answered by a fact of a kind that can answer it. Dimensions are never added up or weighed against each other; they exist so a reader can see which questions were asked.
- Context window — How much input and output length does the documentation state?
- Documented limits — What limits and intended uses does the documentation state outright?
- Modality coverage — Which inputs and outputs are documented?
- Access and licensing — How can it be obtained, and what does its licence permit?
- Lifecycle stability — What lifecycle stage do the vendor records place it in?
- Cost structure — What rates are recorded, in which unit and currency?
- Measured benchmark evidence — What did a recorded evaluation measure, under which setup?
- Usage evidence — What source-qualified usage has been observed, over which population?
The evidence threshold
A statement must rest on at least one structured fact already recorded here — a release or family field, a lifecycle event, an evaluation result, a usage observation, or a pricing record — and it must be a fact about this release. It may cite only sources those facts already cite, so guidance cannot pull in a source no recorded fact carries, and it cannot be dated earlier than the evidence beneath it.
Conflicts are shown, not resolved
When two statements about this release contradict each other on the same dimension, both are published and linked to one another. ModelTree does not decide between them, and the underlying disagreement between sources is left visible.
What this cannot tell you
Recorded facts are mostly documentation, and documentation states what an interface accepts rather than how well a model behaves. Where a dimension has no qualifying evidence, it is listed as a gap instead of being filled by inference. The date shown on a statement is the verification date of the newest fact beneath it, not a record that an editor re-read the reasoning; once that evidence is more than 180 days old the statement is marked stale. Nothing here is a recommendation tailored to a reader, and nothing here ranks models.
Primary sources
- official announcementIntroducing Stable Video Diffusion
Stability AI Ltd. | Published Nov 21, 2023 | Checked Sep 11, 2026
Stability AI Ltd. | Checked Sep 11, 2026
- official docsOSI Approved Licenses
Open Source Initiative | Checked Sep 11, 2026
Found something wrong on this page?Report incorrect data for svd-img2vid-xt. Corrections are filed against the record slug so the fix lands on the right release.
Every fact above is recorded with the primary source it came from and the date it was last checked. This record was last verified on.How we verify and keep facts fresh.