We introduce A2A-Video: an any-to-any multimodal video model that models the world across different representation spaces (e.g., RGB, depth, optical flow, semantic features, bounding boxes, etc.). A2A-Video can predict any modality at any point in time, given any combination of input modalities at any time-steps.
Beyond this prediction capability, the flexible modality-time traversal capability enables higher-level inference instantiations: namely, generation through a chain of modalities, future prediction in multiple modalities, and counterfactual generation with modalities as actions.
We introduce A2A-Video: an any-to-any multimodal video model that models the world across different representation spaces (e.g., RGB, depth, optical flow, semantic features, bounding boxes, etc.), each capturing a different aspect of the underlying observation. A2A-Video can predict any modality at any point in time, given any combination of input modalities at any time-steps.
Beyond this prediction capability, the flexible modality-time traversal capability enables higher-level inference instantiations:
A2A-Video allows predictions from any input modality to any target modality with a single model. Below, we demonstrate this by showing all-to-all prediction, where A2A-Video sequentially predicts each modality from a single input.
A2A-Video extends the any-to-any multimodal framework of 4M [13][17] to the video domain. It is built on three components: a multimodal video dataset constructed via pseudolabeling, a unified tokenization scheme that maps all modalities into discrete tokens, and a multimodal masked modeling objective that enables any-to-any prediction across modalities and time. We discuss each in the following.
No existing video dataset has ground-truth alignment across many modalities at the scale we need. We construct K600-MM, based on the Kinetics-600 dataset [1]: ~400K videos aligned across 12 modalities: depth, normals, optical flow, poses, feature maps, bounding boxes, captions, transcriptions, etc. These modalities are generated via pseudolabeling using task-specific specialist models. This gives us 34B tokens of aligned multimodal video data for pretraining. K600-MM is open-sourced and publicly available on Hugging Face.
Examples from K600-MM. Below we show data examples from the K600-MM pretraining dataset. Each example consists of an RGB video and corresponding aligned modalities.
Click a numbered tile to switch example sets.
Different modalities have fundamentally different structures (e.g., feature maps vs. RGB vs. bounding boxes) and are not directly compatible with a shared modeling architecture. We eliminate modality-specific design choices by mapping all modalities into discrete tokens, which can be jointly modeled with a single pretraining objective. Unless otherwise specified, all tokenized modalities in each data sample represent a 128×128 resolution video clip covering 17 frames with a 4-second duration, sampled at 4 FPS.
Modalities with a 3D spatiotemporal structure. For modalities with a 3D spatiotemporal structure, such as RGB, depth, normals, optical flow, and neural network feature maps, we train FSQ-VAE [11] video tokenizers that compress inputs both spatially and temporally to map them into discrete tokens.
Modalities that are represented sequentially. For modalities such as captions, transcriptions, and bounding boxes, we treat them as text tokens and encode them using WordPiece tokenization [12]. In addition, special tokens are added to the vocabulary to represent second delineators, frame-level boundaries, and modality-specific values, such as quantized bounding-box coordinates for the bounding-box modality, etc.
A2A-Video is an encoder-decoder transformer trained with a multimodal masked modeling objective [13]. During the pretraining, one subset is randomly selected from all modality tokens (across both space and time) of a training sample as input to the encoder and another random subset as target tokens which are predicted by the decoder. For modalities with a 3D spatiotemporal structure, tokens are sampled randomly for input and target, while for sequential modalities, we use span masking [14]. The model is trained with a cross-entropy loss to predict the output tokens. This pretraining objective allows cross-modal prediction capabilities at arbitrary time points, and can be scaled to any number of modalities.
Beyond the any-to-any predictions, the arbitrary modality-time traversal characteristic of A2A-Video enables higher-level inference instantiations that unlock new possibilities for temporal world modeling. We discuss these below.
By traversing across the modality axis, A2A-Video allows for chained generations, where the model is iteratively tasked to generate one modality at a time, feeding each prediction back to the model as additional context before predicting the next target modality.
We can use this property to decompose a fixed one-to-one task (e.g., text to video in our case) into a chain of simpler sub-tasks of predicting intermediate modalities. We demonstrate this in the animation below.
We show generations with a coarse-to-fine chain design: caption → transcription → SigLIP 2 → bounding boxes → depth → RGB. This chained inference yields higher-quality generations than direct generation.
Click a numbered tile to switch example sets.
In chained generation, semantic modalities such as SigLIP 2 and bounding boxes define the high-level layout and object-level motion trajectories. This is followed by geometric modalities in the chain, such as depth, which generate low-level structural details. This factorization leads to final RGB generations with coherent motion and fine-grained details, in contrast to videos generated directly from the captions. In our paper, we provide further ablations and analysis on the impact of the order of modalities and the modality types used in the chain.
Predicting how a scene will evolve given present observations (i.e., the current state of the world) is a core ability of any predictive world model and allows an agent to plan and act in it. Video generation models have received much attention for their potential as world models due to their ability to predict future frames.
However, it remains an open question whether world modeling should happen in raw sensory space, such as RGB video, in a latent space, as in JEPA-style world models, or in other abstractions, such as semantics.
Rather than settling for a single either-or answer, this work proposes joint multimodal modeling as a principled approach.
By treating one underlying physical reality as observable through multiple modalities at once (RGB pixels, semantics and learned latents alike), the model can traverse fluidly between these spaces and predict the future in any one of them.
Given initial past frames as input, A2A-Video can therefore predict the future across multiple abstract modalities simultaneously, as shown in the demonstrations below. Note that blue frames mark the observed input frames, while pink frames are A2A-Video's future predictions.
Click a numbered tile to switch examples.
Selecting the representation space. Based on the information of interest to be predicted, the representation space can be chosen accordingly. For instance, we can predict motion in optical flow space, rather than generating a future video from which motion must then be inferred. The latter is redundant, since it requires modeling details (e.g., background pixels) in RGB that are unnecessary for motion prediction.
Future prediction can also benefit from chaining. Just as chaining improves direct generation, the same idea carries over to predicting the future: modalities can be chained together to produce better future rollouts in the target modality space. We compare against direct (single-hop) prediction below, and quantify the comparison further in the energy-score analysis that follows.
For each modality pair below, the left is the direct prediction and the right is the chained prediction.
Click a numbered tile to switch example sets.
Setup. We experiment with predicting the future through different modalities by giving as input some observed RGB frames and using A2A-Video to generate the future frames of a chosen target modality. We test both direct prediction of the target modality (e.g., RGB → future Depth) and a single-hop chained prediction (e.g., RGB → future DINOv2 → future Depth).
Evaluation. There are many valid rollouts into the future given a single past sequence, which poses a problem for evaluation: each video in the test set only provides one realized future. We use Gneiting and Raftery's energy score [15], which quantifies the quality of a predictive distribution against a single realized outcome. Given sampled rollouts $\hat{x}^{(1)},\ldots,\hat{x}^{(K)}$ and the observed rollout $x$, the energy score is calculated as:
\[ \mathrm{ES} = \frac{1}{K}\sum_{i=1}^K d(\hat{x}^{(i)},x) - \frac{1}{2K^2}\sum_{i=1}^K\sum_{j=1}^K d(\hat{x}^{(i)},\hat{x}^{(j)}), \]
where $d$ is a distance metric.
Predicting (a) Depth and (b) Bounding boxes through different modalities. Lower is better. Each line represents the average quality, measured by energy score, of the predictions made through a particular modality across all test samples. The "Best chain" line represents the performance of oracle-selected best modality for each sample. The shaded area shows the gap between the per-sample best modality chain for prediction and the best single-modality prediction. The shaded area is significant, especially in the context of random, retrieval, and static baselines, suggesting that the best modality (representation space) for prediction is highly sample-dependent.
Results. Energy scores generally increase with the prediction horizon, as farther-future frames are harder to predict. The static baseline (repeating the last observed frame) does well only in the first few frames before degrading quickly, while the random and retrieval baselines stay roughly flat across the whole horizon. These baselines serve as calibration guidelines to gauge how much difference is meaningful. Ground-truth conditioned prediction (e.g. past GT depth-conditioned depth future prediction) is added as another baseline that performs well for the near future, since past ground-truth frames still tightly constrain the near future. Takeaway: best per-sample modality chain substantially outperforms any single fixed chain, even overtaking the ground-truth-conditioned prediction at later frames, with much slower degradation. This suggests that the optimal modality chain is highly sample-dependent, and that choosing the right abstraction (i.e., modality) to predict through can meaningfully improve prediction quality. We refer the reader to the paper for a more comprehensive breakdown.
Different modalities capture complementary spatio-temporal information of the world. A2A-Video leverages this to enable a diverse set of input controls for steering generation (in any output space), ranging from coarse (e.g., caption) to fine-grained (e.g., depth or bounding boxes) control. Through flexible modality-time conditioning,
this controllability connects directly to counterfactual generation and world modeling: starting from a present state (the observed frames) and given an action (the conditioning modality, i.e., modalities as actions), A2A-Video generates the resulting future state; different actions from the same starting state yield different plausible rollouts. Both the action and the representation space used for the current and future states can be chosen from among these abstract modality spaces.
We show generation results with various input combinations below. Choose a first-frame input, a conditioning modality, and an output space.
A2A-Video's multimodal conditioning enables factorized control of different aspects of videos. By decoupling geometry, motion and appearance through different modalities, we can edit each component independently. For instance, we can extract depth from the input video and use it alongside another conditioning modality to change attributes like color, while still preserving the geometry of the original video. Similarly, we can keep the object-level motion of the input video unchanged while altering its semantics, which can be achieved by using bounding boxes and changing their labels.
Changing Attributes while Fixing Depth Sequence
Changing Semantics While Keeping Motion Fixed
We can use A2A-Video for any-to-any retrieval in the video domain: retrieving samples in one modality given a query in any other modality. This is achieved by mapping all modalities of all samples in the database into the SigLIP 2 embedding space, and performing embedding-based similarity search in that same space. This allows retrieval across any pair of modalities, in either direction.
Choose the query modality and the retrieval modality to see retrieval results across different input-output modality spaces.
While A2A-Video demonstrates the benefits of modeling the world across time and modalities, we see several promising directions for future work:
We thank Jason Toskov for his discussion and feedback on the pretraining pipeline, Rishubh Singh for early discussions on motivating the problem formulation during the early stages of the project, and Kunal Singh for help preparing the any-to-any visualization figure. We also thank Ahmad Jarrar, Saqib Javed, and Muhammad Zakwan for valuable discussion and feedback on the project outside the lab.
We thank our sponsors DVPS, Google, and ELLIOT for their support. This work was also supported under project ID a08 as part of the Swiss AI Initiative, through a grant from the ETH Domain, and by funding from the Swiss State Secretariat for Education, Research and Innovation (SERI). Computational resources were also provided by the Swiss National Supercomputing Centre (CSCS) under the Alps infrastructure.