MSP: Modality Self-Play for Training Multimodal LLMs Without Paired Cross-Modal Data
Think about how you understand a video of a dog barking in a park. Your eyes see the dog, your ears hear the bark, and your brain effortlessly puts the two together. Now imagine teaching an AI to do the same. The standard recipe says you need thousands of hours of videos where the sound and the picture are carefully matched, plus text labels on top.
That recipe works, but it is expensive, and it does not scale well. In our new paper, MSP: Modality Self-Play for Training Multimodal LLMs Without Paired Cross-Modal Data, we explore a different path: what if the model could imagine the missing modality in words, and learn from that?
The Bottleneck: Paired Data
Large multimodal models have improved a lot in the last few years, mostly by getting bigger and by training on more paired data. But there is a hidden wall here. Aligned multimodal datasets are:
Expensive to collect. Someone has to gather, sync and often annotate them.
Hard to scale across combinations. Vision plus language is one thing. Vision plus audio plus language plus something else? The number of combinations you need data for grows quickly with each new modality.
Often domain-specific. A dataset that works for one setting may not transfer to another.
Meanwhile, plenty of two-modality data is freely available: image-caption pairs, audio-caption pairs, and so on. That made us wonder:
Can a model learn to reason over several modalities together without ever seeing real paired non-text data?
Our Idea: Modality Self-Play (MSP)
The core idea is simple to say. During training, the model only ever sees one real modality at a time, called the anchor. For the modality it does not see, we give it a text description instead, and we call this a proxy.
For example:
The model gets a real image of a street scene (the anchor).
Instead of the real audio from that scene, it gets a short written description of what the audio would probably sound like, like “cars passing and people talking”.
That proxy description is converted into the same kind of representation the model would use for real audio.
So the model learns to work with an image and something that stands in for audio, in the same slot where real audio would go. The “self-play” part comes from who writes the proxy: it is the model itself. A first-stage model looks at the anchor and generates a description of the unseen modality. So the model is effectively playing both sides, one part inventing the missing piece, the other part learning to use it.
The key hypothesis is that proxy text, combined with strong pretrained contrastive encoders (the CLIP and CLAP families), carries enough meaning for multimodal composition to emerge on its own.
How It Works, Step by Step
MSP has two training phases.
The building blocks
Frozen encoders turn inputs into compact feature vectors: LAION-CLAP for audio and MetaCLIP ViT-B/16 for images.
Small trainable projector networks map those features into the language model’s space.
The language model itself is Qwen3-8B, adapted with lightweight LoRA adapters so we do not have to retrain the whole thing.
The two training phases
Phase 1: Learn each modality separately. We train the audio and vision projectors independently on ordinary caption datasets (audio with audio captions, images with image captions). At the end of this phase we have a model that can describe an image or a sound. That same model will later act as the proxy writer.
Phase 2: Self-play composition. For each training example we pick one real anchor modality and have the Phase 1 model write a short description of the other modality. These descriptions are generated once and cached, so training stays cheap.
There is a wrinkle here. Text embeddings and audio or image embeddings from contrastive models live in related but slightly shifted regions of the same space. If we plugged a text embedding straight into a projector trained on audio features, it would not fit well. So we learn a small text-to-modality bridge, a linear mapping fitted with ridge regression on just 1,000 caption-and-modality pairs. It is trained once and then frozen.
The real anchor and the bridged proxy are then inserted into the prompt at the reserved audio and vision slots, and the model is trained with the standard next-token prediction loss. To the language model, real and proxy inputs look the same shape and are inserted the same way.
We also keep a strict rule, the one-real-anchor rule: a given sample is only ever used with one of its real modalities, never both. This makes sure the model truly never trains on real paired audio-vision data.
At inference time. We do not generate any proxies. Real audio and real vision go through the same encoder-projector pathway and the model reasons over them together. In our analysis, even though the language model can tell proxy tokens from real ones internally, it converges to the same output tokens, which suggests it treats proxies as good substitutes for the real thing.
What We Found
Zero-shot modality composition works. Models trained only with a real anchor and proxy descriptions could integrate audio and vision at test time, even though they never saw them jointly during training.
Competitive or state-of-the-art results. We evaluated on multimodal benchmarks including AVQA, MUSIC-AVQA and OmniBench, and MSP matched or beat existing approaches without any real paired non-text supervision.
The bridge matters. Ablation studies show the learned bridges improve the alignment between proxy embeddings and real modality representations, which leads to better reasoning.
Proxy descriptions are doing real work. The anchor-supporting proxy descriptions are what enable effective cross-modal composition.
An Honest Note on What MSP Does Not Do
It is worth being clear about what we are claiming. MSP removes the need for real paired non-text data, like audio-plus-video from the same recording. It still relies on ordinary modality-text pairs (image-caption, audio-caption), pretrained contrastive encoders, and the small bridge calibration step. We are not saying that cross-modal knowledge appears out of nowhere. We are saying you do not need the most expensive kind of data to get it.
Why This Matters
If it holds up as we scale it, this approach changes the economics of building multimodal systems. Instead of asking “where do we find aligned data for every combination of modalities?”, the question becomes “can we describe the missing pieces in words?”. Text is the most abundant and flexible modality we have, and using it as a semantic bridge could make it much easier to add new modalities to existing models.
This connects nicely with our other work on making multimodal AI practical. In Cross-Modal Proxy Tokens , we used learned stand-ins for missing modalities to keep models robust. In Masked Modality Projection , we projected available modalities to estimate the missing ones. MSP takes a similar spirit to a different setting, the training of large multimodal language models. If you want to see how we approach the merging side of the same problem, check out our SSAM paper on merging multimodal LLMs without training .
Limitations
MSP is only as good as its proxies. If the generated descriptions are noisy or only weakly match the anchor, performance can suffer. It also builds on contrastively pretrained encoders in the CLIP and CLAP family, so it inherits their coverage gaps and biases.
What’s Next
We see several directions worth exploring:
Better proxies. Stronger generative models, or retrieval-based methods, could produce more accurate and better grounded descriptions.
Closing the training-inference gap. During training the model sees proxies, but at test time it sees real inputs. Integrating proxy information better at inference, or learning alignment that does not need proxies, could help.
More modalities and harder tasks. We want to test MSP on a wider range of modalities and more complex reasoning, and to study how it behaves at larger scale and with many more training tokens.
Theory. A deeper understanding of when and why multimodal composition emerges without real paired non-text data.
A Note on Broader Impact
Reducing the need for paired data can help in data-limited or privacy-sensitive settings. At the same time, better cross-modal synthesis could make it easier to produce multimodal deepfakes or fabricated content. To reduce these risks, we encourage the use of ethically sourced training data, controlled access to models, and provenance tracking for any synthesized outputs.
More Details About the Paper
- Accepted By: NeurIPS 2026
- Paper Link: NeurIPS 2026
- Authors: Robert Moseley, Md Kaykobad Reza , Ameya Patil, Edward Ayrapetian, and M. Salman Asif
Related Posts
SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models
Imagine you have two AI assistants. One is really good at looking at pictures and talking about them. The other is really good at listening to sounds and talking about them.
Read moreDualSwinFusionSeg: Multimodal Martian Landslide Segmentation via Dual Swin Transformer with Multi-Scale Fusion and UNet++
Mars has landslides. Just like on Earth, slopes collapse, dust and rock slide downhill, and the scars stay visible for a long time.
Read moreRobust Multimodal Learning via Cross-Modal Proxy Tokens
Imagine an AI designed to understand the world through multiple senses—like sight and hearing. It can identify a cat by both its picture (vision) and its “meow” (audio).
Read more


