SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models
Imagine you have two AI assistants. One is really good at looking at pictures and talking about them. The other is really good at listening to sounds and talking about them. Now you want a single assistant that can do both, and maybe handle video and 3D scans too. What do you do?
The usual answer is to collect a huge pile of data where every example has an image, a sound and some text, all matching each other, and then train a new model, or fine-tune an existing one, for a long time on expensive GPUs. That is slow and costly, and often not even possible, because that kind of perfectly matched data barely exists.
In our new paper, SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models, we ask a different question: what if we simply merge the models we already have?
The Problem: Multimodal Models Are Expensive to Build
Multimodal large language models (MLLMs) are AI systems that can take in more than plain text. Some read images, some listen to audio, some watch videos. They are usually built by connecting special “encoders” (which turn an image or a sound into numbers) to a large language model that does the reasoning and the talking.
This works well for pairs of modalities, like vision and language, or audio and language, and plenty of strong public models already exist for these. But the moment you want one model that handles vision, audio and language together, things get hard:
You need large datasets where different modalities line up with each other, for example the audio and the video from the same recording plus a text description. These are expensive to collect and often impractical to synchronize.
Every time you want to add a new modality, you have to retrain or extend the model again, which takes serious compute.
Even clever shortcuts, like mapping every modality into an image-based space first, still depend on high-quality image-plus-other-modality pairs.
If you have read my earlier posts on missing modalities , you know I care about making multimodal systems practical outside of clean benchmark settings. SSAM comes from the same mindset: resources are limited in the real world, so let’s make better use of what already exists.
The Idea: Merge Instead of Retrain
Many pretrained vision-language, audio-language, video-language and point cloud-language models are publicly available. So we asked whether we could merge them into a single model that handles all the modalities, without training on any paired multimodal data.
Model merging is not a new idea. If you fine-tune a model for different tasks, the changes it went through can be written down as “task vectors”, and you can often combine them with simple arithmetic to get a model that does both tasks, with no training data at all.
The catch is that most of this work merges models of the same kind, such as several vision models or several vision-language models. Merging models built for different input modalities is messier. Their internal representations differ, and when you add their parameter updates together they can interfere with each other and wash out the very skills you wanted to keep.
Our Solution: Singular Subspace Alignment and Merging (SSAM)
SSAM is a training-free framework. There is no fine-tuning and no multimodal dataset involved. Here is the intuition, without the heavy math.
Keep the modality-specific parts separate. Each specialist model has parts that only make sense for its own modality, such as the audio-specific or vision-specific components. SSAM leaves those alone. There is no reason to blend the way you process sound with the way you process pictures.
Focus on the language part. The language model is the shared piece. Every specialist changed it a little during its own training. We measure that change as the difference between the specialist’s language weights and the original base model’s weights, and call it a language vector. These are what we need to combine.
Find what the models agree on. Earlier research showed that these updates are inherently low-rank. In plain words, most of what changed lives in a small number of important directions, and different specialists’ directions partly overlap. SSAM builds covariance-style matrices from all the language vectors and uses singular value decomposition (SVD) to pull out a shared low-rank “consensus” subspace, a small set of directions capturing the update patterns the models have in common.
Project, then merge. Each language vector is projected onto that shared subspace before merging. This keeps the dominant, consistent directions and filters out noisy or conflicting ones. The aligned updates are then merged into one.
A way to picture it: imagine several people editing the same document, each in their own style. If you paste all the edits together, you get a mess. SSAM first works out what kind of edits they mostly agree on, keeps those, and trims the contradictory bits.
What We Merged
We used four specialists, all built on Vicuna-7B as the language model:
- Vision-language: CLIP-ViT-L
- Audio-language: BEATs
- Video-language: LanguageBind
- Point cloud-language: PointLLM
Merging them gives one model that can handle any combination of image, audio, video and 3D point cloud inputs.
What We Found
We tested SSAM on four multimodal benchmarks and compared it with two groups of methods: jointly trained models (ImageBind-LLM, X-InstructBLIP, Proj-Only and OneLLM) and other training-free merging methods (Task Arithmetic, TSV, WUDI, OptMerge, NaiveMC and DAMC).
MUSIC-AVQA (audio-visual questions about music performances): SSAM 54.97%, best training-free baseline (DAMC) 52.60%
AVQA (audio-visual questions about real-world scenes): SSAM 81.29%, DAMC 80.29%
MCUB-3 (three-modality combinations): SSAM 60.35%, DAMC 59.80%
MCUB-4 (all four modalities together): SSAM 62.27%, DAMC 60.08%
The main takeaways:
State-of-the-art on all four benchmarks. SSAM beat every existing training-free merging method we compared with.
It also beat jointly trained models. Compared with the Proj-Only jointly trained baseline, SSAM improved accuracy by roughly 8 to 19 percentage points, without using any multimodal training data. That surprised us a little.
Specialist skills stay intact. On three domain-specific benchmarks, MMLU (general knowledge), OCRBench (reading text in images) and MMAU (audio understanding), the merged model matched or beat the individual specialists.
The subspace step is what makes the difference. In ablation studies, projecting onto the shared subspace clearly beat plain averaging or adding of the updates. A subspace rank of 128 worked best. Going higher let conflicting updates back in.
Qualitative results look sensible. The merged model can pick up shared semantic cues across image, audio, video and point cloud inputs and reason about them together.
Why This Matters
Lower cost. No massive training run and no paired data collection, which makes multimodal AI more accessible to teams without giant compute budgets.
Easy to extend. If a good specialist for a new modality shows up tomorrow, you can merge it in rather than starting over.
Reuse of open models. The community has already published many strong specialist models. Merging turns them into building blocks.
In short, aligning models in parameter space can be a practical and resource-efficient alternative to conventional joint multimodal training.
Limitations and What’s Next
SSAM is not magic, and we want to be upfront about its limits:
Our experiments used models of around 7 billion parameters. Larger models still need to be tested.
We only merged models that share the same architecture. Merging heterogeneous architectures is an open problem.
There are few benchmarks made specifically for evaluating MLLM merging, so we hope the community builds more.
We did not study the safety or alignment of merged models. How merging affects factual consistency, ethical behavior and robustness is important and worth investigating.
Next, we would like to try additional modalities such as depth, thermal and medical imaging, and to work toward merging across different architectures.
If you are interested in the related problem of what happens when a modality disappears at test time, check out our earlier work on Masked Modality Projection and parameter-efficient adaptation for missing modalities . And for another way of building multimodal language models without paired data, see Modality Self-Play .
More Details About the Paper
- Status: Under review
- Date First Available on arXiv: 23 March 2026
- Paper Link: arXiv
- Authors: Md Kaykobad Reza , Ameya Patil, Edward Ayrapetian, and M. Salman Asif
Related Posts
Robust Multimodal Learning via Cross-Modal Proxy Tokens
Imagine an AI designed to understand the world through multiple senses—like sight and hearing. It can identify a cat by both its picture (vision) and its “meow” (audio).
Read moreMMP: Towards Robust Multi-Modal Learning with Masked Modality Projection
In real-world applications, input modalities might be missing due to factors like sensor malfunctions or data constraints. Our recent paper addresses this challenge with a method called …
Read moreU2A: Unified Unimodal Adaptation for Robust and Efficient Multimodal Learning
Imagine you are using an AI system that analyzes both images and text to classify food items. It works great—until suddenly, the text data is missing.
Read more


