DualSwinFusionSeg: Multimodal Martian Landslide Segmentation via Dual Swin Transformer with Multi-Scale Fusion and UNet++
Mars has landslides. Just like on Earth, slopes collapse, dust and rock slide downhill, and the scars stay visible for a long time. Scientists study them to understand the planet’s geology and climate, and to find safe places for future missions. But the planet is huge, and nobody can look through every satellite image by hand.
That is where AI comes in. In our paper, DualSwinFusionSeg: Multimodal Martian Landslide Segmentation via Dual Swin Transformer with Multi-Scale Fusion and UNet++, we built a model that automatically outlines landslides on Martian terrain. We developed it for the PBVS 2026 MARS-LS Challenge, a competition on exactly this task.
The Task: Draw the Outline of Every Landslide
The technical name is semantic segmentation. Given an image, the model has to label every pixel, in this case as either “landslide” or “not landslide”. The output is a map that highlights exactly where the landslide is.
If you have read my post on MMSFormer , this will sound familiar. That paper was about fusing different sensor types for segmentation of everyday scenes and materials. Here we take the same fundamental idea, that different data sources tell you different things, all the way to another planet.
What Makes It Difficult?
The challenge dataset (called MMLSv2) is made of small 128×128 pixel tiles, and each tile has seven co-registered channels coming from several instruments orbiting Mars:
- Regular color imagery (RGB, from a Viking mosaic)
- CTX grayscale imagery
- THEMIS thermal inertia (roughly, how quickly the ground heats up and cools down)
- Terrain slope
- Elevation (DEM)
That sounds great, but it comes with real problems:
Wildly different resolutions. Some channels have a resolution of around 6 meters per pixel, while others are around 232 meters per pixel. Lining them up introduces artifacts.
Very little labeled data. We had only 531 labeled tiles in total (465 for training and 66 for validation). Modern deep networks usually want far more.
Huge variety. Landslides look different depending on the terrain, and some are old, faded and easy to confuse with craters or dust.
Class imbalance. Around 35% of the pixels are landslide, so the model cannot just guess “no landslide” everywhere and be mostly right, but the imbalance still needs handling.
Our Approach: Two Encoders, One Fusion, One Decoder
DualSwinFusionSeg has three main parts.
Two Swin Transformer encoders. Instead of pushing all seven channels through a single network, we use two separate Swin Transformer V2-Small encoders. One handles the RGB imagery. The other handles the auxiliary geophysical channels (thermal, slope, elevation and so on). Each encoder produces features at four scales, from fine detail to a broad view. The reasoning is that color images and terrain measurements are different kinds of data with different statistics, so it helps to let each have its own specialist. For the auxiliary encoder, we initialize the input layer by averaging pretrained weights across channels, so it starts from a sensible place despite the unusual input.
A simple, lightweight fusion. At each scale, we concatenate the features from both encoders and mix them with a 1×1 convolution. This is cheap and avoids the heavy cost of cross-attention. It also turned out to work better than a weighted sum in our tests.
A UNet++ decoder. To turn those features back into a sharp pixel-level map, we use UNet++, which has dense, nested skip connections that help preserve fine boundaries, important when the edge of a landslide is what you care about.
The whole model has about 113 million parameters.
Training details. We used a combination of class-weighted binary cross-entropy and Dice loss, which handles the class imbalance while directly optimizing for region overlap. Before training, we normalize each channel in two steps: clip extreme values (1st to 99th percentile), then standardize.
Results
- Development phase: mIoU of 0.867 and F1 of 0.905.
- Test phase: mIoU of 0.783 and F1 of 0.800.
Our model also outperformed MMSFormer, our earlier segmentation model, by 4.1% mIoU on this task. The gap between the development and test results is worth noting. We believe it mostly reflects a domain shift between the two sets rather than plain overfitting. Precision dropped more than recall on the test set.
What the Ablation Studies Taught Us
We ran a lot of experiments to understand what actually matters. A few of the most interesting findings:
Terrain geometry is the strongest signal. Elevation and slope were the most informative channels, and using all seven channels together gave the best mIoU (0.743 in that setup).
Two encoders beat one. Giving the auxiliary channels their own encoder added about 2.6% mIoU over a single-encoder baseline.
UNet++ won the decoder comparison. It beat FPN, UPerNet, SegFormer’s MLP decoder and DeepLabV3+ (0.824 versus 0.800 to 0.818 mIoU).
Concatenation beat weighted-sum fusion, by about 0.011 mIoU on average.
The loss function matters. Among eleven options, class-weighted BCE plus Dice was the best and the most stable.
Channel attention did not help much in this small-data regime. Every configuration landed within a tiny range.
Exponential moving average (EMA) of weights gave a +0.010 mIoU boost.
Test-time augmentation needs care. Rotating the input actually hurt, because slope features are direction-sensitive. A 5-fold ensemble with 4-view test-time augmentation reduced prediction variance.
The lesson we took away: with just a few hundred training tiles, simple, well-chosen design decisions beat fancy ones.
Where the Model Still Struggles
We looked at the failures too. The model has trouble with flat, dust-covered surfaces, very old landslides on nearly zero slope, and crater rims that look a bit like landslide scars. A 128×128 tile sometimes just does not contain enough surrounding context to decide.
What’s Next
Promising directions include larger context windows or progressive training to give the model more surroundings, domain adaptation across instruments, scaling up to bigger datasets (10,000+ tiles), and adapting large foundation models for planetary imagery.
Beyond Mars, the general recipe of using separate encoders for different kinds of sensors, fusing them simply, and decoding carefully could be useful anywhere that data comes from several instruments, from Earth observation to disaster monitoring.
More Details About the Paper
- Accepted By: 22nd IEEE/CVF Workshop on Perception Beyond the Visible Spectrum (PBVS) @ CVPR 2026
- Paper Link: arXiv
- Code and Pretrained Models: GitHub
- Authors: Shahriar Kabir, Abdullah Muhammed Amimul Ehsan, Istiak Ahmmed Rifti, and Md Kaykobad Reza
Related Posts
MMSFormer: Multimodal Transformer for Material and Semantic Segmentation
Leveraging information across diverse modalities is known to enhance performance on multimodal segmentation tasks. However, effectively fusing information from different modalities remains …
Read moreSSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models
Imagine you have two AI assistants. One is really good at looking at pictures and talking about them. The other is really good at listening to sounds and talking about them.
Read moreRobust Multimodal Learning via Cross-Modal Proxy Tokens
Imagine an AI designed to understand the world through multiple senses—like sight and hearing. It can identify a cat by both its picture (vision) and its “meow” (audio).
Read more


