Hello World, I Am Back!

Hello world, I am back!

If you have visited this blog before, you may have noticed that things went quiet for a while. It was not because nothing was happening. It was the opposite. The last year and a half went by so fast that I kept telling myself I would write an update “next week,” and then a few months would pass.

So this is that update. It covers the time since January 1, 2025: an internship, a tutorial, a bunch of papers, a PhD defense, a new job, and a lot of people who helped me along the way. I am writing it mostly for myself, so I do not forget how it felt. But if you are a student who is wondering what a PhD and a research career can look like from the inside, I hope some of it is useful to you too.

Let me start from the beginning of this stretch.

Phase 1: Interning at Amazon Lab126

The year started with a big opportunity: I joined Amazon Lab126 as a research intern in January 2025. I had spent a long time doing research in academia, and this was a great chance to work on problems that reach real customers.

The internship was later extended through the spring quarter, so I ended up spending six months there, from January to June 2025, working alongside some brilliant people who were building AI models to make everyday life a little easier. I found my projects both impactful and intellectually engaging, and I was thrilled to get the extension.

What did I learn? Less about a single technology and more about how to work:

  • Writing better code. Code that only I can read is not enough when other people depend on it.
  • Explaining my work clearly. A good idea that nobody understands does not go very far.
  • Managing my time. Deadlines feel different when a team is counting on you.
  • Building quick prototypes. Sometimes a rough version tells you more in two days than a perfect plan does in two weeks.
  • Solving real problems for customers. It changes the way you pick what to work on.

I owe a big thank you to my mentors, my manager and the whole team. They were always ready to help and to share what they knew, and I am still carrying a lot of what I learned from them. Several of them also became collaborators on the papers I mention below, which tells you something about how good the experience was.

Phase 2: Back to Riverside, and a Lot of Research

When I returned to UC Riverside, I started on several projects at once. I wanted to work in several different directions to broaden my research.

  • Cross-Modal Proxy Tokens. Soon after I returned, this work got accepted. The idea is to help a multimodal model keep working when one of its inputs, like audio or an image, goes missing at test time. Instead of building extra networks to generate the missing data, we let a small learned token approximate it by looking at the modality that is still there. The shorter version of the work was accepted at the Asilomar Conference on Signals, Systems, and Computers in 2025, and later a longer complete version was accepted at Transactions on Machine Learning Research (TMLR). I wrote about it in detail here: Robust Multimodal Learning via Cross-Modal Proxy Tokens .

  • Masked Modality Projection. Around the same time, our paper on MMP was accepted at the 2025 IEEE International Conference on Big Data. This one was led by one of my labmates who put in relentless effort and deserves most of the credit. You can read the story in MMP: Towards Robust Multi-Modal Learning with Masked Modality Projection .

  • If you want the longer story of this line of work on missing modalities, there are two earlier posts as well: one on parameter-efficient adaptation and one on U2A .

  • A tutorial at ICIP 2025. My advisor and I were invited to give a half-day tutorial, “Foundations and Recent Trends in Robust Multimodal Learning,” at the IEEE International Conference on Image Processing (ICIP) 2025 in Anchorage, Alaska, in September 2025. Preparing a tutorial means stepping back from your own papers and explaining the whole field: the basics, the challenges, the recent advances and the open questions. It was a good exercise in seeing the bigger picture of the area I work in.

Phase 3: Widening the Lens

Somewhere in this period I made a deliberate choice. I started working on several projects in different directions to increase the breadth of my research. A few of them were accepted at different venues:

  • Bengali video summarization. A collaboration led by a group of talented people, where we built the first end-to-end pipeline that summarizes Bengali academic videos. It was accepted to Findings of EACL 2026, which felt like a lovely way to start the new year. I wrote about it in Abstractive Summarization of Bengali Academic Videos Based on Audio Subtitles .

  • Landslides on Mars. Yes, really. With three great teammates, I took part in the first Mars Landslide Segmentation Challenge (MARS-LS) at the PBVS workshop at CVPR 2026 (we joined virtually), and our team finished 3rd place. Our paper, DualSwinFusionSeg, was accepted at the 22nd PBVS Workshop. Detecting landslides on Mars is tough because there is no vegetation to give contrast, and we had very little labeled data to learn from. Here is the full write-up: DualSwinFusionSeg: Multimodal Martian Landslide Segmentation .

  • Merging models instead of training them. Our SSAM paper asks whether we can combine independently trained vision, audio, video and point cloud models into one, without any paired multimodal training data. It is currently under review. The story is in SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models .

I am very grateful to my collaborators on these projects for their hard work and dedication. Working with such committed, friendly groups made this phase a lot of fun.

Phase 4: Little Milestones I Did Not Expect

Early in 2026 I reached 100 citations for the first time. Citations are not everything, but it was inspiring to see that other people had started to build on top of my research.

I got to present the Cross-Modal Proxy Tokens paper at the Information Theory and Applications (ITA) Workshop in San Diego, California, at the Bahia Resort. It was a fantastic week of talks and conversations with people who care about the theory underneath all this machine learning.

And then a nice surprise: MMSFormer , the first research project of my PhD, was recognized by the IEEE Signal Processing Society as one of the Top 25 Most Downloaded Articles (2024-2025). It was special because it was the project that started everything. Special thanks to my supervisor and my co-authors for that one.

Phase 5: Looking for the Next Step

Not all of this time was paper acceptances, so let me mention the quieter part too. While the research was going on, I was also applying and interviewing with multiple teams across multiple companies, trying to secure a full-time role in industry. If you have done this, you know it takes a lot of time and energy on top of your regular work. I am thankful to everyone who took the time to talk to me and give advice along the way.

Phase 6: Defending My PhD

In May 2026 I defended my PhD dissertation at the University of California, Riverside, after about 3 years and 9 months in the program. Even I was surprised by how quickly that came around!

The title was Robust Multimodal Learning With Heterogeneous and Missing Modalities . The dissertation looks at three questions that I kept coming back to:

  1. How can a model reason over any combination of input modalities?

  2. How can it stay reliable when some modalities are missing at inference time?

  3. How can we combine knowledge from independently trained multimodal models without expensive retraining or paired multimodal data?

MMSFormer, Cross-Modal Proxy Tokens and SSAM are the main answers to those three questions, and all of them have their own posts on this blog.

A PhD may have one name on the cover, but it is never one person’s achievement.

I am deeply thankful to my advisor for his mentorship, guidance and patience throughout this journey. I am also grateful to my committee members, collaborators, colleagues, mentors and teachers, and to the people I met during my Amazon internship. And most of all, to my family: my parents for a lifetime of encouragement, my wife for her unwavering support, and my daughter, whose joy and inspiration carried me through both the hard days and the good ones. I could not have done this without them.

Phase 7: A New Chapter at Amazon

In late June 2026, I joined Amazon as an Applied Scientist. It feels a bit like coming full circle, since this is where my industry journey started with the internship. I am looking forward to learning from talented colleagues, working on meaningful and challenging problems, and building things that actually help people.

And Then NeurIPS

Just a few days ago, on September 24, we heard that our paper “MSP: Modality Self-Play for Training Multimodal LLMs Without Paired Cross-Modal Data” was accepted to the main track of NeurIPS 2026 as a poster. The idea is that a multimodal language model can learn to combine modalities by using text descriptions as stand-ins for the modalities it never sees during training. Thank you to my co-authors for their hard work, feedback, encouragement and guidance. I am really looking forward to sharing it at the conference. You can read a plain-language version here: MSP: Modality Self-Play for Training Multimodal LLMs Without Paired Cross-Modal Data .

Oh, and one more small thing: I recently passed 200 citations on Google Scholar . Thank you to everyone who read, used or built on our work.

What I Would Tell My Younger Self (and You, If You Are a Student)

I am not an expert on careers, and I am still learning. But if a few things from these eighteen months are worth passing on, here are mine:

  • Say yes to collaborations. Almost every project above involved people I am grateful to. Research is a team sport.

  • Get experience outside your lab. My internship changed how I write code, communicate and choose problems. It made my research better.

  • Do not be afraid to change direction. Working on different topics widened my thinking and gave me new ideas for my main line of research.

  • Write about your work. Blog posts and talks force you to explain things simply. If you can explain it to a friend, you understand it.

  • Thank people, out loud and often. Nobody gets anywhere alone.

  • Be patient with the messy parts. Rejections, job interviews and slow weeks are all part of it.

If you want more practical tips on getting your work seen, I wrote about that too: Increase Research Visibility .

What’s Next

Time really did move in the blink of an eye. Looking back, most of the good things in this stretch happened quietly, in the background, long before anyone saw the result. There is a line I like about this:

Destruction has noise, but creation is quiet. Grow silently. ~ often attributed to Confucius

I do not feel like I am finished. There are still many exciting ideas I am working on, and there are so many open questions out there. Multimodal AI is nowhere near solved, and I would love to keep chipping away at it, hopefully in ways that make our lives, and the planet we live on, a little better, one beautiful moment at a time.

I will try to write here more often. If you are working on something similar, or you just want to say hello, I would love to hear from you:

Thank you for reading, and thank you to everyone who has been part of this journey.

Hello world, I am back. This is not the end. This is not even the beginning of the end. This is just the end of the beginning. Let’s keep going!

comments powered by Disqus

Related Posts

MSP: Modality Self-Play for Training Multimodal LLMs Without Paired Cross-Modal Data

Think about how you understand a video of a dog barking in a park. Your eyes see the dog, your ears hear the bark, and your brain effortlessly puts the two together.

Read more

DualSwinFusionSeg: Multimodal Martian Landslide Segmentation via Dual Swin Transformer with Multi-Scale Fusion and UNet++

Mars has landslides. Just like on Earth, slopes collapse, dust and rock slide downhill, and the scars stay visible for a long time.

Read more

SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models

Imagine you have two AI assistants. One is really good at looking at pictures and talking about them. The other is really good at listening to sounds and talking about them.

Read more