A Multimodal Hello World
A compact smoke test for Markdown, equations, code highlighting, tables, and local images.
Research notes, engineering stories, and open-source updates—feeling and building multimodal intelligence.
Blog index
A compact smoke test for Markdown, equations, code highlighting, tables, and local images.
The next generation of fully-open multimodal training — pushing the boundary of recipe transparency, native-resolution understanding, and end-to-end reproducibility.
Codec Patchification processes only the 3.1%-25% of visual regions rich in signal entropy, improving video understanding while using substantially fewer patches.
Applying reinforcement learning post-training to enhance reasoning capabilities in multimodal models with significant improvements on STEM, coding, and reasoning tasks.
LongVT introduces a novel paradigm that natively interleaves multimodal tool-augmented Chain-of-Thought with on-demand clip inspection over hours-long videos, enabling large multimodal models to perform more effective and reliable long-video reasoning.
OpenMMReasoner introduces a systematic study on constructing high-quality SFT and RL datasets for multimodal reasoning, demonstrating that both source diversity and answer diversity are crucial for building reliable supervision signals.
LLaVA-OneVision1.5 introduces a novel family of fully open-source Large Multimodal Models (LMMs) that achieves state-of-the-art performance with substantially lower cost through training on native resolution images.
A family of generative critic VLM trained through GRPO using pairwise critic data, achieving SoTA policy performance at 7B scale while excelling at both evaluation and generation
We improve MMSearch-R1 by integrating improved reasoning capabilities into the model
Assessing masked diffusion language models viability for multi‑token prediction, diagnosing their efficiency and learning challenges, and surveying emerging solutions within the broader multi‑token prediction landscape.
A framework that allows you to apply Sparse AutoEncoder on any models - inspired by PEFT design for seamless integration
The first end-to-end RL-based solution designed to equip LMMs with the capability to perform search on demand in real-world internet environments
MGPO enables LMMs to iteratively focus on key image regions through automatic grounding, achieving superior performance on high-resolution visual tasks without requiring grounding annotations
Aero-1-Audio is a 1.5B compact audio model capable of handling a range of audio tasks, including speech recognition, audio understanding, and audio instructions following.
A benchmark that asks: If a model 'goes to class,' can it learn from lectures and apply knowledge to MMMU-style exam problems?
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models - First demonstration of SAE feature interpretation in the multimodal domain
LLaVA-Video: Video Instruction Tuning With Synthetic Data
The first single model that can simultaneously push the performance boundaries of open LMMs in three important computer vision scenarios: single-image, multi-image, and video
Reality Check on the Evaluation of Large Multimodal Models
Long Context Transfer from Language to Vision - An innovative solution towards long video LMM, leveraging long context capabilities of language models
Created by GPT 4.5 to think about the future relationship between humans and machines.