LMMs-LabResearch journal

LMMs-Lab Blog

Research notes, engineering stories, and open-source updates—feeling and building multimodal intelligence.

Blog index

Latest entries

21 entries
  1. A Multimodal Hello World

    A compact smoke test for Markdown, equations, code highlighting, tables, and local images.

  2. LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

    The next generation of fully-open multimodal training — pushing the boundary of recipe transparency, native-resolution understanding, and end-to-end reproducibility.

  3. OneVision Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

    Codec Patchification processes only the 3.1%-25% of visual regions rich in signal entropy, improving video understanding while using substantially fewer patches.

  4. LLaVA-OneVision-1.5-RL: Unlocking Multimodal Reasoning via Lightweight Reinforcement Learning

    Applying reinforcement learning post-training to enhance reasoning capabilities in multimodal models with significant improvements on STEM, coding, and reasoning tasks.

  5. LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

    LongVT introduces a novel paradigm that natively interleaves multimodal tool-augmented Chain-of-Thought with on-demand clip inspection over hours-long videos, enabling large multimodal models to perform more effective and reliable long-video reasoning.

  6. OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe

    OpenMMReasoner introduces a systematic study on constructing high-quality SFT and RL datasets for multimodal reasoning, demonstrating that both source diversity and answer diversity are crucial for building reliable supervision signals.

  7. LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

    LLaVA-OneVision1.5 introduces a novel family of fully open-source Large Multimodal Models (LMMs) that achieves state-of-the-art performance with substantially lower cost through training on native resolution images.

  8. LLaVA-Critic-R1: Unified Critic and Policy Model Through Reinforcement Learning

    A family of generative critic VLM trained through GRPO using pairwise critic data, achieving SoTA policy performance at 7B scale while excelling at both evaluation and generation

  9. Improved MM-Search-R1: Reasoning and Action in Multimodal Search

    We improve MMSearch-R1 by integrating improved reasoning capabilities into the model

  10. Assessing Diffusion LM in the view of Multi-token Prediction

    Assessing masked diffusion language models viability for multi‑token prediction, diagnosing their efficiency and learning challenges, and surveying emerging solutions within the broader multi‑token prediction landscape.

  11. SAE Made Easy: Simplified Sparse Autoencoder Integration

    A framework that allows you to apply Sparse AutoEncoder on any models - inspired by PEFT design for seamless integration

  12. MMSearch-R1: Multimodal Search with Reinforcement Learning

    The first end-to-end RL-based solution designed to equip LMMs with the capability to perform search on demand in real-world internet environments

  13. High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

    MGPO enables LMMs to iteratively focus on key image regions through automatic grounding, achieving superior performance on high-resolution visual tasks without requiring grounding annotations

  14. Aero-1-Audio

    Aero-1-Audio is a 1.5B compact audio model capable of handling a range of audio tasks, including speech recognition, audio understanding, and audio instructions following.

  15. Video-MMMU: Evaluating Knowledge Acquisition from Educational Videos

    A benchmark that asks: If a model 'goes to class,' can it learn from lectures and apply knowledge to MMMU-style exam problems?

  16. Multimodal-SAE: Interpreting Features in Large Multimodal Models

    Large Multi-modal Models Can Interpret Features in Large Multi-modal Models - First demonstration of SAE feature interpretation in the multimodal domain

  17. LLaVA-Video

    LLaVA-Video: Video Instruction Tuning With Synthetic Data

  18. LLaVA-OneVision: Easy Visual Task Transfer

    The first single model that can simultaneously push the performance boundaries of open LMMs in three important computer vision scenarios: single-image, multi-image, and video

  19. LMMs-Eval

    Reality Check on the Evaluation of Large Multimodal Models

  20. LongVA: Long Context Transfer from Language to Vision

    Long Context Transfer from Language to Vision - An innovative solution towards long video LMM, leveraging long context capabilities of language models

  21. The Dream of the Digital Tide

    Created by GPT 4.5 to think about the future relationship between humans and machines.