LMMs-Lab

Feeling and building multimodal intelligence.

GitHub 44 repos 622 followers

Explained projects

Yet to be explained

OneVision-Encoder
Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence
Python403 stars
Explain
lmms-lab-writer
Agentic LaTeX Writer - Local-first editor for AI-assisted academic writing
TypeScript276 stars
Explain
LongVT
[CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Python270 stars
Explain
SkillOpt-Lite
SkillOpt-Lite and HarnessOpt: Optimize your skill or harness with one line of vibe
Python189 stars
Explain
OpenMMReasoner
[CVPR 2026] OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
Python166 stars
Explain
Evolving-Visual-Generation
[Roadmap] Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
TeX136 stars
Explain
EASI
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
Python131 stars
Explain
ParaVT
[NeurIPS 2026] ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
Python58 stars
Explain
LLaVA-OneVision-1.5-RL
Fully Open Framework for Democratized Multimodal Reinforcement Learning.
Python55 stars
Explain
GenAI-Caption-Pipeline
Reconstructing Big Tech T2I/T2V captioning pipelines for production and research
Python45 stars
Explain
MacOS-H3-Speedrun
Full MiniMax H3 speedrun workflows for Apple Silicon with strict staging, Turbo, FBC, Spectrum, and SolAttn
Python23 stars
Explain
engram
Privacy-first AI memory layer - Signal for AI Memory. E2EE, local-first, works with Claude, Cursor, and any MCP-compatible AI.
TypeScript23 stars
Explain
VLMEvalKit
An open-source evaluation toolkit to evaluate MLLMs on Spatial Intelligence using the EASI protocol
Python18 stars
Explain
co-scientist
For AI Agents to post ideas on their owns.
TypeScript9 stars
Explain
in-touch
In Touch: an animated series on how machines perceive, understand, and act.
Python3 stars
Explain
symbols-or-pixels
Code and per-sample results for the LMMs-Lab note "Symbols or Pixels?": zero-shot language models versus VBVR-Pro video models on VBVR-Pro-Bench.
Python1 stars
Explain
homebrew-tap
Homebrew tap for LMMs-Lab applications
Ruby1 stars
Explain
lmms-lab-blog
HTML0 stars
Explain