# Multimodal AI
Multimodal AI is capable of understanding and processing multiple forms of data inputs and outputs such as text, images, and video.
11/17/2025 AI Express | What's New in AI: MiroThinker Open Source, Gemini and Grok Features Upgraded
MiroMind team released the open source bAgent model MiroThinker v1.0, proposing the concept of "Deep Interaction Scaling". Google for ...

Baidu MuseSteamer in-depth analysis: a new milestone in domestic AI video generation
Baidu's commercial R&D team launched MuseSteamer, a multimodal generative large model, which achieved the world's first place in the VBench graph-generated video review, and synchronized audio and video in Chinese...

Qwen-VLo: A major release in the field of multimodal AI from AliCloud
AliCloud recently released its latest multimodal AI model, Qwen-VLo, whose image generation and editing capabilities were highly rated by users and even surpassed GPT-4o. The model has a detailed...

OmniGen2: A breakthrough in next-generation multimodal AI
OmniGen2 is a multimodal generative model based on the Qwen-VL-2.5 architecture with 7 billion parameters, of which 3 billion are used for text processing and 4 billion for...

Google Gemini 2.5 Pro: a multimodal evolution from video to interactive apps
Google releases Gemini version 2.5 Pro, a major realization in the field of multimodal understanding and code generation. The model surpasses competitor Cl ...
