想讓 AI 數字人不只是「會打字」,而是真的「邊說邊演」?香港中文大學(深圳)與 LIGHTSPEED 合作開源的 Ex-Omni-2D,正是針對這個目標設計。它是一個開源對話框架,接受多模態查詢、參考圖像與參考音訊後,先由語言模型輸出一份結構化的 Visual Thought Plan(VTP),再依此生成文字回應、個人化語音,並產出嘴形與動作同步的頭像影片。
測試範圍包括145段來源影片、196條經人工核實的音畫指令、2,688項細緻 checklist,以及28種編輯操作,涵蓋聯合音畫、語音、純影片和純音訊修改。它以 Multimodal Large Language Model (MLLM)-as-Judge 配合跨模態、影片及音訊自動指標,分開計算 Instruction Following、Fidelity Preserving、Editing Intent 和 Realism。
Prompt:
4 NFE PDD on Wan2.1 14B: A joyful child,
with a big smile and arms spread wide,
swings energetically on a rusty old swing set in a sunlit backyard. The swing set, with peeling paint and creaking chains,
contrasts against the vibrant green grass and blooming flowers surrounding it.
The child's laughter echoes as they swing higher and higher,
their feet barely touching the ground at the bottom of each arc.
The scene is captured from a low angle,
emphasizing the height of the swings,
with the sun casting a warm glow over everything.
Medium shot focusing on the child and the swing set.