Digital Human 3

A real-time, video-first AI companion with persistent memory and multimodal expressions.
Beni AI is a multimodal AI companion platform designed for real-time, video-first interactions. Unlike text-only AI, Beni responds with voice, motion, and expressions through video calls. The system features persistent memory that allows the companion to remember past interactions and adapt over time. It is built to serve as a 'presence-native' companion that can be expanded into a creator engine, allowing users to bring any imagined IP to life and scale it into short-form content with the help of action plugins and perception awareness.

Open-source AI generating lip-synced talking videos from a single photo and audio/text.
daVinci-MagiHuman is an advanced, open-source 15B-parameter AI model developed by Sand.ai and GAIR Lab at Shanghai Jiao Tong University. It is designed to generate high-quality, lip-synced talking videos from a single portrait image and a script or audio file. Unlike traditional methods that combine separate text-to-speech and video pipelines, daVinci-MagiHuman utilizes a unified single-stream Transformer to jointly denoise video and audio tokens simultaneously. Released under the Apache 2.0 license, it allows users to inspect weights, run inference locally, and use the technology for commercial purposes. It is optimized for speed, capable of generating short clips in just seconds on professional-grade hardware like the NVIDIA H100.

Real-time interactive AI avatar platform with streaming video capabilities.
Vidu S1 is a real-time interactive AI avatar platform powered by a streaming video generation model. It allows enterprise product teams to build and evaluate digital human interactions using custom personas, voice control, and bidirectional perception. The system is designed to work efficiently across WebRTC media transport and WebSocket control channels, offering end-to-end low latency for live conversations.