Transformer architecture 2

AI model creating realistic videos from text, images, or existing videos.
Sora is an AI model developed by OpenAI that can create realistic and imaginative scenes from text instructions. It is designed to understand and simulate the physical world in motion, generating videos up to a minute long while maintaining visual quality and adherence to the user’s prompt. Sora uses a diffusion model and a transformer architecture, similar to GPT models, allowing it to generate complex scenes with multiple characters, specific types of motion, and accurate details. It can also generate video from existing still images and extend or fill in missing frames of existing videos. Sora aims to be a foundation for models that can understand and simulate the real world, a step towards achieving AGI.

Deepseek's unified multimodal AI model for understanding and generating images and text.
Janus Pro AI is a unified multimodal understanding and generation model developed by Deepseek. It is an advanced version of Janus, incorporating an optimized training strategy, expanded training data, and scaling to a larger model size. Janus Pro AI excels in both multimodal understanding and text-to-image instruction-following capabilities, while also enhancing the stability of text-to-image generation. It supports bidirectional image understanding and generation via an autoregressive framework with a unified Transformer architecture.