SeedRealtime brings full-duplex audio and vision to live AI
ByteDance Seed has released SeedRealtime, a full-duplex model that continuously processes audio, video and text while deciding when to speak or act. The model is designed for fluid, interruption-aware interaction, with demonstrations spanning noisy group conversations, visual guidance and proactive assistance. Access and pricing details remain limited.
MiniMax Launches H3 for Multimodal 2K Video Generation
MiniMax has launched H3, an omni-modal generation model that accepts text, image, video and audio context and produces video with native stereo sound. The service supports clips up to 15 seconds at up to 2K resolution, with open model weights promised after release checks.
Black Forest Labs Opens FLUX 3 Multimodal Model Early Access
Black Forest Labs has opened early access to FLUX 3, a multimodal foundation model spanning image, video, audio and action prediction. Video generation is available first, while image, robotics and open-weight components remain on a staged release plan.
Baseten Adds Day-One Hosting for Inkling Multimodal Model
Baseten has added day-one access to Thinking Machines Lab’s Inkling model through its managed Model APIs and Dedicated Inference service. The launch gives developers a hosted route to a very large open-weight, multimodal model, while leaving performance, cost and production suitability to be tested in real workloads.