Baseten has added NVIDIA Nemotron 3.5 ASR Streaming models to its model library, offering English and 40-locale multilingual transcription through NVIDIA NIM. Baseten reports support for 100 concurrent real-time WebSocket streams on one H100, with finalisation latency below 140 milliseconds in its test. Production results will vary by workload.
Baseten Adds Hosted Inkling Small Model API Access
Baseten has added Inkling Small to its hosted Model APIs and dedicated-deployment service. The open-weight mixture-of-experts model has 276 billion total parameters, 12 billion active parameters, a one-million-token context window and native text, image and audio inputs.
Baseten Launches Distribution Platform for Closed Model Labs
Baseten has launched a distribution and monetisation platform for developers of closed-weight AI models. Baseten for Model Labs combines managed inference, Model Library distribution and commercial support, giving specialist labs another route to production customers without building a complete serving business themselves.
Baseten Cuts Wan 2.2 Video Inference Below Three Seconds
Baseten says its optimised Wan 2.2 runtime can generate a video clip in 2.75 seconds, 53.6 times faster than its baseline. A guarded public demo runs until 31 July, while the company acknowledges that four-step distillation can trade some visual quality for speed.
Baseten has made NVIDIA's Nemotron 3 Embed 8B and 1B models available for dedicated inference. The pair targets enterprise and code retrieval with different balances of accuracy, throughput and indexing cost, giving developers a managed deployment option while published performance figures remain vendor-reported.
Baseten Adds Day-One Hosting for Inkling Multimodal Model
Baseten has added day-one access to Thinking Machines Lab’s Inkling model through its managed Model APIs and Dedicated Inference service. The launch gives developers a hosted route to a very large open-weight, multimodal model, while leaving performance, cost and production suitability to be tested in real workloads.