By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Reka Releases Rho-1: A 19B Omni-Reasoning Model
Reka has unveiled a research preview of Rho-1, a 19-billion parameter omni-reasoning model developed from the ground up. This single neural network is designed to process and generate text, images, and video, perform reasoning across these modalities, and directly output robot actions. Reka positions Rho-1 as a potential replacement for current agentic pipelines that rely on passing tasks between multiple modality-specific models. The current approach in multimodal systems often involves a central model for planning, which then delegates specific tasks to specialized models for image processing, video analysis, or object detection. Each transfer of information between these models introduces latency and limits the scope of information each specialist model can access. Rho-1 aims to eliminate these handoffs by integrating text, vision, and robotic actions as tokens within a unified context window.
According to Reka's research, an unedited demonstration session illustrates the model's end-to-end capabilities. In this session, Rho-1 was shown to draw a lighthouse, apply bounding boxes to it, animate the drawing, modify the resulting video to depict a snowstorm, and then provide an explanation of the changes. This entire sequence was completed in five turns without requiring external tool calls or the involvement of a secondary model. The architecture of Rho-1 employs two distinct streams, each operating with a shared KV cache. Inputs and outputs are processed in one of two native formats: discrete tokens for text, symbolic reasoning, and high-level commands, and continuous tokens for image latents, video frames, robot actions, and proprioception. Each transformer block within the model contains two expert weight streams. The understanding stream is responsible for language and visual parsing, while the generation stream focuses on denoising latents to produce images and video. Both streams share attention mechanisms and operate on the same KV cache. When the model needs to produce visual output, the understanding stream emits a discrete handoff token, signaling the generation stream to render the output based on the accumulated state. The training process for Rho-1 combines next-token prediction for discrete sequences with flow matching for continuous outputs. This architectural design has practical implications, such as bounding boxes being outputted as coordinate tokens rather than requiring a separate detection model. Furthermore, the first frame of a video reuses the in-context image representation, avoiding the need for re-encoding.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.