In brief
- Rho-1 processes text, images, video, and robot control actions within the same model.
- The model handles all modalities in a shared context window rather than routing tasks to external models or making tool calls.
- Reka AI trained Rho-1 for about three months on 320 H100 GPUs.
- To address the shortage of robot training data, the company developed an inverse dynamics model that extracts control signals from internet videos.
Reka AI released a research preview of its multimodal model, Rho-1, on 5 October 2026. The company’s new model can process and generate text, images, video, and robot control actions within a single neural network with 19 billion parameters. According to The Decoder, rather than distributing tasks across different specialist models, Rho-1 handles all modalities in a shared context window.
What sets Rho-1 apart?
While many multimodal systems rely on separate models or tool calls for specific tasks, Rho-1 represents all modalities within the same model. According to The Decoder, the system does not route tasks to external models or use tool calls to complete them. This brings outputs related to text, images, video, and robot movements within the scope of a single model.
Reka AI says the model can generate continuous video and operate in real time. Its announced capabilities also include responding to new instructions while a stream is ongoing. Users can steer the process without having to restart video generation for each new request. The source article does not provide details on the use cases for these features or their performance limits.
How does it learn robot movements?
Rho-1 uses the same weights both to predict camera images and to guide robot movements. Because data for training robot control is limited, Reka AI has built an inverse dynamics model that extracts control signals from everyday videos found online.
The company trained Rho-1 for about three months using 320 H100 GPUs. While this gives an indication of the scale of its development, it does not specify the hardware needed to use the preview or which hardware the model will run on. The article also does not state access conditions, pricing, or a timeline for general availability.
What it means for architecture and visualization workflows
In architecture, interior design, and archviz, bringing text, image, and video capabilities together in one model could increase interest in tools that work across different content types. Real-time video generation and the ability to respond to new instructions mid-stream suggest the possibility of interactive work with moving visual content. However, Reka AI has not specifically listed architectural design or visualization among the model’s stated use cases, so there is not enough information to conclude that it directly benefits existing workflows.
For offices and visualization teams, the points to watch will be the model’s availability, licensing and usage terms, hardware requirements, and compatibility with existing tools. The source article provides no details on these topics. For now, Rho-1 should therefore be viewed not as a proven production tool for architecture, but as a research preview of an approach that brings text, images, video, and robot control together in one system.
Sources
1 sourceSource texts are not republished; short quotes are marked, everything else is our own summary and commentary.
What makes Rho-1 notable is that it brings real-time video and robot control into the same model as text and image generation. This approach could eventually pave the way for tools where visual content creation and interaction are less disconnected, but no direct use case for architecture or archviz has been provided yet.
For offices in Turkey, the access, licensing, pricing, and hardware details needed to make a decision have not been announced. It is therefore more sensible to monitor the scope of the preview and any updates on these conditions than to add the model to an existing production pipeline. In particular, it remains to be seen how the claims about video generation and real-time responses to instructions hold up in actual design tasks.
Frequently asked questions
What types of data does Reka AI Rho-1 process?
The model can process and generate text, images, video, and robot control actions within a single neural network.
When was Rho-1 released?
Reka AI released the research preview of Rho-1 on 5 October 2026.
Have Rho-1’s price and hardware requirements been announced?
The source article does not specify pricing, access conditions, or the hardware requirements for users.



