Interestana
Home/News/Generalist AI Releases GEN-1.5 Robot Foundation Model
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Generalist AI Releases GEN-1.5 Robot Foundation Model

Generalist AI has released GEN-1.5, a robot foundation model capable of learning new physical tasks from a single, brief demonstration. This model can process 3 to 12 seconds of sensorimotor data within a 30-second context window, enabling a robot to perform the demonstrated task without requiring gradient updates, fine-tuning, or task-specific programming. In evaluations across 10 diverse manipulation tasks, this "physical prompting" method achieved an average success rate of 59% (with a standard deviation of 10%) directly from the pretrained model. When subjected to ten gradient steps using five minutes of data per task, the success rate increased to 83% (±9%).

The company describes the core mechanism as "physical prompting," which reportedly emerged organically during over eight months of continuous pretraining on physical interaction data. This capability was not explicitly designed into the model through architectural changes, meta-learning loops, or auxiliary objectives. While the tasks demonstrated are acknowledged as simple and short-horizon, Generalist AI asserts that GEN-1.5 is the first known model to exhibit one-shot learning of physical skills at scale. However, the model is currently a research release and is not yet deployable for general use. There are no public weights, no API, and no pricing information available. Generalist AI utilizes GEN-1.5 on its own infrastructure and data engine, and any external access requires a direct partnership.

GEN-1.5 is characterized as a large multimodal model that integrates video, sensor, language, and proprioceptive inputs. It possesses a 30-second memory capacity and generates action trajectories at a frequency of 100 Hz. The extensive pretraining phase involved continuous learning on physical interaction data collected from various environments, including homes, warehouses, and factories. The "physical prompting" mechanism involves inserting a sensorimotor example, comprising sensor streams and the corresponding action trajectory, into the model's 30-second context window via a drag-and-drop interface. The remaining portion of the context window processes rolling observations, allowing the model to execute the task immediately without any gradient steps or fine-tuning. Generalist AI emphasizes that this in-context learning capability was an emergent property, not a result of specific architectural modifications designed to promote it.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next