Zeno-1: Collaborative Intelligence for Robots That Work Together
A foundation model for decentralised collaborative physical intelligence.
Contents
Robots are becoming increasingly capable on their own. Modern robot policies can perceive their surroundings, follow instructions, manipulate objects, navigate complex spaces, and execute increasingly sophisticated behaviours. But much of the physical world is organised around people working together, and many tasks become more practical when capability can be distributed across multiple agents. Rather than making every robot larger, more complex, or more expensive, teams of robots can work in parallel, share physical demands, and expand capacity incrementally as needs grow. Realising this potential requires physical intelligence that scales beyond the individual robot.
This makes collaboration a fundamental capability rather than an edge case. Carrying, assembling, handing over, manipulating deformable objects, and operating tools with others require more than individual capability. They require robots to stay in sync: to recognise what the agents around them are doing, adapt their own motion and timing, and remain coordinated as the interaction changes. This becomes harder as tasks extend over minutes rather than seconds. Small differences in timing accumulate. Contacts change. Objects move into new configurations. Other agents hesitate, recover, or behave differently from what was expected. In distributed systems, each robot must respond from its own observations rather than relying on a central controller with complete knowledge of the team. Decentralised control makes collaboration naturally extensible: new robots can join the team without redesigning a central controller or joint action space.
Today we are introducing Zeno-1, the first foundation model built specifically for decentralised collaborative physical intelligence among robots. At 3B parameters, Zeno-1 extends physical intelligence from robots that act alone to robots that can act together. Zeno-1 follows a staged progression: learn the world, learn itself, then learn to act with others. Large-scale video pretraining gives the model broad physical priors. Individual embodiment grounding turns those priors into robot capabilities. Finally, through Closed-loop Partner Interaction (CPI), independently acting robots become one another’s training partners. Its central idea is that Robots learn to collaborate by collaborating. During CPI, each robot acts from its own observations while interacting with a partner that is also observing and acting. Their motions change the shared world, their timing changes what the other robot should do next, and coordination must continually adapt as contacts and object configurations evolve. Targeted interventions focus additional learning on the moments when coordination breaks down. In this way, CPI transforms scalable individual robot experience into collaborative intelligence without relying on centrally choreographed multi-robot demonstrations.
In real-world experiments, the same Zeno-1 policy is deployed independently on every robot, without a central controller. The resulting teams sustain dexterous collaboration across shared-object manipulation, changing contacts, and variation in partner timing and behaviour. A single Zeno-1 policy can execute continuously for more than ten minutes, progressing through multiple subtasks and interacting with diverse objects without task resets or policy switching. Zeno-1 extends physical intelligence from what a robot can do alone to what robots can accomplish together. Collaboration is no longer a script imposed on the team. It is a capability learned by each robot.
From individual competence to collaborative intelligence
A robot that is highly capable on its own is not necessarily a capable collaborator. Two robots may each possess the dexterity required to grasp, lift, manipulate and transport an object, yet still fail together when their timing, motion or contact forces become incompatible.
The reason is fundamental: once multiple agents share a physical task, the appropriate action for any one robot depends on the evolving behaviour of the others. A motion that is correct in isolation may become incorrect a moment later as a partner moves, hesitates, changes contact or alters the configuration of a shared object. Collaborative competence is therefore inherently relational: a robot must learn not only what actions are useful, but when they are appropriate and how they should change in response to other agents.
Collaboration also offers a path to scalable physical capability. Rather than building increasingly large, specialised robots for heavier or more complex tasks, multiple general-purpose robots can work independently when possible and combine their capabilities when needed, allowing capacity to scale incrementally with better hardware utilisation and lower system cost.
Zeno-1 is designed around this transition from individual competence to collaborative intelligence. Its whole-body capabilities span dexterous manipulation, tool use and mobile interaction, while remaining responsive to the agents sharing the task.
A collaborative foundation model
Most visuomotor foundation models generate actions for a single robot from observations of the environment, language and the robot's own state. Zeno-1 extends this formulation to physical interaction between agents. Each robot generates whole-body actions from its local visual observations, proprioceptive state and a persistent representation of interaction history. This context provides information about how the task has evolved, how other agents are progressing and how the shared physical state is changing.
Zeno-1 is decentralised by design. The same model runs independently on each robot, without a central policy that jointly determines team actions. Each agent acts from its own visual observations and proprioceptive state, without access to another robot's internal state or planned actions. Through an optimised deployment stack, Zeno-1 runs closed-loop visuomotor inference locally at 30 Hz despite its 3B-parameter scale.
Instead, the shared physical world itself becomes the coordination interface. Partner motion, object response and task progress condition each robot's next action, allowing coordination to emerge through the evolving task rather than a fixed synchronised trajectory or explicit exchange of internal policy state.
This formulation extends naturally across multiple agents: adding another robot does not require constructing a larger central action space. Each robot continues to act from its own viewpoint while adapting to the partners and physical changes it can observe.
Learning collaborative physical intelligence at scale
The central training challenge is not simply teaching robots how to perform physical skills, but teaching those skills to remain compatible with the evolving behaviour of other agents. Zeno-1 addresses this through a four-stage training recipe: broad physical pretraining, individual embodiment grounding, closed-loop partner interaction and targeted correction of collaborative failures. The first two stages establish broad physical competence; the latter two transform that competence into collaborative behaviour.
Broad physical pretraining. Large-scale egocentric and exocentric video provide complementary views of physical interaction. Egocentric observations capture manipulation close to action, including grasp formation, contact, occlusion and tool use, while exocentric observations capture object trajectories, task progression and interactions between agents. Zeno-1 brings these views into a shared representation, providing a broad prior over physical behaviour before training on the robot itself.
Individual embodiment grounding. Real-robot demonstrations ground this experience in the geometry, dynamics and contact characteristics of an individual robot. At this stage, training focuses on acquiring high-quality whole-body manipulation skills rather than collaboration. We collect forty hours of demonstrations through teleoperation with force feedback and use them for supervised fine-tuning. This establishes the individual physical competence on which collaborative behaviour is later built.
Closed-loop partner interaction (CPI). Collaborative adaptation then takes place through direct interaction with deployed learned partners. Each robot encounters the timing, motion and execution variation produced by autonomous partners rather than observing only fixed jointly demonstrated trajectories. We use four hours of this interaction data for collaborative adaptation, allowing Zeno-1 to learn actions conditioned on a partner's realised behaviour. This distinction matters empirically. Synchronised demonstrations show what successful coordination looks like, but provide limited supervision for how behaviour should change once an interaction departs from that sequence.
Targeted correction. Deployment reveals where collaboration remains fragile. Human intervention is concentrated around these coordination-critical moments, providing corrective demonstrations that are folded back into training. This focuses additional supervision on failures exposed by autonomous interaction rather than behaviour the model already performs well.
Together, these stages separate learning how to perform physical skills from learning how those skills should adapt to a partner. Zeno-1 can therefore acquire collaborative behaviour without learning its underlying physical capabilities from jointly coordinated multi-robot demonstrations from the outset.

Collaboration in the physical world
These properties become most consequential when collaboration moves beyond task allocation and spatial coordination into physical interaction. Robots must coordinate how they move and make contact, preserve relevant state over time, adapt when partners deviate, and anticipate how their actions will affect what happens next.
Dexterity is evolving into a collaborative field
The challenge is most demanding when robots interact through the same physical objects. Handovers, shared grasps, deformable-object manipulation, tool use and collaborative assembly couple agents through contact: how one robot grasps, moves, holds or presents an object changes what actions remain possible for the others.
Zeno-1 extends collaboration into contact itself. A small pull can alter the tension of a sheet. A change in grasp can rotate a shared payload. Moving an object a few centimetres can open or remove a viable grasp for a partner. Robots therefore have to regulate their actions relative to one another continuously.
Zeno-1 learns this fine-grained partner contingency directly in its whole-body behaviour. It slows when another robot has not completed a grasp, holds position while a partner repositions, and changes its motion as a shared object deforms. Its manipulation therefore remains coupled to the evolving actions and contacts of its partners.
When robots manipulate together, dexterity becomes a property of the interaction, not just of the individual robots.
Sustaining collaboration over long durations
As tasks extend over time, the right action may depend not only on what a robot sees now, but on how it reached that state: which grasp was established, how a shared object was moved, whether a previous attempt succeeded, or which stage of the task has already been completed. Two moments can therefore look nearly identical while requiring different actions because their histories are different.
Zeno-1 addresses this with a persistent interaction memory system that maintains a compact representation of task-relevant visual and robot-state evidence throughout execution. Instead of retaining an ever-growing observation history, it preserves information that remains useful for future decisions, including prior contact, object manipulation and partner coordination.
This memory lets Zeno-1 both disambiguate visually similar states and sustain coherent behaviour over extended tasks. In our evaluations, a single Zeno-1 policy runs continuously for more than ten minutes, progressing through eight subtasks and interacting with diverse objects without task resets or policy switching.
Staying coordinated when partners deviate
Remembering how an interaction unfolded is only part of sustaining collaboration; robots must also remain coordinated when partner behaviour departs from what was expected. A partner may be delayed, become obstructed, lose contact, or approach a shared object differently than expected. In physically coupled tasks, these deviations change what actions are appropriate for every robot involved, and continuing along a nominal trajectory can turn a local mismatch into a team failure. We test this directly by perturbing partner timing during deployment.

This robustness emerges from closed-loop partner interaction, where coordination-critical deviations become part of the training distribution. Zeno-1 learns to slow, wait, yield, maintain contact, re-time its motion, and resume once the interaction becomes compatible again.
Recovery is therefore not a separate capability invoked after failure, but part of maintaining coordination as the interaction unfolds.
From adaptation to anticipation
Adapting to a partner can prevent local mismatches from becoming failures. But some failures are better avoided before they occur. Zeno-1 therefore uses predictive introspection to evaluate how an intended action is likely to affect what happens next. It reasons at two complementary levels: how an action may influence a partner's response, and how it may change the shared physical state.
At the level of partner interaction, an action can be locally correct but leave another robot poorly positioned for what comes next. A short wait or repositioning may instead enable a more compatible response. Zeno-1 therefore evaluates actions partly by the partner behaviour they are likely to enable. Across held-out collaborative decision points, Zeno-1 selected the action that led to better subsequent partner behaviour in 87% of cases, compared with 61% for an ablation considering only the action's immediate effect.
At the level of physical contact, an action-conditioned latent world model predicts how the shared state is likely to evolve under a planned action chunk. This allows likely misses, slips, poor grasps and object disturbances to be identified before execution. Across 200 live-robot rollouts, the world model achieves a failure-prediction AUC of 0.94 with a 0.5 s pre-contact warning window, compared with 0.81 AUC when predicting failure from the current observation alone.

Together, these predictions let Zeno-1 reason about both its partners and the shared physical state before committing to an action. It can avoid coordination failures rather than only recover from them.
Looking ahead
Physical intelligence has largely treated the individual robot as the fundamental unit of capability. We scale what one machine can perceive, understand and do. Zeno-1 asks a different question: what changes when the right action depends not only on the task, but on what other robots are doing?
With Zeno-1, we demonstrate that collaborative physical intelligence can be learned within a single decentralised policy. The same policy runs independently on each robot, coordinating from local observations without a central controller or access to another robot’s internal state. Through closed-loop partner interaction, Zeno-1 learns to adapt its motion, timing and contact to other agents, enabling dexterous shared-object manipulation and robust coordination when partners deviate. Persistent interaction memory allows this collaboration to remain coherent over extended tasks, while predictive introspection enables the model to anticipate how its actions may affect both its partners and the shared physical state. Together, these results show a path from individual physical competence to collaborative intelligence learned directly through interaction.
The next era of physical intelligence will be collaborative.
The goal is not simply to build robots that can do more alone, but teams whose capabilities expand when they act together: coordinating without perfect synchronisation, fixed roles or centralised control; adapting continuously to one another; and remaining effective as teams grow larger, more heterogeneous and more dynamic.
Zeno-1 is a first step toward that future. The next frontier of physical intelligence is not what a robot can do alone, but what intelligence can accomplish together.