← Research & News Ludi

Ludi₀.₁: An Agentic System for Socially Intelligent Robots

Published: August 2, 2026  ·  Contact: research@ludorobotics.ai

Wooseong Chung, William Cong, Jakub Dworakowski, Ethan Ewer, Try Wahyu Guntara, Yeonwoo Jeong, Tianchong Jiang, Chaewon Kim, Hyunseo Kim, Jinwoo Kim, Jinyeon Kim, Yea-Seul Kim, Jack Kunde, Kangwook Lee, Sangheon Lee, Robert Nowak, and Junha Roh.

Demonstration 1 — Human–Robot Interaction
Demonstration 2 — Fluid Interaction

Robotics has advanced rapidly in recent years. Action foundation models, including models based on pretrained large language models (LLMs) and world models, have demonstrated increasingly capable perception, navigation, and manipulation.

Much of the field has focused on expanding what robots can physically accomplish. At Ludo Robotics, we believe that physical capability alone is not enough. Robots also need social intelligence: the ability to let everyday people guide and collaborate with them naturally through speech, gesture, and shared understanding.

We believe now is the right time to address this problem directly.

Social Intelligence in Action

Consider a table with two soda cans, one Coca-Cola and one Pepsi. A person asks a robot:

“Pick up a soda can for me.”

An action model might treat this as a straightforward manipulation task, selecting whichever can is closest, easiest to reach, or assigned the highest action probability.

But the instruction is ambiguous. A socially capable person would likely ask:

“Which one would you like, Coca-Cola or Pepsi?”

The important behavior is not merely choosing the correct grasp trajectory. The robot must recognize that the user has not provided enough information, determine that the distinction may matter, and ask an appropriate follow-up question.

A more difficult case arises when the user changes their mind after the robot has already begun acting.

Suppose the robot remembers that the user usually prefers Coca-Cola. It begins reaching for the Coca-Cola can and says:

“Sure, I’ll bring you a Coca-Cola.”

While the robot is moving, the user changes their mind:

“Actually, I’ll try Pepsi today.”

The robot should interpret this correction immediately and adjust its behavior. Depending on its current state, it might redirect its reach toward the Pepsi can or safely set down the Coca-Cola before retrieving the Pepsi.

Despite the tremendous breakthroughs in robotics in recent years and the eye-catching demonstrations they have produced, robotic systems cannot yet perform this kind of interaction reliably across varied environments and tasks.

Doing so requires these capabilities to operate together seamlessly and in real time:

Ludo Robotics was founded to make this kind of fluid, natural human–robot collaboration possible.

Ludi₀.₁: An Agentic Approach

Ludi₀.₁ is our first system designed to integrate social reasoning, dialogue, memory, and physical action in a functioning robot.

We deliberately chose an agentic architecture. Rather than relying on a single end-to-end model, Ludi₀.₁ coordinates multiple specialized models and tools through a vision-language model (VLM) and a control harness.

At the center of the system is a 4-billion-parameter VLM responsible for:

The available tools include custom vision-language-action (VLA) models, indoor navigation systems, and speech-generation components.

All inference and system operation can run locally, without relying on external model APIs or cloud-based inference. In the demo, Ludi₀.₁ ran on an onsite GPU workstation connected to the robot.

The VLM operates within a harness responsible for several system-level functions:

We refer to the complete system, including the harness and its underlying models, as Ludi₀.₁.

What Ludi₀.₁ Can Do

Ludi₀.₁ currently supports:

Together, these capabilities allow the robot to handle interactions that cannot be reduced to isolated perception or control tasks.

For example, the robot can receive a spoken request, inspect its surroundings, decide whether clarification is needed, explain what it intends to do, begin an action, and revise that action when the user provides new information.

The goal is not simply to make robots more conversational. Speech must remain grounded in what the robot perceives, remembers, plans, and does, so communication and physical action unfold as one continuous interaction in the real world.

The Limits of Agentic Robotics

Ludi₀.₁ takes an agentic approach to robotic intelligence. A central agent maintains shared context and continuously decides whether the robot should speak, clarify, wait, stop, navigate, or manipulate, while specialized models carry out the underlying tasks. This architecture enables innovative capabilities that are difficult to achieve with conventional action models, including real-time clarification, interruption, memory-informed behavior, and coordination between conversation and physical action.

But the approach is fundamentally limited. Dialogue, perception, memory, navigation, and control remain distributed across separate components, and their coordination can introduce context loss, latency, brittle handoffs, and fragmented behavior. The system may appear integrated at the interaction layer, but it does not learn a unified representation of the user, the environment, and the robot’s evolving physical state.

Overcoming these limitations will require new foundation models designed specifically for robots and people. Such models could connect language, perception, memory, reasoning, and action more directly, allowing behavior to emerge from a shared understanding of the person, their intent, the physical environment, and the evolving interaction between them, rather than from coordination among loosely coupled modules. Ludi₀.₁ provides both a practical system for exploring advanced human–robot interaction today and a mechanism for collecting the continuous training traces needed to build this more deeply integrated model.

Toward Ludi 1.0: A Foundation Model for Robots and People

The interaction traces collected by Ludi₀.₁ provide the training data for Ludi 1.0, a foundation model for robots and people that more deeply integrates perception, dialogue, memory, reasoning, and control. Rather than treating conversation as an interface layered on top of robot behavior, Ludi 1.0 will ground language and physical action in a shared representation, allowing each to continuously inform the other.

A robot operating in this way must understand not only what it is doing, but why it is doing it, what the user currently expects, how the user’s intent is changing, and how its physical state constrains what it can do. It must also know when to explain, revise, pause, or stop its behavior in real time.

The goal is a robot that understands people, communicates naturally, and acts as a collaborative partner in one coherent, ongoing interaction. Ludi₀.₁ gives us a working environment in which to develop these capabilities, expose the limits of current systems, and identify the representations a foundation model for robots and people will require. Ludi 1.0 is a step toward a new class of machines: robots that do not simply execute commands, but understand, adapt, and collaborate with people in the open-ended complexity of the physical world.

We plan to release Ludi 1.0 later this year. Stay tuned.

CITATION

Ludi₀.₁: An Agentic System for Socially Intelligent Robots

Ludo Robotics, 2026.

contact@ludorobotics.ai copied to clipboard