Skip to content
HN On Hacker News ↗

HomeBody: A Humanoid That Explores, Remembers, and Acts on Its Own

▲ 39 points • 4 comments • by famouswaffles • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

31 %

AI likelihood · overall

Mixed
70% human-written 30% AI-generated
SEGMENTS · HUMAN 2 of 11
SEGMENTS · AI 0 of 11
WORD COUNT 1,260
PEAK AI % 62% · §7
Analyzed
Sep 26
backend: pangram/v3.3
Segments scanned
11 windows
avg 115 words each
Distribution
70 / 30%
human / AI fraction
Verdict
Mixed
Pangram v3.3

Article text · 1,260 words · 11 segments analyzed

Human AI-generated
§1 Mixed · 58%

01Long Horizon Humanoid Loco-ManipulationLong-horizon loco-manipulation takes a humanoid beyond what its ego view can show. To act across a room, it needs an internal spatial model that connects visible objects with remembered locations and helps resolve ambiguous requests. This context guides where to go, what to act on and how to sequence walking and manipulation. HomeBody combines that spatial memory with composable skills whose execution feedback lets the VLM revise its next decision when an action or transition fails.

§2 Human · 24%

More coming soon!1 / 2Tidy the kitchenClean up all of the coffee bags and put them in the middle, and throw away all of the milk and orange juice cartons that have gone bad.Cleaning the kitchen means deciding what to keep, what to throw away and how to move each object to the right place.

§3 Mixed · 56%

HomeBody gathers the coffee bags on the island and discards the specified cartons, coordinating repeated trips, grasps and placements across the room. Its view changes with every move, so memory and action feedback help it track what is done and what still needs attention.Retrieve the medicineI forgot my medicine, can you get it for me?

§4 Human · 22%

Also throw out the bad carton while you are at it.The medicine is initially out of view. HomeBody uses stored keyframes to locate the drawer, retrieves the medicine, hands it to the person, and then discards the carton. Across the room, it uses the right hand to grasp the drawer handle and open the drawer, then the left hand to discard the carton. Selecting between the two arms lets it access targets on both sides of the body.02How to DeployStep 1. ExploreFirst, we give the humanoid context about its role and let it explore an unseen environment. HomeBody collects 0.5× iPhone video, D435i camera observations, LiDAR scans with SLAM, joint poses and waypoints chosen by Astra. Exploration captures the room from the humanoid’s own viewpoint, grounding its spatial context in what it can see as it moves and interacts with the space. HomeBody retains this context when objects leave the ego view.Exploration instructionYou are a kitchen robot, please explore the space!HomeBody guides exploration of the kitchen, choosing useful viewpoints and saving observations for later tasks.Step 2. Real2SimHomeBody uses Astra as its Real2Sim agent to build a digital twin in Isaac Sim [2] from the humanoid’s own collected data. This grounds its observations in a spatial model of the world, helping the high-level VLM reason about locations beyond the ego view. See what accurate Real2Sim reconstruction requires and how the reconstructions compare.FROM EXPLORATION TO A DIGITAL TWINGrounding HomeBody’s Real2Sim agent in data from its own exploration, including SLAM geometry, ego views, joint states and waypoints, helps it build a geometrically, semantically and visually accurate digital twin for reasoning at deployment.Step 3. Give your humanoid an everyday taskGive the robot an instruction such as “tidy up the kitchen.” HomeBody uses its spatial context to choose actions and targets without an action-level script. Send the instruction below to replay an illustrative cleanup in the digital twin.Reconstructed kitchenIllustrative rolloutClick the room or Send to explore in 3DHomeBodyInstructionTidy up the kitchen. Put the coffee bags on the island and throw away the spoiled cartons.03Our ImplementationExpandable skill libraryWatch the humanoid navigate, grasp, place and open drawers. Explore skillsClose libraryNavigation, picking, placing and drawer opening form HomeBody’s action vocabulary. Skills share an interface for targets and execution results, so the VLM can compose them at deployment. Local retries and visual feedback correct execution errors.SKILL 01REAL ROBOT · 1.25×PickGrasp and lift the object selected in the ego image.SKILL 02REAL ROBOT · 1.25×PlaceMove a held object to a selected 3D release point.SKILL 03REAL ROBOT · 1.25× · retryOpen drawerVisually align with the handle, hook it, and walk backward to open the drawer.SKILL 04REAL ROBOT · 14× → 4×Pick from drawerReach into an open drawer and lift the selected object clear of its edge.SKILL 05REAL ROBOT · 1.25×NavigateFollow a planned route to a location in the Real2Sim map.Add your own skillConnect a learned policy, a classical algorithm or another controller through the shared interface.

§5 Mixed · 59%

HomeBody’s VLM can chain it with existing skills to carry out new tasks.Most skill previews are sped up to about 7–10 seconds. Drawer opening and grasping include recorded reattempts. The drawer-pick preview starts with the drawer open and keeps the retries continuous, slowing down for the final successful grasp.Our ArchitectureFAQWhat is required for accurate Real2Sim reconstruction?

§6 Mixed · 33%

Human-recorded videoSLAM referenceHomeBody (Ours) Human-recorded video provides rich visual detail for aligning a reconstruction with the room’s appearance. HomeBody also needs accurate geometry to navigate through the room, position the humanoid and reach objects. HomeBody uses the humanoid’s own exploration data as grounding for the Real2Sim agent, including camera observations, measured SLAM geometry, joint poses and selected waypoints.

§7 Mixed · 62%

This grounding supports geometrically accurate spatial reasoning alongside semantic understanding of the space.With video alone, the Real2Sim agent estimates the room’s dimensions from appearance. HomeBody also provides measured geometry from SLAM to constrain those dimensions. The middle and right panels use the same viewing angle and scale so their layouts can be compared directly.

§8 Mixed · 40%

Colors in the SLAM map distinguish points at different heights.How does HomeBody localize in a known environment? To plan how to complete a task, HomeBody needs to know both where relevant objects are and where the G1 is relative to them. This spatial context helps the planner choose where to move and how to sequence actions across the room. To ground these decisions in a shared coordinate frame, we localize the G1 with Super Odometry [1] and align its SLAM map with the reconstructed simulation using iterative closest point (ICP) registration. HomeBody stores ego camera observations in this shared frame, together with descriptive content. HomeBody can use this spatial information to return to the recorded position.How does HomeBody turn a VLM decision into physical action?

§9 Mixed · 61%

HomeBody uses spatial targets to connect task reasoning to physical execution. The VLM selects a skill and its target from the current ego view, map context, gripper state, recalled observations and the previous result. It passes this selection through a structured tool call, leaving the skill to plan and execute the motion. The VLM therefore does not need to know the skill’s low-level implementation.For picking, the call specifies an image point normalized to 0–1000 and which hand to use. The point prompts segmentation [3], while Fast-FoundationStereo [4] estimates depth from D435i stereo images. Camera calibration projects the masked geometry into 3D, where we predict the grasp analytically. To reach that pose, the arm planner builds a spline reference with minimum-jerk timing, solves inverse kinematics along the path and checks the swept motion for collision clearance.Other skills use targets suited to their actions. Navigation takes a 2D goal and facing point in map coordinates, measured in meters. A placing call specifies which hand to use, a 3D release target in the torso frame and a release distance. The skill moves the held object to the target and opens the hand. Drawer opening combines handle alignment, a hooking posture and backward walking into one skill, coordinating the transition from reaching to pulling.How does HomeBody correct mistakes and retry?

§10 Mixed · 42%

The target can shift in the camera view as the humanoid approaches. Segmentation identifies the object, and SAM 2.1 tracking with SAMURAI memory selection [5] follows it in subsequent frames.

§11 Mixed · 58%

We integrate visual servoing to use these tracking updates to correct alignment during the approach, without requiring a new VLM decision for each adjustment.If a grasp closes without contact, the pick skill can try another grasp candidate or adjust its stance and replan. These local retries are bounded.