BeyondCSe grasps occluded past-event objects at 77% success
TL;DR
- BeyondCSe hits 77% grasp success on occluded targets referenced by past events, versus 55% for the strongest baseline in real-robot tests.
- On initially visible targets the system reaches 76% versus 40% baseline, using pretrained models without task-specific training.
- In heavy-occlusion scenes success rises from 75% to 95% while mean camera views drop from 3.35 to 2.20.
A robot that watches someone handle objects can later be asked to grab 'the one he was using' and succeed, even when that item is no longer in view. That is the claim of a new paper on arXiv introducing BeyondCSe, a zero-shot grasping system.
The authors frame the problem directly: "Such requests may specify a grasp target by the role it played in a past event rather than by its name or appearance. Moreover, the target may no longer be visible when the robot is asked to act."
On a real robot with a single wrist-mounted RGB-D camera, BeyondCSe reports grasp success rates of 76% and 77% for initially visible and occluded targets, versus 40% and 55% for the strongest baseline. On four heavy-occlusion scenes it raises success from 75% to 95% while cutting the mean number of camera views from 3.35 to 2.20, compared with an active-perception baseline that was handed the target's ground-truth 3D bounding box.
The system uses pretrained models without additional task-specific training. When the target is hidden, it combines an event prior recovered from the history with current scene geometry to pick where to look next. The abstract publishes no per-condition trial counts behind those headline rates.
Originally reported by paper
Read the original article →Original headline: Robot Grasps Past-Event-Referenced Objects at 77% Success Even When Occluded, vs. 55% Baseline