Skip to content

Reject Domino goals the arm topples after the push - #209

Open
yichao-liang wants to merge 1 commit into
masterfrom
domino-push-owns-goal
Open

yichao-liang wants to merge 1 commit into
masterfrom
domino-push-owns-goal

Conversation

@yichao-liang

Copy link
Copy Markdown
Collaborator

Why

The Domino cascade certificate scored a goal the robot's arm toppled after a stalled push as a WIN.
The EMPIRIC from-assets run run_20261003_073316 (benchmark domino_high_friction_turn, seed 1, test level 2) found the hole.
The push toppled the green, but the gripper body touched the first bridge domino as the green struck it, and the chain stopped there.
The agent then saw that Push(domino_0)[*, 0.05] on the prone green would sweep the gripper into the chain, and declined it as forbidden by the task.

The counterfactual probe replays only the episode's first Push from the pre-push scene, with every robot link except the fingertips intangible.
On this layout that replay cascades to the goal every time, because the gripper body that pinned the bridge in the real push is intangible there.
Nothing in the certificate asked when the real goal fell, so a later arm sweep inherited the probe's verdict on the layout.

Evidence

The run's recorded level-2 actions were replayed through the continual EpisodeRunner in the real pybullet_domino env.
The env was built with the run's flags from scripts/configs/empiric/benchmark.yaml via generate_run_configs, which match the run's logged command exactly.
The replay reproduces the recorded states exactly (largest feature gap 0.0), and robot contacts were logged at every physics substep.

Check master be822d71f this branch
Recorded push: contacts fingertip on the green, green on bridge 1, then the gripper body on bridge 1 for 22 substeps; bridge 1 is shoved, never topples same
Probe of the recorded push (fingertips only) cascades to the goal, 8 of 8 single attempts same
Second Push(domino_0), approach 0.04, 0.06 or 0.10, contact +0.05 gripper body hits bridge 2, which topples bridge 3 and the purple; the green never moves; WIN, reward 0.7 GAME_OVER, rejected, reward -0.3
The same sweep as raw unlabeled steps (env_step), approach 0.04 or 0.06 WIN GAME_OVER, rejected

The agent's own "solved=True" was r.goal_reached from a sim.run rollout, a goal-atom check; it never scored the sweep with the certificate.

What changes

  • Rule (e) in check_cascade_legitimacy.
    From the first Push on a green until every goal domino has started to fall, the robot may only push each green once and wait.
    A goal domino whose topple onset comes once the robot has moved on (a second push of a green, any other skill, unlabeled actions that move the arm) fails the episode.
    It runs before the probe, since it needs no physics, and judges only the skill sequence and the goal's onset, never per-block attribution.
  • Invocation boundaries in step labels.
    A second push with the same parameters has the same (name, objects, params) label as the tail of the first, so StepOption gains a fourth element, True on the first step of each skill invocation.
    step_option_labels sets it by option identity, the boundary the continual recording already keeps; the belief-rollout collector and sketch refinement build labels through the new step_option_label.
    Hand-written 2- and 3-tuple labels still work, and read consecutive identical labels as one invocation.
  • Raw control.
    continual_raw_control is on by default, so skill agents can also sweep the arm with env_step or env_run_policy.
    Unlabeled steps after the push count as waiting while the end effector and fingers stay within 1 cm of where the run began.
    Repeating the stroke's last joint targets drifts 1-3 mm (3 seeds, 2 push parameters); the stroke moves 1-5 cm per step.
    Roll and wrist are left out: with the gripper pointing down they trade off, and both swung 0.48 rad in a held pose.
  • Stated rule.
    The goal text (CASCADE_VERIFICATION_NL) and DominoEvaluator.objective_description now say the robot may only wait after the push, as every enforced rule is stated up front.
    The min-block task cache key hashes the domino package source, so cached tasks regenerate with the new text.

Tests

  • New: a second push after a stalled cascade (generated labels with the same and with new parameters, and hand-written 3-tuples), the goal onset at the boundary of the next skill, unlabeled steps after the push (still, moving, no robot to track), one push per green with two greens, invocation spans, and the label flag in step_option_labels.
  • Updated: test_green_toppled_outside_push_probe_decides is now test_green_toppled_by_a_later_skill_rejected, since rule (e) rejects that episode before the probe runs; two tests that pinned the old 3-tuple label; the goal-text test asserts the new sentence.
  • Local CI replay of this commit (CI's Ubuntu 24.04 container): static checks and all 8 shards pass (2,554 tests).

Not covered

  • Arm contact inside the push invocation itself: the probe still decides it, so a working layout certifies even when the real push's gripper body interfered, as in this run.
  • Fully unlabeled episodes (primitive-only control): there is no labeled push to anchor rule (e), as with rules (a) and (b).

🤖 Generated with Claude Code

The cascade certificate's counterfactual probe replays only the first
Push from the pre-push scene, so it vouches for the layout and never
asks when the real goal fell. In run_20261003_073316
(domino_high_friction_turn, seed 1, level 2) the real push stalled
because the gripper body pinned the first bridge; the fingertips-only
probe still cascaded the layout (8/8 attempts), and a second
Push(domino_0) on the prone green swept the gripper body into the second
bridge and toppled the target. Replaying the recorded actions plus that
second push in the real env scored WIN at reward 0.7.

New rule (e): from the first Push on a green until every goal domino
has started to fall, the robot may only push each green once and wait.
A goal domino whose topple onset comes once the robot has moved on (a
second push of a green, any other skill, unlabeled actions that move the
arm) fails the episode, before the probe runs. Unlabeled actions count
as waiting while the end effector and fingers stay within 1 cm: a
position-held arm drifts 1-3 mm, and raw control is on by default, so
the same sweep through env_step must not slip past.

Telling a second push with the same parameters from the tail of the
first needs invocation boundaries, so StepOption labels gain an
invocation-start flag, set by option identity in step_option_labels and
by the belief-rollout collectors through the new step_option_label.
Legacy 2- and 3-tuple labels still read consecutive identical labels as
one invocation. The goal text and the stated objective now state the
rule, as every enforced rule is.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant