The reporting

1 articles

Topic: reward models

Research illustration showing the same recorded robot trajectory receiving different progress scores under two equivalent instructions about placing a radish in a pink bowl.

ANALYSIS Robotics

RoboRMBench Exposes Wording-Sensitive Reward Scores in Robot Learning

A new preprint tests vision-language reward models on 2,390 recorded real-robot trajectories and finds that equivalent instructions can produce contradictory judgments, revealing a measurable reliability gap before physical deployment.