Modern methods reconstruct or generate simulation-ready articulated objects, predicting not only their geometry but also how their parts are connected and allowed to move. Evaluating the geometry is straightforward, but evaluating the predicted articulation is not, because articulation specifies a motion rather than a shape, and there is no agreed distance between two motions. More specifically, existing protocols score joint type, axis direction, origin, and motion limits separately, although these parameters jointly describe a single physical motion, and the same motion can be written as different parameter values. As a result, a joint can score maximally wrong against an equivalent encoding of itself, and several component errors are ill-conditioned or undefined exactly where predictions become accurate. We propose ArticulateArena, a representation-invariant counterpart of Chamfer distance for articulation that compares the motions one-DOF joints induce rather than the parameters that encode them. It represents each joint by the unordered pair of its Lie-algebra endpoint twists, and we prove that the resulting quotient distance is a metric. It unifies fixed, revolute, prismatic, and helical joints, brings continuous joints into the same score through a compactification, and reads as the RMS motion of the moving part in meters when weighted by its mass distribution. A motion-aware tree edit distance lifts the metric to full kinematic trees, pricing structural errors such as spurious or missing joints in the same motion units as joint errors, and for a fixed inner product it remains a metric on trees up to relabeling. Alongside the metric we release ArticulateArena-20K, a new suite of 19,977 articulated objects with verified kinematics, and we re-evaluate published reconstruction methods on it under the new metric.
Articulated objects move through a set of canonical joint types, each of which constrains how one part moves relative to another. Our metric covers the five one-DOF types below (fixed, prismatic, revolute, continuous, and helical), shown with an example object, its kinematic tree (the highlighted node is the moving part), and the motion the joint allows.
From a URDF joint to our metric. (1) A joint is specified by type, axis, origin, and limits. (2) The twist ξ = (ω, v) absorbs type, axis, and origin at once, and the limits give the unordered endpoint pair z± = l±ξ. (3) The metric E compares two endpoint pairs under the better of the direct and the swapped assignment. (4) Radial compactification maps unbounded ranges into the unit ball, so continuous joints become boundary points and every joint type is scored on one scale.
The animations below illustrate different articulation errors against the ground-truth motion. More error cases will be added here.
Qualitative comparison over four benchmark cases rendered with the viewer's per-part palette, so every part keeps its own flat color. Columns are object instances and rows are the methods, starting from the ground truth, each cell animating that method's articulated output over its predicted joints. The label under each cell counts the movable joints the method predicts, and hovering a cell magnifies its clip. The quantitative evaluation aggregates all valid outputs in the evaluation set, not only these examples.