日本語 · English
Purpose: In building the vision function components for evis (the MS-Human-700 musculoskeletal humanoid),
a sorting exercise to strictly hold the line of “don’t reinvent what OSS/ROS2 already covers.” Each perception module is
weighed against the standard stack actually used in ROS2 robotics to decide between
reinvents (a numpy re-implementation of existing OSS = little value in building) and genuine gap (absent from OSS = worth building ourselves).
The ROS2 ecosystem was confirmed by hands-on investigation (see Sources at the end).
Grasping (eating with chopsticks): stereo_rectify → disparity_sgm → depth_to_points + normals_from_depth
→ pcseg.remove_ground/euclidean_clusters → ppf.find_surface_pose(6-DoF) → grasp
Locomotion (hillco): depth → terrain.elevation_map → slope/step_edges/foothold_candidates
→ locomotion.support_polygon + com_support_margin
| module (ops) | Role | ROS2/OSS standard (in real use) | Verdict |
|---|---|---|---|
camera (21) |
projection/backproject/PnP/essential/rectify | image_geometry (PinholeCameraModel) + OpenCV (solvePnP/findEssentialMat/stereoRectify/Rodrigues) |
reinvents OpenCV |
stereo (11) |
census/SGM disparity, depth | image_pipeline/stereo_image_proc (SGBM), depth_image_proc, NVIDIA Isaac ROS (deep stereo) |
reinvents image_pipeline |
pcseg (17) |
RANSAC plane/sphere/cylinder, clustering, OBB, curvature | PCL (SACSegmentation/EuclideanClusterExtraction/MomentOfInertiaEstimation) via perception_pcl |
reinvents PCL |
pointcloud |
normals, voxel, outlier, FPFH | PCL (NormalEstimation/VoxelGrid/StatisticalOutlierRemoval/FPFHEstimation) |
reinvents PCL |
registration |
ICP, Kabsch, FPFH register | PCL (IterativeClosestPoint/SampleConsensusPrerejective) / Open3D |
reinvents PCL/Open3D |
ppf |
Drost PPF 6-DoF | OpenCV surface_matching (ppf_match_3d); frontier = deep (FoundationPose/GraspNet) |
reinvents OpenCV (frontier is deep) |
terrain (13) |
elevation/foothold/traversability/slope/step_edges | ANYbotics grid_map + elevation_mapping + leggedrobotics traversability_estimation (ANYmal = the legged-robot standard; holds elevation/foothold quality/traversability as layers) |
reinvents grid_map (core of legged robots) |
locomotion (5) |
support polygon, COM margin, gait phase | No standard ROS2 perception pkg. Scattered across legged-robot control (OCS2/TOWR/WBC) | partial gap (control-side) |
odometry (5) |
RGBD/PnP odometry, umeyama, trajectory | rtabmap_ros / ORB-SLAM3 / robot_localization |
reinvents rtabmap |
sceneflow (7) |
FoE, TTC, looming, scene flow | OpenCV optical flow (sparse/dense). scene-flow/TTC are research-leaning, thin ROS standards | partial gap (niche) |
features (5) |
Harris/FAST, descriptors, match | OpenCV (goodFeaturesToTrack/FAST/ORB/BFMatcher) |
reinvents OpenCV |
pose (3) |
silhouette posture descriptors | (simple silhouette-derived descriptors; thin direct OSS coverage) | partial gap (simple) |
occupancy |
grid, inflate, clearance | nav2 costmap_2d |
reinvents nav2 |
The complete loop where evis “uses” vision is perceive → plan → execute. The second half is where the important components lie, and here ROS2 has a thick standard.
Perception (stereo/PCL) → 6D grasp pose (GPD/AnyGrasp) → MoveIt2 MTC (grasp pose → IK → collision-free trajectory
→ move-to-pick/grasp/lift/place) → ros2_control (position/velocity/effort I/F) → robot
| Component | Role | ROS2/OSS standard (in real use) | Verdict |
|---|---|---|---|
| Motion planning/IK/collision avoidance | grasp pose → collision-free trajectory | MoveIt2 (OMPL/STOMP/Pilz, 150+ robots in production) + MoveIt Task Constructor (staged pick&place) | use OSS (not worth self-building) |
| grasp generation | cloud/RGB-D → 6-DoF grasp candidates + scores | GPD / AnyGrasp / SuctionNet (MoveIt integration), frontier = deep (GraspNet) | use OSS |
| Hardware abstraction/low-level control | position/velocity/effort I/F | ros2_control (the foundation of MoveIt2/Nav2; a humanoid’s ROS2 exposure is almost entirely through here) | use OSS |
| Navigation | mapping/pathing/obstacle avoidance | Nav2 (costmap_2d/BT) | use OSS (evis doesn’t need it for now) |
| grasp force/force-closure | antipodal grasp quality | GraspIt!/grasp (Ferrari-Canny) |
partial (the existing grasp op suffices) |
MoveIt2/ros2_control assume URDF position/torque joints + a standard gripper. evis is driven by MuJoCo’s 700 muscles (Hill type),
and handles chopsticks (a tool, not a gripper) with an articulated hand. → No OSS layer exists that realizes the joint trajectory MoveIt2 emits with the activation of 700 muscles.
This “joint plan → muscle activation” step = QP / static optimization / WBC (the QP+osqp you already have in reference_wbc_qp_control) is precisely the
evis-specific component that OSS/ROS2 cannot fill. The final step of vision (6D pose) → MoveIt2 (trajectory) → muscle realization (QP) is the real gap.
The tables so far leaned heavily on algorithms (PCL/OpenCV/grid_map/MoveIt2) and were missing the “see and confirm” layer. HDevelop is strong at displaying 2D images/BLOBs, but Physical AI vision requires 3D display (point clouds, depth, 6D pose axes, TF trees, grasp markers, elevation maps) as a must. The ROS2 standard here is RViz2. This is a visualization fidelity/feature reference so that Fullseye Studio (an HDevelop-style IDE) can be used to understand, test, and put things into practice; it is not an algorithm, so it is not a re-implementation target but rather serves as the requirements map for Studio exposure (F6).
| Subject | What to see | ROS2/OSS standard (in real use) | Handling in fullseye |
|---|---|---|---|
| point cloud | PointCloud2 color/intensity/normals | RViz2 PointCloud2 display / Open3D viewer | Reference for Studio’s 3D viewer requirements (F6). No re-implementation; integrate with an existing viewer or thin rendering |
| depth/image | depth colormap, camera image | RViz2 Image/DepthCloud, image_view |
Place Studio’s 2D panel (HDevelop equivalent) alongside 3D |
| 6D pose/grasp | pose axes, grasp posture markers | RViz2 Pose/PoseArray/InteractiveMarker, moveit_visual_tools |
Render ppf/grasp op outputs as axes in Studio (core of evis debugging) |
| coordinate frames | TF tree, link-relative poses | RViz2 TF display | Confirm the camera↔hand↔object pose chain |
| terrain | elevation/traversability layer | RViz2 + grid_map_rviz_plugin | Visualize footholds from the terrain op (hillco locomotion) |
★ Implication: Studio = a fusion of HDevelop (2D image-processing IDE) + RViz2 (3D perception visualization) is the right shape. evis vision debugging (is the 6D pose correct? / is the point-cloud segment reasonable? / is the foothold sitting on the terrain?) cannot be judged honestly without 3D visualization. Each op of the unified I/F should carry in its meta (F3) “how it draws in Studio” (2D image / point cloud / pose / grid_map layer), so that Studio can automatically select RViz2-equivalent rendering.
Fullseye = a comprehensive library that holds every image-processing/vision algorithm as a “skill,” ready to use instantly (a dedicated HALCON). HALCON-grade coverage is the goal. Fullseye Studio (HDevelop-style) = an IDE to understand, test, and use these functions in real work. → The ROS2 investigation in this doc is not “don’t build it,” but a fidelity/priority reference for building it correctly, comprehensively, and faithfully to real-world vocabulary.
reference_wbc_qp_control). The most important component absent from OSS/ROS2.sim.MuJoCo / sim.Gazebo / sim.IsaacSim behind the unified I/F with the same verbs
(.frames() / .depth() / .intrinsics() / .ground_truth()), so that vision ops can be composed regardless of the input source.
The difference in their roles =
sensor_msgs)
= the entry point where the unified I/F can be validated with real ROS2 wiring.sim.*.ground_truth(): true 6D pose, segmentation, contact) —
an evaluator (along the lines of CONSUMER_APPLICATIONS.md). Isaac Sim’s synthetic data + randomization is the proper source of evaluation data. App-specific.★ Unified interface principle (user-confirmed 2026-08-17): whether the internals are self-built numpy or an OSS wrapper, the caller can invoke it through fullseye’s identical I/F (facade op naming, signature conventions, Studio exposure, honest gate/meta). Even where OSS is used, don’t hit raw PCL/OpenCV directly; tuck it away as a thin adapter behind fullseye’s unified I/F (= the same as how HALCON offers diverse internal implementations under a single operator vocabulary). The substance of “instantly usable as a skill” is this consistent I/F.
To fit fullseye’s purpose (comprehensive, instantly usable, skill-ified, unified I/F, Studio practicality), run two lines in parallel:
To confirm: should tonight’s autonomous work go to (A) coverage of perception ops + Studio exposure, or (B) the evis muscle-drive bridge / sim vision? Do not fall back to general CS (algo-c).