Yi Wei
(she/her)
2026
- arXivFaster Visuomotor Policy Learning on Action Manifolds via Riemannian MeanFlowUnder Review
Visuomotor policies learn a direct map from raw sensory observations to robot action sequences. Policies based on Diffusion and Flow Matching capture the multimodal distribution over action sequences in an end-to-end manner. This expressivity comes at the cost of multi-step numerical integration of the learned vector field for action generation, which can be expensive and time-consuming, impeding fast control rates required in robotics applications. Furthermore, robot action sequences are usually defined on a smooth, differentiable manifold, requiring that the learned policy respects the intrinsic geometry of the robot’s action space. Here, we present Riemannian MeanFlow Policy (RMFP), which learns the conditioned flow map of the probability path on the robot action manifold. Our formulation employs a flow map consistency objective grounded in the data by a Riemannian Conditional Flow Matching anchor. The flow map consistency condition is stable to train and constrains the learned model to finite-time transport, which yields on-manifold action sequence generation with as few as one network function evaluation. We present results on the spherical LASA and Push-T benchmarks, on the Tool Hang and Transport tasks of the Robomimic suite, and on the Franka Kitchen task with manifold-constrained action generation, and demonstrate that RMFP attains performance competitive with prior work at a lower sampling cost. We also employ RMFP on a real-world robotic manipulation task to demonstrate fast action generation under imperfect sensor measurements in the physical world.
@misc{bukhari2026rmfp, archiveprefix = {arXiv}, author = {Bukhari, S. Talha and Garrett, Austin and Wei, Yi and Ni, Ruiqi and Kingston, Zachary and Bera, Aniket}, eprint = {2609.30127}, note = {Under Review}, primaryclass = {cs.RO}, title = {Faster Visuomotor Policy Learning on Action Manifolds via {R}iemannian {MeanFlow}}, year = {2026} } - arXiv
Fast Generative Grasping via Lie Group-Constrained MeanFlowUnder ReviewGrasp synthesis is a core task in robotic manipulation, for which the solution typically forms a multimodal distribution rather than a point estimate. Generative robotic grasping aims to learn this distribution with deep generative models such as diffusion and flow-based approaches. The iterative nature of such generative models makes them flexible and generalizable; however, multi-step sampling impedes the time-critical operation required in robotics. We devise an approach to fast generative grasping based on MeanFlow on the product Lie group SO(3) x R^3. The training objective couples a purely algebraic semigroup consistency condition with Riemannian Conditional Flow Matching on the product Lie group that anchors the average velocity to the data distribution. The resulting Lie Group-constrained MeanFlow formulation samples reliable grasps in at most 5 network evaluations, matching the grasp generation performance of state-of-the-art diffusion and flow-based models on the ACRONYM dataset at millisecond-scale inference latency (up to 39 times speed-up). We further demonstrate that the approach directly translates to real-world robotic grasping without additional training or domain adaptation, exhibiting robust grasp synthesis under observation noise.
@misc{bukhari2026meanflow, archiveprefix = {arXiv}, author = {Bukhari, S. Talha and Wei, Yi and Ni, Ruiqi and Kingston, Zachary and Bera, Aniket}, eprint = {2608.26076}, note = {Under Review}, primaryclass = {cs.RO}, title = {Fast Generative Grasping via {L}ie Group-Constrained {MeanFlow}}, year = {2026} }