Ph.D. Student (he/him) Co-advised with Aniket Bera
I am a PhD student at CoMMA Lab, advised by Dr Zachary Kingston and Dr Aniket Bera, where I focus on Generative learning for Grasp and Motion Optimization. Before coming to Purdue, I worked at Pointivo, Inc., developing and deploying ML pipelines for 3D asset inspection. I did my Bachelors from UET Lahore and my Masters from LUMS.
2026
arXiv
Fast Generative Grasping via Lie Group-Constrained MeanFlow
Grasp synthesis is a core task in robotic manipulation, for which the solution typically forms a multimodal distribution rather than a point estimate. Generative robotic grasping aims to learn this distribution with deep generative models such as diffusion and flow-based approaches. The iterative nature of such generative models makes them flexible and generalizable; however, multi-step sampling impedes the time-critical operation required in robotics. We devise an approach to fast generative grasping based on MeanFlow on the product Lie group SO(3) x R^3. The training objective couples a purely algebraic semigroup consistency condition with Riemannian Conditional Flow Matching on the product Lie group that anchors the average velocity to the data distribution. The resulting Lie Group-constrained MeanFlow formulation samples reliable grasps in at most 5 network evaluations, matching the grasp generation performance of state-of-the-art diffusion and flow-based models on the ACRONYM dataset at millisecond-scale inference latency (up to 39 times speed-up). We further demonstrate that the approach directly translates to real-world robotic grasping without additional training or domain adaptation, exhibiting robust grasp synthesis under observation noise.
@misc{bukhari2026meanflow,archiveprefix={arXiv},author={Bukhari, S. Talha and Wei, Yi and Ni, Ruiqi and Kingston, Zachary and Bera, Aniket},eprint={2608.26076},note={Under Review},primaryclass={cs.RO},title={Fast Generative Grasping via {L}ie Group-Constrained {MeanFlow}},year={2026}}
Grasp synthesis is a fundamental task in robotic manipulation which usually has multiple feasible solutions. Multimodal grasp synthesis seeks to generate diverse sets of stable grasps conditioned on object geometry, making the robust learning of geometric features crucial for success. To address this challenge, we propose a framework for learning multimodal grasp distributions that leverages variational shape inference to enhance robustness against shape noise and measurement sparsity. Our approach first trains a variational autoencoder for shape inference using implicit neural representations, and then uses these learned geometric features to guide a diffusion model for grasp synthesis on the SE(3) manifold. Additionally, we introduce a test-time grasp optimization technique that can be integrated as a plugin to further enhance grasping performance. Experimental results demonstrate that our shape inference for grasp synthesis formulation outperforms state-of-the-art multimodal grasp synthesis methods on the ACRONYM dataset by 6.3%, while demonstrating robustness to deterioration in point cloud density compared to other approaches. Furthermore, our trained model achieves zero-shot transfer to real-world manipulation of household objects, generating 34% more successful grasps than baselines despite measurement noise and point cloud calibration errors.
@article{bukhari2025graspdiff,author={Bukhari, S. Talha and Agrawal, Kaivalya and Kingston, Zachary and Bera, Aniket},doi={10.1109/LRA.2025.3645521},journal={IEEE Robotics and Automation Letters},title={Variational Shape Inference for Grasp Diffusion on SE(3)},year={2025}}