I am a Master's student in Electrical and Computer Engineering (Computer Vision) at the University of Michigan, Ann Arbor. I build robotics software, perception, and machine learning systems for real-world robots — visual servoing and pose estimation for industrial manipulation, imitation learning and vision-language-action policies for robot arms, SLAM with language-guided navigation, and world models for long-horizon planning.
Most recently I was a Robotics Software Intern at Mowito, where I built a visual-servoing auto-fastening stack that derives target poses from a single demonstration — enabling CAD-free robotic bolt fastening on parts moving on a conveyor — and deployed it on FANUC and UR10e arms. As a Graduate Research Assistant at UMich I work on JEPA-based world models for planning, and previously at Carnegie Mellon University I trained multimodal audio-language models for deepfake speech detection (under review at Interspeech 2026). My B.Tech is in Electrical Engineering from IIT (ISM) Dhanbad.
I graduate in December 2026 and am looking for full-time robotics software, perception, and ML engineering roles — ideally where research ideas get shipped onto real hardware.
Pose estimation from a single demonstration, detector-free semi-dense matching and real-time tracking, MPC tuning via preferential Bayesian optimization, and Isaac Sim camera-placement optimization for CAD-free robotic fastening
Action Chunking with Transformers (ACT) policies from teleoperated demonstrations, LoRA fine-tuning of vision-language-action models (OpenVLA), and LeRobot data pipelines for robot arms
Cartographer / LiDAR SLAM in ROS2, persistent semantic mapping with vision foundation models (Grounding DINO), and zero-shot object-goal navigation through Nav2 without environment-specific training
JEPA-based world models for long-horizon planning, multi-view 3D scene reconstruction, deepfake audio detection, and differentiable physical simulation with gradient-based inverse design
Wrapped up my Robotics Software Internship at Mowito — shipped visual-servoing auto-fastening on FANUC CRX10iA-L and UR10e arms
Started as Robotics Software Intern at Mowito, working on visual-servoing-based CAD-free robotic fastening
Joined UMich research on JEPA-based world models for planning and differentiable electromagnetic-scattering simulation (Prof. Raj Rao Nadakuditi); released the freegaussianizer open-source package
Deepfake speech detection paper submitted to Interspeech 2026 (CMU, under review)
Started MS in ECE (Computer Vision) at University of Michigan
Mowito
Enhanced the full software stack for a visual-servoing-based auto-fastening system that derives target poses from a single demonstration, enabling CAD-free robotic bolt fastening on parts moving on a conveyor; demonstrated the system to investors, customers, and integrators. Added a detector-free semi-dense matching method and a real-time tracker so pose estimation stays robust through occlusion and low-texture parts, raising vision-stack throughput from 14.5 Hz to 21 Hz via compressed image transport. Built an Isaac Sim camera-mount pose optimization pipeline that renders and scores candidate mounts across CAD parts (90–565 mm) with multi-plane bolts via staged search over 1,000 candidates. Tuned the MPC controller with automatic weight tuning via preferential Bayesian optimization across 30 bolt positions, cutting mean settle time from 4.17 s to 2.67 s (31%) on a FANUC arm. Extended working conveyor speed from 60 to 100 mm/s on a UR10e, cutting cycle time up to 5.9× (21.3 s to 3.6 s) while holding 98% success across 150 runs. Ported the stack to a FANUC CRX10iA-L, resolving driver-level jerk instability to sustain a continuous 159-cycle run, and authored a state machine for multi-plane fastening reaching 95% success across 200 runs.
University of Michigan — Prof. Raj Rao Nadakuditi
Evaluating LeWM, a JEPA-based world model, with free-energy and coding-rate regularizers as replacements for the SIGReg anti-collapse regularizer — improving hard-protocol (100-step planning) accuracy on a two-room navigation dataset from 16% to 75% with the coding-rate variant. Showed that a free-energy (Gaussianizing) loss improves reconstruction and classification across CNNs, MLPs, and autoencoders, built a joint Gaussianizing classifier, and published freegaussianizer as an open-source package. Separately, translating the CyScat electromagnetic-scattering simulation codebase from MATLAB to Julia and Python and implementing differentiable simulation via forward-mode automatic differentiation through the T-matrix solver, enabling gradient-based optimization of refractive index and wavelength for maximum wave transmission through photonic structures; validated to machine precision against finite-difference checks.
Carnegie Mellon University — Dr. Arun Balajee Vasudevan
Engineered an end-to-end data generation and labeling pipeline to curate a 1M+ sample dataset of real/synthetic audio pairs in structured JSON. Fine-tuned the CLAP audio-language model achieving 97% accuracy on deepfake speech detection, outperforming human listeners. Paper submitted to Interspeech 2026 (under review).
Mowito
Developed Multi-head Mask R-CNN models for instance segmentation and object detection in warehouse environments and integrated them into production robot perception pipelines. Implemented perceptual hashing for large-scale dataset deduplication (30% size reduction) and built order-fulfillment optimization using Leiden clustering.
CAD-free robotic bolt fastening on parts moving on a conveyor, driven by single-demonstration pose estimation; 14.5→21 Hz vision stack and 60→100 mm/s conveyor speed at 98% success
Learn More →Cartographer SLAM + Grounding DINO in ROS2/Gazebo; a persistent semantic map lets natural-language commands trigger Nav2 goals for zero-shot object-goal navigation
Learn More →LoRA fine-tuning of OpenVLA (7B vision-language-action model) on LIBERO manipulation benchmarks; deployed on UMich HPC (NVIDIA A40), training loss 19.08→1.16
Learn More →Intro to Algorithmic Robotics (EECS 465), Mobile Robotics (ROB 530), Advanced Multi-Robot Systems (ROB 516), 3D Robot Perception (ROB 599)
Machine Learning (EECS 553), Advanced Topics in Computer Vision (EECS 542), Matrix Methods for ML & Signal Processing (ECE 551), Imaging & Advanced Machine Learning (BIOINF 590)
Probability & Random Processes (ECE 501), Computational & Data-Driven Methods in Engineering (MECHENG 599), Medical Imaging Systems (ECE 516)