I am a Master's student in Electrical and Computer Engineering (Computer Vision) at the University of Michigan, Ann Arbor. I build robotics software, perception, and machine learning systems for real-world robots — imitation learning and vision-language-action policies for robot arms, SLAM with language-guided navigation, and world models for long-horizon planning.
As a Graduate Research Assistant at UMich I work on JEPA-based world models for planning, and previously at Carnegie Mellon University I trained multimodal audio-language models for deepfake speech detection (under review at Interspeech 2026). My B.Tech is in Electrical Engineering from IIT (ISM) Dhanbad.
I graduate in December 2026 and am looking for full-time robotics software, perception, and ML engineering roles — ideally where research ideas get shipped onto real hardware.
Instance segmentation and object detection for warehouse robots, multi-view 3D scene reconstruction, LiDAR-based state estimation, and 3D perception pipelines in ROS2
Action Chunking with Transformers (ACT) policies from teleoperated demonstrations, LoRA fine-tuning of vision-language-action models (OpenVLA), and LeRobot data pipelines for robot arms
Cartographer / LiDAR SLAM in ROS2, persistent semantic mapping with vision foundation models (Grounding DINO), and zero-shot object-goal navigation through Nav2 without environment-specific training
JEPA-based world models for long-horizon planning, multi-view 3D scene reconstruction, deepfake audio detection, and differentiable physical simulation with gradient-based inverse design
Joined UMich research on JEPA-based world models for planning and differentiable electromagnetic-scattering simulation (Prof. Raj Rao Nadakuditi); released the freegaussianizer open-source package
Deepfake speech detection paper submitted to Interspeech 2026 (CMU, under review)
Started MS in ECE (Computer Vision) at University of Michigan
University of Michigan — Prof. Raj Rao Nadakuditi
Evaluating LeWM, a JEPA-based world model, with free-energy and coding-rate regularizers as replacements for the SIGReg anti-collapse regularizer — improving hard-protocol (100-step planning) accuracy on a two-room navigation dataset from 16% to 75% with the coding-rate variant. Showed that a free-energy (Gaussianizing) loss improves reconstruction and classification across CNNs, MLPs, and autoencoders, built a joint Gaussianizing classifier, and published freegaussianizer as an open-source package. Separately, translating the CyScat electromagnetic-scattering simulation codebase from MATLAB to Julia and Python and implementing differentiable simulation via forward-mode automatic differentiation through the T-matrix solver, enabling gradient-based optimization of refractive index and wavelength for maximum wave transmission through photonic structures; validated to machine precision against finite-difference checks.
Carnegie Mellon University — Dr. Arun Balajee Vasudevan
Engineered an end-to-end data generation and labeling pipeline to curate a 1M+ sample dataset of real/synthetic audio pairs in structured JSON. Fine-tuned the CLAP audio-language model achieving 97% accuracy on deepfake speech detection, outperforming human listeners. Paper submitted to Interspeech 2026 (under review).
Mowito
Developed Multi-head Mask R-CNN models for instance segmentation and object detection in warehouse environments and integrated them into production robot perception pipelines. Implemented perceptual hashing for large-scale dataset deduplication (30% size reduction) and built order-fulfillment optimization using Leiden clustering.
An ACT (Action Chunking with Transformers) policy trained on teleoperated demonstrations collected through HuggingFace LeRobot, driving screwdriver pick-and-place from a single front-facing camera
Learn More →Cartographer SLAM + Grounding DINO in ROS2/Gazebo; a persistent semantic map lets natural-language commands trigger Nav2 goals for zero-shot object-goal navigation
Learn More →LoRA fine-tuning of OpenVLA (7B vision-language-action model) on LIBERO manipulation benchmarks; deployed on UMich HPC (NVIDIA A40), training loss 19.08→1.16
Learn More →Intro to Algorithmic Robotics (EECS 465), Mobile Robotics (ROB 530), Advanced Multi-Robot Systems (ROB 516), 3D Robot Perception (ROB 599)
Machine Learning (EECS 553), Advanced Topics in Computer Vision (EECS 542), Matrix Methods for ML & Signal Processing (ECE 551), Imaging & Advanced Machine Learning (BIOINF 590)
Probability & Random Processes (ECE 501), Computational & Data-Driven Methods in Engineering (MECHENG 599), Medical Imaging Systems (ECE 516)