Skip to main content Skip to secondary navigation

mmWave Sensing and Sensor Fusion

Main content start

Note: Please view the desktop site in order to see images.

A Kinematics-Basis View of the World with mmWave Sensing

Perception in artificial systems has traditionally been grounded in geometry—reconstructing positions and structures to emulate human visual reasoning. Data is typically processed and analyzed for human interpretation. While effective, this paradigm is inefficient: recovering detailed geometry is computationally expensive and often unnecessary for tasks that rely primarily on predicting motion and interaction, especially when the end consumer is a machine rather than a human.

Biological systems demonstrate that motion, not static form, lies at the core of perception. Frogs’ retinas respond selectively to velocity patterns; bats and dolphins navigate by sensing Doppler shifts; and the human vestibular system encodes acceleration as the substrate of balance. Even in daily life, humans anticipate actions through unfolding kinematics rather than frozen shapes. This is not a biological shortcut but a reflection of a deeper physical truth: velocities and accelerations, not positions alone, govern how systems evolve through time.

A kinematics-basis perception framework builds on this view by shifting emphasis from reconstructing geometric scenes to modeling dynamical processes for prediction and control, thereby aligning artificial systems with both biological efficiency and physical law. This perspective parallels state-space models in control theory and Hamiltonian formulations in physics, where systems are described jointly by positions and velocities or momenta. A kinematics-basis framework is thus akin to treating perception as a phase-space embedding of the world, where the primary goal is not static reconstruction but modeling trajectories through time.

mmWave radar is well suited to this perspective because it directly and efficiently encodes velocity through Doppler shifts, whereas cameras and LiDAR capture geometry and recover motion only indirectly through methods such as optical flow. In this way, radar provides a natural foundation for kinematics-basis perception and exemplifies how sensing can align with the dynamical principles that truly govern the world.

By directly capturing motion and velocity, mmWave sensing enables high-resolution, efficient, and robust digitization of the dynamic world from a motion basis.

When Motion Meets Structure: Modality Fusion toward Complete World Models

Optical sensors such as cameras and LiDAR provide high-resolution spatial detail, geometry, and semantic context, while radar delivers motion cues and temporal dynamics that remain robust to distance, occlusion, and lighting. These modalities reinforce one another: optical sensors ground perception in spatial structure, while radar enriches it with temporal dynamics. Their fusion enables true spatial–temporal perception, enhancing robustness and efficiency across applications from human sensing and robotics to healthcare and smart environments.

Fusing structure from images with motion from Doppler signatures unlocks the potential for full spatiotemporal understanding of the world.

Publications

[1] S. Hor, S. Yang, J. Choi,  A. Arbabian, “MVDoppler: Unleashing the Power of Multi-View Doppler for MicroMotion-Based Gait Classification,” NeurIPS 2023.

[2] J. Choi, S. Hor, S. Yang, A. Arbabian, “MVDoppler-Pose: Multi-Modal, Multi-View mmWave Sensing for Long-Distance, Self-Occluded Human Walking Pose Estimation,” CVPR 2025.

[3] S. Yang, S. Hor, JH. Choi, A. Arbabian, "High-Resolution Gait Micro-Doppler Synthesis from Videos Over Diverse Trajectories, ICASSP, 2025.