mmHRITowards Privacy-Preserving Human-Robot Interaction
with Millimeter-Wave Radar

Under Review

Junqiao Fan1, Yuxuan Hu2, Bofan Lyu2, Yanshuo Lu2, Pengfei Liu2, Jiarui Zhang1,
Fangqiang Ding3, Lihua Xie1, Gen Li4,†, Jianfei Yang2,†

1 School of Electrical and Electronic Engineering, Nanyang Technological University

2 School of Mechanical and Aerospace Engineering, Nanyang Technological University

3 The Hong Kong University of Science and Technology (Guangzhou)

4 School of AI and Robotics, Hunan University

† Corresponding authors: Gen Li, Jianfei Yang

Motivation of privacy-preserving mmHRI using mmWave radar for human sensing. Unlike conventional vision-based HRI that fails under occlusion, mmHRI senses human intent through privacy curtains to coordinate object delivery without directly observing humans. This supports privacy-sensitive applications such as hospitals, restaurants, and unmanned stores.
Figure 1

Human intent, beyond visual barriers.

Motivation of privacy-preserving mmHRI using mmWave radar for human sensing. Unlike conventional vision-based HRI that fails under occlusion, mmHRI senses human intent through privacy curtains to coordinate object delivery without directly observing humans. This supports privacy-sensitive applications such as hospitals, restaurants, and unmanned stores.

mmHRI Video Demo

Radar-guided, privacy-preserving human–robot interaction.

Abstract

Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on RGB cameras that continuously observe humans to respond to non-verbal commands, such as hand gestures. This raises privacy concerns in privacy-critical environments, such as hospital wards or restaurants, where direct camera observation of humans is restricted. To develop privacy-preserving HRI, we leverage millimeter-wave (mmWave) radar, which can sense human motion through privacy barriers without identifiable imagery. We propose mmHRI, the first multi-modal robot manipulation framework that achieves mmWave radar-guided privacy-preserving HRI. mmHRI introduces two key designs to mitigate the sparsity and temporal inconsistency of radar data in cluttered robot manipulation environments. First, we propose a dual-stream architecture that jointly learns from unfiltered raw radar tensors and radar point clouds to estimate both human actions and 3D poses. To mitigate signal inconsistency, mmHRI further incorporates a memory-based state-space model (MSSM) that retains historical radar features to reduce abrupt changes in pose/action. These estimated human states are then converted into structured textual robot instructions, which control a vision-language-action (VLA) policy for closed-loop robot manipulation and human-aware reactions. Our evaluation covers human action recognition and closed-loop delivery and retrieval. In the privacy-preserving curtain setting, mmHRI achieves 85.09% action-recognition accuracy, outperforming existing radar-based alternatives. Robot trials further demonstrate successful delivery and retrieval under visual occlusion, with stable task performance across unseen subjects, clutter configurations, and environments.

System Setup

Cameras observe the robot workspace. mmWave radar senses human intent in the private area.

Left: Physical robot system setup and the mmWave sensing unit. Middle: Two privacy-preserving scenes, i.e., hospital ward and office. Right: Three gesture/pose based human-robot-interaction decision-making tasks. We show different radar signals that the robot may refer to for different reactions.
Figure 2

A robot platform for privacy-preserving interaction

Left: Physical robot system setup and the mmWave sensing unit. Middle: Two privacy-preserving scenes, i.e., hospital ward and office. Right: Three gesture/pose based human-robot-interaction decision-making tasks. We show different radar signals that the robot may refer to for different reactions.

The mmHRI Framework

From complementary radar measurements to human-aware robot actions.

Overview of mmHRI. Radar measurements are first preprocessed into RDT tensors and RPC. The dual-stream radar human perception (DRP) then extracts motion and geometry patterns from both modalities, which are fused to jointly predict action and 3D poses. The human-aware text reasoning (HTR) then converts the estimated human states into structured robot instructions, controlling the downstream VLA policy for different actions.
Figure 3

Dual-stream perception meets robot manipulation

Overview of mmHRI. Radar measurements are first preprocessed into RDT tensors and RPC. The dual-stream radar human perception (DRP) then extracts motion and geometry patterns from both modalities, which are fused to jointly predict action and 3D poses. The human-aware text reasoning (HTR) then converts the estimated human states into structured robot instructions, controlling the downstream VLA policy for different actions.

Dual-stream radar perception (DRP Fusion)

Raw radar tensors capture motion patterns, while radar point clouds preserve spatial structure for joint action and 3D pose estimation.

Memory state-space model (MSSM)

Historical radar features reduce temporal inconsistency and abrupt changes in the estimated human state.

Human-aware text reasoning (HTR)

Human states become structured instructions that condition a vision-language-action policy for closed-loop manipulation.

Visualization of the complementary radar modalities used by the dual-stream model. The DT and RT maps preserve motion-sensitive Doppler patterns, whereas the RPC retains 3D spatial structure for pose and root estimation. Synchronized RGB images are shown for visual reference.
Figure 4

Complementary motion and geometry

Visualization of the complementary radar modalities used by the dual-stream model. The DT and RT maps preserve motion-sensitive Doppler patterns, whereas the RPC retains 3D spatial structure for pose and root estimation. Synchronized RGB images are shown for visual reference.

Radar Perception Results

mmHRI achieves 85.09% action-recognition accuracy under privacy occlusion and 74.91% under cross-environment occlusion.

Table 1Performance on the radar perception dataset. HPE is evaluated under clear visibility, and HAR is evaluated under clear visibility, privacy occlusion, and cross-environment occlusion.
MethodsHPE · Clear VisibilityHAR · Clear VisibilityHAR · Occlusion (Privacy)HAR · Occlusion (Cross-Env)
MPJPE (cm)TE (cm)ACC (%)FPR (%)Static-F1 (%)HPE JitterACC (%)FPR (%)Static-F1 (%)HPE JitterACC (%)FPR (%)Static-F1 (%)HPE Jitter
RGB
RGB (YOLOv.) + AGCN——96.024.5597.06—————————
mmWave Radar
PointTrans. (RPC) + AGCN7.406.2686.1816.1889.583.5272.4923.1188.363.9131.0640.8467.916.08
RadHAR (RPC)——87.0611.9577.11—62.1826.0982.50—27.6133.1061.81—
Waveman (RT)——91.654.4791.33—79.6512.6786.95—62.310.0082.28—
Ours
mmHRI (Fusion)7.546.0192.564.5595.143.5481.455.0088.693.4568.9420.4285.835.15
mmHRI (Fusion + MSSM)4.253.1296.511.3897.252.8385.099.9090.173.2574.9111.2786.094.42

ACC: accuracy. FPR: false-positive rate. HPE: human pose estimation. HAR: human action recognition. MPJPE: mean per-joint position error. TE: translation error. Jitter is reported in centimeters; — denotes an unreported or unavailable result.

Real-World Robot Trials

Side decisions, object delivery, box retrieval, and collision avoidance under clear visibility and visual occlusion.

Table 2Performance on real-world robot trials. PA and Success report the perception-only and end-to-end stage success rates. MA reports the success rate over 30 manipulation-only trials with ground-truth human states.
MethodsSide Decision & GraspingObject DeliveryBox RetrievalCollision Avoidance
PAMASuccessPAMASuccessPAMASuccessPAMASuccess
Clear Visibility
RGB + π0.530/3026/3026/3030/3027/3027/3027/3030/3027/3030/3030/3030/30
mmWave + π0.530/3026/3029/3026/3029/3029/3030/3030/30
Occlusion
RGB + π0.50/3026/300/300/3027/300/300/3030/300/300/3030/300/30
mmWave + π0.528/3025/3028/3025/3029/3029/3030/3030/30
Occlusion (Cross-Environment)
RGB + π0.50/3023/300/300/3026/300/300/3028/300/300/3030/300/30
mmWave + π0.528/3021/3025/3022/3026/3025/3025/3025/30
Qualitative examples of the human states and corresponding robot operations. The columns show an azimuth wave, waiting, approaching, and a radial wave; the camera views show object grasping, tray delivery, and tray retrieval during the interaction.
Figure 5

Human intent guides the interaction sequence

Qualitative examples of the human states and corresponding robot operations. The columns show an azimuth wave, waiting, approaching, and a radial wave; the camera views show object grasping, tray delivery, and tray retrieval during the interaction.

Robustness & Generalization

Evaluating radar perception with unseen subjects and changes in tabletop clutter.

Table 3Generalization to unseen subjects.
SubjectDelivery ActivationRetrieve ActivationCollision Avoidance
Seen Subject28/3029/3030/30
Unseen Subject 130/3030/3030/30
Unseen Subject 227/3026/3030/30
Table 4Generalization to unseen tabletop clutter.
ClutterDelivery ActivationRetrieve ActivationCollision Avoidance
No Additional Clutter28/3029/3030/30
Metal Box28/3026/3030/30
Tall Cup Box26/3030/3030/30
Real-world robustness evaluation setups. (a) Cross-table-clutter settings with a metal box and an occluding paper box. (b) Cross-subject settings with two unseen subjects.
Figure 6

Testing beyond the training setup

Real-world robustness evaluation setups. (a) Cross-table-clutter settings with a metal box and an occluding paper box. (b) Cross-subject settings with two unseen subjects.