m2C2: mmWave-Camera Calibration via Human Pose

1Keio University   2Kyoto University
IROS 2026 MIRU 2026
IEEE/RSJ International Conference on Intelligent Robots and Systems (Pittsburgh, PA, USA) · Meeting on Image Recognition and Understanding (Nagasaki, Japan)

m2C2 calibrates an mmWave radar and an RGB camera using only the pose of a person moving in the scene, with no calibration target. The video plays muted; unmute for narration.

Abstract

This paper presents m2C2, a markerless calibration method for estimating the 6-DoF extrinsic transformation between mmWave radar and RGB cameras using natural human joint poses. Despite the growing demand for robust multi-sensor perception systems, calibration remains challenging, not only because traditional calibration methods rely on controlled environments and manual, target-based setups, but because mmWave radar measurements are inherently sparse and noisy, making stable geometric correspondences difficult. m2C2 addresses these limitations by utilising natural human motion as a shared calibration domain, enabling automated recalibration without dedicated calibration targets. Our framework jointly optimises the extrinsic translation and rotation by minimising an objective function over 2D and 3D human keypoints. Our approach mitigates the noise and sparsity of mmWave data through a novel bone-length scaling mechanism, jointly optimised with the extrinsics, that absorbs systematic skeleton-proportion errors from the radar pose estimator, supported by a confidence-weighted reprojection loss. Verified through experiments on a real-world dataset, m2C2 achieves an average translation error of 13.23 cm and a rotation error of 1.51°, relative to the ground truth, outperforming baseline PnP methods. Sensitivity analysis further demonstrates the method's robustness by displaying comparable results across different parameters, configurations, and subjects. These results demonstrate that m2C2 provides a robust and practical alternative to traditional marker-based calibration.

Method

m2C2 estimates the rigid transformation \(T = \{R, \mathbf{t}\}\) that maps points from the mmWave radar frame to the camera frame. Instead of matching raw radar points to pixels, both sensors are first converted into the same representation: the joints of an SMPL human skeleton. The joint streams are aligned in time, and the extrinsics are then solved by minimising a confidence-weighted reprojection error while jointly optimising per-bone length scale factors. Since the method only requires human pose, it can be applied to any sensor pair that supports SMPL pose estimation.

Left: conventional calibration with a ChArUco board and four corner reflectors. Right: m2C2 aligns the radar and camera using the joints of a person walking in the scene.

Conventional target-based calibration (left) requires a board with a ChArUco pattern for the camera and corner reflectors for the radar. m2C2 (right) removes the physical target and uses natural human joint poses as the shared reference for aligning the two sensors.

m2C2 pipeline: data representation with mmMesh, mmJoints and HybrIK; temporal stream alignment with cross-correlation and PCHIP resampling; weighted optimisation with EPnP initialisation, bone length scaling and reprojection error.

Overview of the m2C2 framework. Raw data from the mmWave radar and RGB camera are processed independently to extract 3D and 2D SMPL keypoints. The streams are temporally aligned using motion-signature cross-correlation and PCHIP resampling. Finally, the extrinsic parameters \((R, \mathbf{t})\) are estimated by minimising the confidence-weighted reprojection error, accounting for optimisable bone-length scaling.

Shared SMPL representation

3D joint positions \(\mathbf{p}^{\mathcal{M}}_{j,k}\) are estimated from the radar point cloud with mmMesh, and mmJoints adds a per-joint confidence \(w^{\mathcal{M}}_{j,k}\). On the camera side, HybrIK provides 2D joint positions \(\mathbf{p}^{\mathcal{C}}_{j,k}\), with a confidence \(w^{\mathcal{C}}_{j,k}\) taken from the heatmap peak. Both outputs are truncated to the same \(N = 22\) joints so every radar joint has a one-to-one partner in the image.

Temporal stream alignment

The two sensors have different latencies, leaving a sub-frame offset \(\Delta\tau\) between the streams. Joint positions cannot be compared directly across 3D and 2D, but joint speeds correlate well once the streams are aligned. We therefore compute a motion signature from the median speed of six robust joints \(\mathcal{J}\) (pelvis, ankles, neck and shoulders), normalise it with the median absolute deviation, and find the offset that maximises the continuous cross-correlation:

\[ \sigma[k] = \operatorname*{median}_{j \in \mathcal{J}} \left\| \mathbf{p}_{j,k} - \mathbf{p}_{j,k-1} \right\|_2, \qquad \Delta\tau^{*} = \operatorname*{arg\,max}_{\Delta\tau \in [-w,\, w]} \sum_{k} \tilde{\sigma}^{\text{3D}}(\tau_k)\, \tilde{\sigma}^{\text{2D}}(\tau_k + \Delta\tau) \]

The radar stream is then resampled at the shifted camera timestamps using piecewise cubic Hermite interpolation (PCHIP).

Confidence-weighted reprojection

Each radar joint is projected into the image with the current estimate of \(T\), and compared with the matching 2D joint. The error of every joint is weighted by the geometric mean of the min–max normalised confidences from both sensors, so a joint that is unreliable in either modality has little influence:

\[ e_{j,k} = \left\| \pi(T, \mathbf{p}^{\mathcal{M}}_{j,k}) - \mathbf{p}^{\mathcal{C}}_{j,k} \right\|_2, \qquad w_{j,k} = \sqrt{\bar{w}^{\mathcal{M}}_{j,k}\, \bar{w}^{\mathcal{C}}_{j,k}} \] \[ \mathcal{L}_{\text{reproj}} = \sum_{k} \sum_{j} w_{j,k}\, \rho(e_{j,k}) \]

where \(\rho\) is the Huber loss. An initial estimate \(T_0\) is obtained with RANSAC-based EPnP over all joint pairs.

Diagram of 3D radar keypoints projected onto the camera image plane, with the reprojection error drawn between projected and detected 2D joints.

Radar keypoints \(P^{\mathcal{M}}\) (red) are projected with \(T\) and compared with the RGB keypoints \(P^{\mathcal{C}}\) (blue).

Bone-length scaling

mmMesh and HybrIK estimate skeletons with systematically different proportions. If left uncorrected, the optimiser compensates for this by moving the camera, which biases the extrinsics. m2C2 instead rescales each bone along the kinematic chain with a learnable factor \(s_j = \exp(\lambda_j)\), which changes bone lengths while preserving joint angles. The scale factors are optimised jointly with \(T\), with a regulariser that keeps them close to the estimator's original lengths:

\[ \hat{\mathbf{p}}_j = \hat{\mathbf{p}}_{\mathrm{pa}(j)} + s_j \left( \mathbf{p}_j - \mathbf{p}_{\mathrm{pa}(j)} \right), \qquad T^{*}, \{\lambda_j^{*}\} = \operatorname*{arg\,min}_{T,\, \{\lambda_j\}} \; \mathcal{L}_{\text{reproj}} + \lambda_{\text{reg}} \sum_{j=1}^{N} \lambda_j^{2} \]

Dataset

We recorded a real-world mmWave–RGB dataset with five subjects (four male, one female) of diverse body types. Each subject performed two motion sequences under four sensor configurations that vary the horizontal and vertical offset and the rotation between the radar and camera. Results are averaged over 20 recorded sequences.

Each sequence is 3,000 frames (about 5 minutes at 10 FPS), recorded at 640×480 with subjects 1.6–5.0 m from the sensors performing lateral and radial walking. Ground-truth extrinsics come from a dedicated calibration take with a target-based board, captured at 1920×1080 before each recording while the sensors stayed fixed.

Calibration board used for ground truth: a ChArUco pattern surrounded by four metal corner reflectors.

Target-based calibration board (ChArUco pattern with four corner reflectors), used only to obtain ground truth.

Quantitative Results

We compare the estimated extrinsics against the target-based ground truth using the translation error \(E_t\) (cm), the rotation error \(E_R\) (°) and the mean reprojection error (MRE, px). Across all 20 sequences, m2C2 reaches 13.23 cm and 1.51°, compared with 23.72 cm and 2.48° for a standard PnP baseline on the same joints.

Ablation study

Table 1. Ablation study averaged over 20 sequences. Bold / underline: best / second best.
Method\(E_t\) (cm) ↓\(E_R\) (°) ↓MRE (px) ↓
Baseline (PnP)23.722.489.05
Baseline (PnP) + Time Shift23.322.396.36
m2C2 (Full)13.231.515.41
w/o Time Shift14.211.548.29
w/o Conf. Weighting13.311.555.41
w/o Bone Scaling18.131.956.36
w/o Conf. + Bone18.281.976.35
w/o Time + Bone18.452.039.03

Bone-length scaling is the largest single contributor, accounting for 4.90 cm of the reduction in translation error. Temporal alignment improves translation error by 0.98 cm but reprojection error by 2.88 px, which reflects the small inter-stream offsets in this dataset (\(|\Delta\tau| \le 0.3\) s). Confidence weighting contributes a further 0.08 cm.

Sensor configurations

Table 2. Mean calibration error across the four offset configurations.
Configuration\(E_t\) (cm) ↓\(E_R\) (°) ↓MRE (px) ↓
Horizontal Offset (L/R)15.141.747.60
Horizontal Offset (F/B)11.270.584.28
Vertical Offset11.310.424.21
Vertical Offset + Rotation14.823.135.33

Errors are higher for the lateral and rotated configurations. The reprojection objective has a depth–scale ambiguity that is only resolved when the subject moves along the viewing direction, so purely side-to-side motion is the main failure case of the method.

Robustness

Table 3. Calibration accuracy by subject.
Subject\(E_t\) (cm) ↓\(E_R\) (°) ↓MRE (px) ↓
Subject A13.771.144.16
Subject B14.870.844.53
Subject C17.462.815.35
Subject D10.572.397.79
Subject E10.550.725.21
Line chart of translation, rotation and reprojection error against number of frames from 50 to all; errors fall steeply up to 150 frames and level off after about 500.

Accuracy versus the number of sampled frames. Errors drop steeply up to about 150 frames and level off after about 500.

Table 4. Sensitivity to the bone-length regularisation weight \(\lambda_{\text{reg}}\) (default \(10^{-1}\)). Bone Dev.: mean absolute deviation of the optimised bone lengths from the estimator's original lengths.
\(\lambda_{\text{reg}}\)\(E_t\) (cm) ↓\(E_R\) (°) ↓MRE (px) ↓Bone Dev. (%) ↓
\(10^{-4}\)13.041.515.4116.93
\(10^{-3}\)13.041.515.4116.92
\(10^{-2}\)13.061.515.4116.81
\(10^{-1}\)13.231.515.4115.96
\(10^{0}\)14.291.585.4511.95
\(10^{1}\)17.251.775.745.81

Results are consistent across subjects, and translation error stays within 13.04–13.23 cm for \(\lambda_{\text{reg}} \le 10^{-1}\). With very strong regularisation (\(10^{1}\)) the scale factors are pushed towards 1 and accuracy approaches the "w/o Bone Scaling" ablation.

Qualitative Results

Four panels: HybrIK 2D pose on the RGB frame, mmMesh 3D pose, radar joints projected onto the RGB frame with the estimated extrinsics, and the raw mmWave point cloud.
Output of the m2C2 pipeline: HybrIK pose estimated from RGB (top left), mmMesh pose estimated from radar (top right), radar joints projected into the image with the estimated extrinsics (bottom left), and the raw mmWave point cloud (bottom right).

BibTeX

@inproceedings{shilton2026m2c2,
  title     = {m2C2: mmWave-Camera Calibration via Human Pose},
  author    = {Shilton, Jack and Isogawa, Mariko and Sakurada, Ken},
  booktitle = {Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
  year      = {2026}
}

Acknowledgements

This work was partially supported by JSPS KAKENHI (B) 26K02929.