FusionCore verification study
Testing three claims from a published ROS 2 sensor-fusion paper against its code and data. Two did not hold up; the third did, on a sequence the paper didn't examine.
- 2026
- Complete
- Independent verification: state estimation
- C++
- ROS 2 Jazzy
- Python
- evo
- NCLT dataset
FusionCore by Manan Kharwar (original repository (opens in a new tab)), Apache 2.0. Data: NCLT dataset, University of Michigan, ODbL.

FusionCore is an open-source ROS 2 package that fuses IMU, wheel encoder, GPS and visual SLAM data with a 23-state Unscented Kalman Filter. Its paper reports lower trajectory error than robot_localization, the standard ROS package for this job, on 10 of 12 sequences of the University of Michigan NCLT dataset. I took three of its claims and tested each one against the source code and the original data.
Problem
I chose this paper because its claims can be checked. I tested three of them:
- That a velocity pre-gate, proposed in the paper's future work, would fix its documented failure on the 2012-08-20 sequence.
- That the 23rd state, an online estimate of the wheel encoder's yaw-rate bias and the paper's main novelty, improves performance. The paper never isolates it.
- That the failure mechanism is covariance growth: during a GPS blackout the filter's uncertainty grows until its chi-squared outlier gate stops recognising corrupted fixes, and accepts them.
The code changes were small. Most of the work was establishing what they did, where full 90-minute runs turned out to be unreliable in a way that became the main finding.
Approach
Validate the build
My first full run on 2012-01-08 scored 61.1 m of trajectory error against the published 18.6 m. I ruled out playback rate, GPS coordinate conversion, ground-truth construction and numerical differences on ARM64, then compared my build directly against the author's own output for a GPS-spike test. The two trajectories agreed to within 2.6 m for the whole run, and to 1.1 m on average. The build was correct. What remained were isolated excursions of a few hundred metres on long runs, which I couldn't yet explain, so short injected-fault experiments became the main evidence and full sequences became supporting material.
Implement the proposed pre-gate
The check compares a fix's implied speed against a limit. I placed it after the existing quality gate and before the delayed/on-time branch, the one point every fix passes through. Its reference position only advances when a fix is accepted, so one rejection cascades through a burst of bad fixes without any cluster-detection logic. It is off by default and covered by two unit tests.
Add a switch for the 23rd state
The state can only adapt through a single process-noise entry, so a runtime scale on that entry freezes it without changing the filter's dimension: still 23 states and 47 sigma points. The repository had no test that exercised the state with a non-zero value. I wrote one, which also pins the sign convention the code actually uses.
Fix the benchmark harness
Recorded odometry was between 38% and 95% duplicate timestamps. The data player never exited, its shutdown raised an error, the launch file never stopped the stack, and the node republished whenever simulated time hadn't moved. I fixed all four. Re-scoring showed the error figures were unaffected, but message counts had been sending me after differences that didn't exist.
Record every gate decision
With the filter's debug topics recorded, I had the gate's decision and the squared Mahalanobis distance (d²) for every GPS fix, and checked each fix against RTK ground truth. This is what overturned the paper's diagnosis of 2012-08-20, and what found the real failure on 2012-01-08.
Results
- 1000 ×
- 157 m off
- 7,340
The proposed pre-gate has nothing to fix
The paper says the fix arriving after a 211-second blackout implies about 3400 m/s. It is 713.8 m from the previous fix, 211.19 s later: 3.4 m/s, which passes any sensible speed threshold. The paper's figure is a thousand times too high, consistent with dividing by 0.211 s instead of 211 s. The check also has a structural flaw. Its limit grows with the time since the last fix, so after a long blackout it goes blind for the same reason the statistical gate is said to.

Transcript
The proposed velocity check measures a new fix against the last accepted fix and rejects anything faster than 20 m/s, so the distance it allows grows with the time since that fix. A cursor sweeps that time from 0.1 s to 1,000 s.
Captions, in order:
- The proposed check divides a new fix's distance by the time since the last accepted fix, and rejects anything faster than 20 m/s.
- Between normal fixes, 0.2 s apart, it allows 4 m of movement.
- As the gap lengthens, the distance it will accept grows with it.
- After a real 112 s signal loss, a fix 157 m from the truth implies 1.48 m/s. It passes.
- After 211 s, the fix the paper cites implies 3.4 m/s, not 3,400. It passes.
- The allowance grows with the gap. After 211 s it is 4.2 km, so the longer the blackout, the less it can catch.
"Passes" means the proposed check would let the fix through. It says nothing about the filter's own chi-squared gate, which rejected the 2012-08-20 fix and accepted the 2012-01-08 one (report sections 3 and 5).
The failure wasn't there to fix either. With every gate decision recorded on 2012-08-20, the chi-squared gate rejected the entire corrupted cluster. The fix the paper names scored d² = 90.6, more than five times the 16.27 threshold. The error in that stretch is dead-reckoning drift while GPS was absent and then refused, correctly for all but the last nine seconds, and no GPS gate can reduce it.
The 23rd state didn't help
I ran four runs on 2012-01-08: the state active or frozen, with and without a 200-second injected GPS outage. robot_localization ran alongside as an untouched control and varied by only 0.8 m across all four, so the runs are comparable.
| 23rd state active | 23rd state frozen | |
|---|---|---|
| No outage | 66.5 m | 60.7 m |
| 200 s outage | 294.6 m | 75.1 m |

Freezing the state gave lower error in both conditions. The large gap comes from a GPS lockout, described below: the active run locked out and never recovered, finishing with an estimated path 12% shorter than the true one. The one direct measurement points to chance: while coasting without usable GPS, the two runs drifted by similar amounts, 198 m and 170 m, in different directions, and that decided which side of the gate each landed on. With one run per configuration I can't say whether the state makes lockouts systematically worse, but the claimed improvement isn't supported.
The paper's failure, on a different sequence
2012-01-08 has natural gaps where the receiver genuinely lost its signal. Checking the first fixes the receiver produced after each gap against ground truth explains every lockout I found.
After a 112-second signal loss, the receiver's first fix was 157 m from the truth, and the next few converged quickly: 121, 84 and 68 m, then 5 m within four seconds. The filter's uncertainty had grown to about 43 m during the gap, so the first fix scored d² = 9.6 and 12.7 in the two outage runs, under the 16.27 threshold, and was accepted. This is the mechanism the paper describes. The proposed pre-gate wouldn't have stopped it either: measured from the last accepted fix, it implies 1.48 m/s.
What the paper doesn't describe is what happens next. Accepting the bad fix collapsed the filter's uncertainty to about 3 m, at the wrong place. The fixes that followed, almost all of them good, were converging on the truth faster than a 3 m uncertainty allows, so the gate rejected every one of them. Checked against ground truth, the rejected fixes were ordinary GPS, a few metres from the truth, much like the accepted ones.

Once a lockout starts, it feeds itself. The filter drifts, so correct fixes look like outliers, so it keeps drifting. Escape depends on a single fix scoring just under the threshold. Every run locked out at this gap for about a thousand fixes, and in one run the filter refused the last 7,340 fixes of the sequence.

Transcript
Two synchronised map panels of the same stretch of NCLT 2012-01-08, one run with the 23rd state active and one with it frozen, both with the 200 s injected outage earlier in the run. Each panel shows the true path (RTK), the true position, the filter's estimate with its 1-sigma uncertainty, and each GPS fix as it arrives: a green dot if accepted, a red cross if rejected. A readout gives the error against the truth, the uncertainty, the latest fix's Mahalanobis distance against the 16.27 gate, and the number of fixes rejected in a row. A timeline underneath marks the GPS gaps and every accepted and rejected fix.
Captions, in order:
- Normal operation. Each GPS fix is accepted and keeps the estimate on the true path.
- The receiver loses its signal. The filter dead-reckons, and its uncertainty grows.
- The signal returns. The first fix is 157 m from the truth, but the enlarged gate lets it through.
- Uncertainty collapses to 3 m at the wrong place. The good fixes that follow are all rejected.
- Locked out. Every fix is refused, so nothing can correct the drift.
- The signal is lost again, this time for 203 s.
- GPS returns and the fixes are good, but both estimates are now hundreds of metres away.
- Frozen run: one fix scores 16.25 against the 16.27 threshold, and the estimate snaps back. Active run: never.
It ends with the active run 536 m from the truth after 1,357 consecutive rejections, and the frozen run 11 m from it.
Time runs at up to 51x through quiet stretches and is slowed to 0.04x around the corrupted fix, with the speed shown on screen. Both runs are shown on one clock (the frozen run's; the two runs' zeros differ by 0.601 s). Positions are drawn exactly as recorded, with no alignment; accepted fixes sit a median 0.64 m from the estimate in normal operation, confirming the frames agree.
This also explained the long-run excursions from the first step: each one I could check is the moment the filter finally accepted GPS after minutes of refusing it. And it explains why none of my injected outages triggered this failure. The injector deletes fixes while the receiver keeps its signal, so the first fix afterwards is accurate. A real signal loss produces a bad first fix, and a test built only on injected outages misses this failure entirely.
Where the paper and the code disagree
- Equation 13 subtracts the encoder bias; the code adds it, consistent with the paper's own equation 11.
- The coast-mode bias subtraction described in the abstract doesn't exist in the code.
- The gate relaxation blamed for the 2012-08-20 failure is switched off in the benchmark configuration, with the author's comment explaining it caused outlier acceptance.
- The 23rd state is only observable through GPS, which the paper doesn't mention.
Reflection
What I learned
Checking claims took far longer than implementing changes. The pre-gate and the ablation switch were a few dozen lines between them. Establishing what they did required validating the build against the author's output, fixing the harness, and recording every gate decision.
My controlled experiments also had a blind spot I didn't see at first. I moved to injected faults because they were repeatable and easy to defend, but an injected outage is easier than a real one, and the failure that mattered only appears after a real signal loss. The data overturned my first reading more than once: I initially took the 46 fixes rejected after an injected outage as the gate working. One was the spike I had injected; the other 45 were good fixes, an early sign of the lockout.
Absolute errors in my runs are inflated by the lockouts, so they shouldn't be compared with the paper's figures.
Next iteration
- Find out why the author's runs, which score 18.6 m on the same sequence, apparently avoid the lockout, even though the corrupted first fix is in the data he used too.
- Repeat the ablation around real signal losses rather than injected outages, across more sequences, to separate a systematic effect of the 23rd state from chance.
- Run 2012-06-15, the paper's longer-blackout loss, which may show the same failure again.
- Explain the
robot_localizationbaseline, which scored 5–20× worse than the published figures in every run.