Hi nuScenes team — disclosure up front: I work with thyn-ai, where we publish clean-room Mojo kernels as optional drop-in accelerators for popular Python packages (https://github.com/thyn-ai/mojo-kernels, Apache-2.0). The devkit is the reference implementation here — it is the oracle we test and benchmark against, and what follows is an optional fast path, not a replacement.
We've just landed nuscenes-eval-mojo on main:
https://github.com/thyn-ai/mojo-kernels/tree/main/python/nuscenes_eval_mojo
It runs the detection eval's greedy center-distance matching fan-out — the matching loop the official evaluator re-runs for every class × distance-threshold pair (10 classes × 4 thresholds), with a NumPy scalar norm per candidate box pair — in one compiled kernel call. Box construction, filtering, and AP/TP assembly stay in shared Python/NumPy code, so both backends produce identical results.
Measured vs the DetectionEval path of nuscenes-devkit 1.2.0 (Apple M4 Max; deterministic seeded synthetic corpora — 8–20 gt annotations and ~25–30 predictions per sample across all 10 classes; median of 5; correctness against the devkit asserted before every timing pass; NuScenes DB loading excluded on both sides):
| corpus |
samples |
cold first evaluation |
warm steady state |
| small |
60 |
68.5× |
6.2× |
| medium |
300 |
25.8× |
7.0× |
| large |
1,200 |
13.8× |
6.7× |
"Cold" is the first evaluation in a fresh interpreter (imports done; on our side this includes the one-time dlopen + ABI handshake); "warm" is steady-state per-evaluation latency. Raw latencies and full methodology are in the package README.
Identical output: the returned dict follows the DetectionMetrics.serialize() schema (label_aps, mean_dist_aps, mean_ap, label_tp_errors, tp_errors, tp_scores, nd_score, eval_time, cfg) and matches the devkit's full metrics dict within 1e-8 on every corpus in the differential battery — measured agreement 8.9e-16 on the benchmark corpora (ulp-level). The battery covers exact confidence ties, zero-gt classes, zero-prediction classes, heavy false positives, bike-rack filtering, NaN velocities, duplicate detections, and a custom DetectionConfig. The whole suite runs twice — native backend and vendored pure-Python fallback — each asserted against the real nuscenes-devkit package.
API shape mirrors what DetectionEval.evaluate() returns:
import nuscenes_eval_mojo
metrics = nuscenes_eval_mojo.evaluate(gt_corpus, submission_json)
metrics["mean_ap"] # mAP
metrics["nd_score"] # NDS
metrics["tp_errors"] # mATE / mASE / mAOE / mAVE / mAAE
Current scope, honestly: 3D detection only (DetectionBox; no tracking / prediction / panoptic), dist_fcn='center_distance' only, and DB loading + rendering stay with the devkit. Everything is reproducible from the repo today: the seeded benchmark (benchmarks/bench_nuscenes_eval.py), the differential suite (tests/test_nuscenes_eval_differential.py, run native and forced-fallback via scripts/test_all_nuscenes_eval.sh), and the wheel-build script (python/nuscenes_eval_mojo/build_wheel.sh) are all in the tree — the wheel is not on PyPI yet, so building from the repo is the path for now.
Eval performance has come up here before on the tracking side (#884, closed — that thread was about motmetrics, a different eval path from this one).
If a faster detection eval would be useful, we're glad to answer questions, run corpora you suggest, or adjust the interface. And either way — thanks for maintaining the devkit; a well-specified reference is what makes an output-identical comparison like this possible.
Hi nuScenes team — disclosure up front: I work with thyn-ai, where we publish clean-room Mojo kernels as optional drop-in accelerators for popular Python packages (https://github.com/thyn-ai/mojo-kernels, Apache-2.0). The devkit is the reference implementation here — it is the oracle we test and benchmark against, and what follows is an optional fast path, not a replacement.
We've just landed
nuscenes-eval-mojoon main:https://github.com/thyn-ai/mojo-kernels/tree/main/python/nuscenes_eval_mojo
It runs the detection eval's greedy center-distance matching fan-out — the matching loop the official evaluator re-runs for every class × distance-threshold pair (10 classes × 4 thresholds), with a NumPy scalar norm per candidate box pair — in one compiled kernel call. Box construction, filtering, and AP/TP assembly stay in shared Python/NumPy code, so both backends produce identical results.
Measured vs the DetectionEval path of nuscenes-devkit 1.2.0 (Apple M4 Max; deterministic seeded synthetic corpora — 8–20 gt annotations and ~25–30 predictions per sample across all 10 classes; median of 5; correctness against the devkit asserted before every timing pass; NuScenes DB loading excluded on both sides):
"Cold" is the first evaluation in a fresh interpreter (imports done; on our side this includes the one-time
dlopen+ ABI handshake); "warm" is steady-state per-evaluation latency. Raw latencies and full methodology are in the package README.Identical output: the returned dict follows the
DetectionMetrics.serialize()schema (label_aps,mean_dist_aps,mean_ap,label_tp_errors,tp_errors,tp_scores,nd_score,eval_time,cfg) and matches the devkit's full metrics dict within 1e-8 on every corpus in the differential battery — measured agreement 8.9e-16 on the benchmark corpora (ulp-level). The battery covers exact confidence ties, zero-gt classes, zero-prediction classes, heavy false positives, bike-rack filtering, NaN velocities, duplicate detections, and a customDetectionConfig. The whole suite runs twice — native backend and vendored pure-Python fallback — each asserted against the real nuscenes-devkit package.API shape mirrors what
DetectionEval.evaluate()returns:Current scope, honestly: 3D detection only (
DetectionBox; no tracking / prediction / panoptic),dist_fcn='center_distance'only, and DB loading + rendering stay with the devkit. Everything is reproducible from the repo today: the seeded benchmark (benchmarks/bench_nuscenes_eval.py), the differential suite (tests/test_nuscenes_eval_differential.py, run native and forced-fallback viascripts/test_all_nuscenes_eval.sh), and the wheel-build script (python/nuscenes_eval_mojo/build_wheel.sh) are all in the tree — the wheel is not on PyPI yet, so building from the repo is the path for now.Eval performance has come up here before on the tracking side (#884, closed — that thread was about motmetrics, a different eval path from this one).
If a faster detection eval would be useful, we're glad to answer questions, run corpora you suggest, or adjust the interface. And either way — thanks for maintaining the devkit; a well-specified reference is what makes an output-identical comparison like this possible.