Mixterio

The boardmeasured 2026-07-16

The numbers, and the means to check them.

Mixterio is the only DJ product that publishes reproducible, held-out analysis accuracy. Every row below states the dataset, the split, and the comparison figures: published ones where they exist, and our own same-dataset measurements, labelled as ours, where a tool publishes none. Where a stronger number exists, it is on the chart next to ours.

Key detection74.80MIREX weighted

72.67 Rekordbox 7measured by us, same 566 tracks

held-out, shipping app
Beat tracking89.23F-measure

89.1 Beat This!Foscarin, ISMIR 2024

held-out, never trained on
Downbeat tracking78.69F-measure

78.3 Beat This!Foscarin, ISMIR 2024

held-out, never trained on
Tempo81.85ACC1, strict

no published comparisonnobody publishes a figure for this metric, so we draw none

raw grid, general material
Structure0.4025boundary HR3F

0.472 Foote (2000)measured by us, same corpus

held-out, coarse HR3F
Row by row

Every figure, with its dataset and its split.

Held-out numbers first. Where we are level or behind rather than ahead, the row says so.

Key detection

MIREX weighted - GiantSteps, 566 tracksabove the published figures
Mixterio, held-out, shipping app74.80
Rekordbox 7measured by us, same 566 tracks72.67
S-KEY paperarXiv 2501.1290772.1
Faraldo 2016Essentia edmm, published72.0
madmompublished71.0
Serato DJ Litemeasured by us, same 566 tracks70.04

74.80 held-out (73.78 full set, 66.08 exact), measured through the shipping analysis path at the 44.1 kHz the application decodes at, and agreeing with the reference 22.05 kHz call on 566 of 566 tracks. On tracks it was never tuned on, the shipping application scores above the strongest published academic figures. Rekordbox and Serato publish no number for this benchmark, so we ran both ourselves on the same 566 tracks with key tags stripped and kept the per-track results; our directly comparable full-set score is 73.78 and also leads them. Standalone key-tagging utilities are a different category of product and are not on this chart.

results files: key_product_conditions_full566_2026-09-10.json - key_product_conditions.py - key_giantsteps_shipped_2026-07-16.json - competitor_key_rekordbox_2026-07-16.json - competitor_key_serato_2026-07-16.json - competitor_key_bench.py

Beat tracking

F-measure - GTZAN, 992 trackslevel with the published figures
Mixterio, held-out, never trained on89.23
Beat This!Foscarin, ISMIR 202489.1
madmompublished88.5

We reproduce the published state of the art on a set our tracker never saw. Strong where DJs live (hip-hop 97.5), weakest on classical (65.9), which is the hard case for every system. We deliberately do not report our Ballroom score, because Ballroom sits in the reference model's training data and would measure memorisation.

results files: beats_gtzan_beatthis_2026-07-16.json - beats_extract_gtzan.py - beats_eval_gtzan.py

Downbeat tracking

F-measure - GTZAN, 992 trackslevel with the published figures
Mixterio, held-out, never trained on78.69
Beat This!Foscarin, ISMIR 202478.3
madmompublished75.6

Downbeats are the ones that matter for phrasing, and they are harder than beats. On top of the model we add a figure no benchmark asks for: how tightly the grid sits on the actual kick, a median of 8.5 ms across our library.

results files: beats_gtzan_beatthis_2026-07-16.json - beats_eval_gtzan.py

Tempo

ACC1, strict - GTZAN, 992 trackslevel with the published figures
Mixterio, raw grid, general material81.85

Right 81.85 percent of the time strictly, 93.48 percent if you allow double or half time. The 11.63 percent gap between those two is the octave problem, which is open for every published system, not just ours. Nobody publishes a comparable figure for this metric, so we draw no comparison.

results files: tempo_gtzan_2026-07-16.json - tempo_eval_gtzan.py

Structure

boundary HR3F - SALAMI IA, 153 held-out of a 361-track corpusbehind the published figures
Mixterio, held-out, coarse HR3F0.4025
Foote (2000)measured by us, same corpus0.472
CBM, TISMIR 2024measured by us, same corpus0.465
Spectral clusteringmeasured by us, same corpus0.383

This was our weakest pillar at 0.221, published rather than hidden. Measured at the operating point and decode rate the application runs: coarse HR.5F 0.1966 / HR3F 0.4025 held-out (n=153, 9.49 boundaries per track). We are level with our own Foote reproduction at the strict tolerance and 0.070 behind it at the lenient one. The denser research operating point scores 0.2341 / 0.4823 on the same tracks and is an ablation, not the product. We do not claim the higher published figures beaten: those come from different, easier slices of the benchmark, so we measured the field's own algorithms on our corpus instead of comparing rulers. Human annotators only agree with each other near 0.6 to 0.7 on this task, which is the honest ceiling.

results files: structure_product_conditions_held_2026-09-10.json - structure_product_conditions.py - structure_compare_held_2026-08-04.json - structure_eval.py

Measured 2026-07-16. Bars are drawn to each benchmark's own scale, so no gap is stretched. Every comparison is either a figure its makers published, or our own measurement on the same tracks, labelled as ours and backed by the per-track results.

How a third party checks this

Reproducible without our source.

Checking these numbers means running the app you install over a fixed public dataset and recomputing the table. How the analyzer works stays ours; the claim does not depend on that being open, because anyone with the same audio can run the same app and get the same numbers.

  1. 01

    Pinned dataset: exact source, version, and audio checksum.

  2. 02

    Declared split: what tuned, what was held out. The headline is always the held-out number.

  3. 03

    Written protocol, published alongside the harness.

  4. 04

    Per-track results: one row per track, prediction and score, regenerated whenever the analyzer changes.

  5. 05

    Verification drives the shipped product over public audio, so our source and models can stay closed while the claim stays checkable.

Coming with the first releaseThe test harness and the per-track results filesEach row above names the files behind it. They publish with the release, together with instructions that let anyone recompute the whole table from the same public audio.
Rules we hold ourselves to
  • The headline is always the held-out number, never the tuning half.
  • We do not report scores on data our models were trained on, even when they look excellent.
  • Comparisons prefer figures their authors published. Where a widely used tool publishes none, we measure it ourselves on the same pinned dataset, say so on the chart, and keep the per-track results.
  • Every figure on this site traces back to the benchmark document behind it. If it is not in there, it does not go on the site.

Get the results files when they go public.

The per-track files and the harness publish with the release. What the analyzer measures.

Nothing to buy today. Payments open at launch.