Mixterio

01 The analyzer and the library

Nothing plays until it has been read.

Every track gets listened to once, properly. Beats, key, loudness, the shape of the arrangement, and the drums, bass, vocals and melody pulled apart. What comes out is not a guess about a record. It is a grid, a key, a shape and a set of numbers we can be held to.

The published accuracy for this pass

One analysis passdiagram
Grid
every beat, every downbeat, the 4-bar phrase
Key
Camelot, with a confidence
Tempo
BPM, and whether it holds
Shape
intro / build / drop / breakdown / outro
Stems
drums, bass, vocals, melody, separated

A drawn diagram of one analysed track, not a screenshot. The sections are the labels the analyzer uses; the fields are what it stores on your disk.

01Per track

Ten things, once, then never again.

It listens to each file once and keeps what it learns on your disk. That is the slow part, and it is the reason everything after it is fast and exact. Tags are ignored; the audio is measured.

Tracks whose beat grid cannot be trusted are held back rather than quietly used, because a bad grid is how a set ends up warped to half tempo in front of a room. You get the list, with the reason for each one.

None of this is asked of you twice. Analyse a folder once and every set you ever build from it starts from the same measurements.

What one analysis pass produces and stores.
GridEvery beat and downbeat, plus the 4-bar phrase grid you count in
TempoBPM, and how steady it stays. A track whose grid cannot be trusted is held back rather than used
KeyKey and how confident we are, in Camelot wheel notation
LevelLoudness, so no track arrives louder or quieter than the one it follows
EnergyAn energy curve across the whole track, second by second
ShapeLabelled sections: intro, verse, breakdown, build, drop, chorus, outro
Per sectionHow much bass, vocal and drum each section carries, and how busy it is
TimbreA fingerprint of the track's texture, so you can look for something that sounds like this
StemsDrums, bass, vocals and melody, pulled apart so a blend can hand over one at a time
FlagsWhether it has vocals, whether the tempo holds, whether the grid can be trusted
02Measured, not asserted

Every claim on this page has a number and a results file.

We run the same public benchmarks the research field uses, hold half the data out, and publish what comes back. Even when the number is against us.

Key detection74.80MIREX weighted

72.67 Rekordbox 7measured by us, same 566 tracks

held-out, shipping app
Beat tracking89.23F-measure

89.1 Beat This!Foscarin, ISMIR 2024

held-out, never trained on
Downbeat tracking78.69F-measure

78.3 Beat This!Foscarin, ISMIR 2024

held-out, never trained on
Tempo81.85ACC1, strict

no published comparisonnobody publishes a figure for this metric, so we draw none

raw grid, general material
Structure0.4025boundary HR3F

0.472 Foote (2000)measured by us, same corpus

held-out, coarse HR3F

Grid quality, per track

Public benchmarks score whether the beats are found. We also check how tightly the grid sits on the actual kick, a median of 8.5 ms across our library. A loose grid is what makes a technically correct mix still feel wrong.

Strong where DJs live

On the public beat set our tracker scores 97.5 on hip-hop and 65.9 on classical. We publish the per-genre split rather than one flattering average, because you should know which of your folders it is good at.

Held out on purpose

Half of every benchmark is never used for tuning, and we do not report scores on data our models were trained on. That rule costs us a much better looking number and we keep it anyway.

03The library

A library that knows which pairs work.

Two tracks are not called compatible because they look alike on paper. The engine builds the actual blend as audio, in advance, one instrument at a time, and grades what comes out.

Pairs that clear the floor become the ones it reaches for. Pairs that fail are remembered as failures, so the same bad idea is not tried on you twice. That is what you browse: not tracks that sound alike, but joins that have already been listened to by an instrument.

While idle, a background supervisor measures more candidate pairings across your library at lowest priority, expanding the map without draining battery or interrupting foreground work. Search still works the way you expect, by tempo, key, energy and sound.

  • 0.85

    Measured-map floor

    A pairing only counts as proven once its transition has been built as audio and scored at or above this. If the engine ever has to reach below the floor, it tells you on screen instead of playing it quietly.

  • 3.0 dB

    Established handoff

    The incoming track has to arrive into a section that is already running. If it would jump in louder than this, the blend is stretched to the next phrase or a different track is chosen.

  • 6%

    Tempo lock

    Tracks too far from the tempo of the night are ruled out before anything else is considered, so nothing gets stretched into a chipmunk.

  • 12%

    Legal warp band

    A hard ceiling on how far any track can be pulled from its own natural tempo, however much the rest of the set would like it to fit.

Measured map, schematicrows out / cols in
  • strong
  • passed the floor
  • tried, rejected
  • not tried yet

A drawn diagram of a pair matrix, not real data.

04The row we rebuilt

Structure was our weakest number.

Section detection was rebuilt in August 2026. Measured at the operating point and decode rate the application runs, it scores coarse HR.5F 0.1966 / HR3F 0.4025 held-out (n=153), up from 0.221. We are level with our own Foote reproduction at the strict tolerance and 0.070 behind it at the lenient one. The denser 0.2341 / 0.4823 is an ablation, not the product.

Higher published figures exist on easier slices of that benchmark, so we measured the field's own algorithms on our corpus instead of comparing rulers. At the strict tolerance we hold level with our own Foote reproduction, and at the lenient tolerance we are 0.070 behind. The per-track results file shows every call either way.

The score was not entirely a failure, either. That benchmark rewards finding every single boundary in a recording, while we deliberately look only for the handful of points a DJ can actually blend on. So we also measure it the way the app uses it, and publish the limits of that too.

Measuring structure the way a DJ uses it
Anchor precision, phrase tolerance0.588when we place an anchor, it lands on a real functional boundary
Major-boundary recall, phrase tolerance0.372the three landmarks we care most about are still missed most of the time
Boundaries exactly on a downbeat39.8%against 25.0% by chance, measured across the library

This is a reframe, not a victory: the external sample is n=24 and off-genre, and the in-library signal is modest. The SALAMI number above stays on the board unchanged.

05Next

The analysis is the half you never see.

The other half is what the engine does with it: a set you can read, edit in the workbench, and export with cue points to your DJ gear - or a performance that holds one tempo for hours.

Composing and performing / The pipeline end to end

Waitlist

Get the per-track results when they go public.

The test harness and the per-track files publish with the first release. Leave your email and you get the link the day they are up.

Nothing to buy today. Payments open at launch.