About the dataset

The story

Live music models need more data, especially stem-separated audio with real improvisation. A lot of jazz datasets out there are either not original recordings, lean on automated transcription, or analyze music someone else already put out. I also already had a band. Most of us have played together since high school, with a few friends joining later. At the start of a college summer we were going to jam anyway, so we decided to record and put the music online.

We did not rent a studio. The setup was a home one, and pretty rough, and for the asynchronous material we mostly tracked in our own rooms. Synchronous sessions were harder mainly because of scheduling: getting everyone free at the same time rarely worked, which is a big reason the asynchronous protocol is in the corpus. It let us keep adding songs on our own time. Drums went first so everyone else had something to play against. On a given part the worse take is usually the first one, often a cold read, and the better take is the one after the form feels familiar. For synchronous recording the pain was less about timing and more about organizing the sessions and not wearing everyone out.

The repertoire is basically our old high school gig book: standards we already knew from lead sheets. A lot of the tracking was cold reading. Who could show up that day was the combo. Saxophones split between alto and tenor across the catalog, and bass was whoever was free. I did the annotation and processing pipeline. Everyone played, me included. Recording took much longer than labeling; once the audio was in, the lane annotations were maybe a week or two. One thing about asynchronous takes that is easy to miss: a worse take is not always a clean leftover. Later instruments always sit on the better takes of whoever recorded earlier.

I care about musician-in-the-loop work in AI music, and I want this release to be honest about how it was made. We were going to play either way, and the research needed original stems with improvisation. Making the corpus was genuinely fun in places and also a long grind.

Thank you to Tornike Karchkhadze for early advice on how to actually record for this project, including audio engineering. Thank you to Zachary Novack for early talks about useful data to collect. Thank you to Landon Andrizzi for advice on recording drums on a budget. Thanks also to my collaborators at MIT who talked through recording choices and next steps while this was coming together.

Phillip Long

Two protocols on purpose

Asynchronous Per-instrument sessions

Musicians record separately. Drums go down first to establish the grid; each instrument has two takes assigned to better or worse tiers. Mixtures are loudness-normalized composites. Annotations are authored on the better mixture and shared with the worse tier via the same bar grid.

Synchronous Full band together

The ensemble records two whole-band takes. Better/worse select entire performances (not per stem). Close-mic bleed is expected; the public release includes original stems plus debleeded derived audio. Live timing varies — use the annotated bar grid, not a fixed BPM.

Annotations

Lead sheets seed rough bars, chords, sections, and soloists. A browser lane annotator aligns the bar grid to audio, then edits chords, sections, and soloists in metrical coordinates. Public CSVs ship under each take’s annotations/ folder with both metrical and absolute-time fields.

Website vs download

People

The combo that recorded JazzSAMBA is on The Band. Coauthors from the University of California, San Diego and the Massachusetts Institute of Technology appear on the home author list.

Catalog statistics

Summary charts from our song catalog. Each bar shows the full count for a category, stacked by recording protocol — Asynchronous and Synchronous. Hover a segment for exact counts.

Recorded Synchronously?: Share of corpus songs recorded synchronously (full band together) versus asynchronously (per-instrument sessions).
Hours per Instrument: Active hours from section musician annotations (both take tiers; silence while sitting out excluded), stacked by recording protocol.
Genre: Distribution of musical genres in the JazzSAMBA repertoire.
Form: Distribution of song forms (for example AABA and ABAC).
Year: Distribution of composition or standard publication years.
Tempo (BPM): Histogram of annotated tempo in beats per minute.
Tempo Text: Counts by tempo marking text (for example Medium Swing or Ballad).
Is Swung?: Share of songs marked as swung versus straight rhythm, stacked by recording protocol.
Horn Lead: Breakdown of songs with trumpet lead, saxophone lead, or both horns leading, stacked by recording protocol.
Key Signature: Distribution of key signatures across the catalog.
Time Signature: Distribution of time signatures across the catalog.