The AI Front Page

Reading signals from this article are folded back into your front page ranking on this device.

Research/arXiv AI/ML/August 4, 2026 at 5:56 PM

arXiv paper: Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

A new arXiv AI paper by Junhao Chen, Mingjin Chen, and Jingjia Mao, and 12 more studies Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation.

Research / arXiv AI/ML
Source

Follow arXiv AI/ML to make it a durable For You signal.

A systematic comparison of music tokenization strategies shows that *how* music is represented matters far more for distributional fidelity than scaling model size. Introducing **Agogic**, a performance-timed tokenizer with 10 ms onset quantization, per-note velocity, and multi-track support (609-symbol vocabulary), researchers fixed all other variables—pretrained Qwen3.5 backbones from 0.8B to 27B, identical data and decoding—and swapped only the representation across seven tokenizations. The result: a 0.8B model using the performance-timed representation achieved a Fréchet Music Distance (FMD) of 159, beating beat‑grid tokenizers that scored 272–286 even when scaled to 27B parameters. The gap persisted (67–129 FMD points) after artificially snapping the performance‑timed onsets to beat‑grid resolution, indicating a fundamental advantage of the class rather than merely finer timing. A lightweight decode‑time constraint further doubled instrument‑F1 (0.28 → 0.60) and Correct‑Key (0.16 → 0.35) without hurting distributional scores. The team also released all checkpoints, training corpora (including the largest captioned music dataset at 6.25M examples), and an imprinting diagnostic,

arXiv paper: Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation | The AI Front Page