YuE2 in your browser

Lyrics → full song with m-a-p/YuE2-3B, running locally on WebGPU
checking WebGPU…
One-click
Length
0:40
Plan
Steps
Seed
Model: 2.2 GB download, cached after first load
1 · score
2 · audio tokens
3 · latents
4 · decode

YuE2 writes an ABC score, then 25 audio tokens per second, then refines them into latents with a flow-matching pass (12 steps by default, 32 matches the reference) and decodes 48 kHz stereo. Everything runs in this tab: 4-bit language model, fp16 audio decoder. A 40 s clip needs about 1,000 tokens.

How it's fast

A hand-written decoder. YuE2 writes the song one token at a time: 25 audio tokens per second of music, plus the ABC score before that. Normally each token goes through ONNX Runtime, then the GPU result is copied back so JavaScript can pick the next token, and the loop starts again. Here, generation runs on a custom WebGPU decoder (dec.js) built from about 175 small WGSL kernels per token. The 4-bit weight matmuls unpack the weights in registers and use GPU subgroups, with RMSNorm folded into the matmul that follows it. The gate and up projections share one pass. Attention over the cached keys and values is split into chunks that run in parallel (flash-decoding).

Sampling stays on the GPU. The repetition penalty, top-k, top-p, temperature and the next token's embedding lookup are kernels too, so the CPU never waits between tokens. Many tokens are queued in a single submit, and their ids are read back in batches only to update the progress bar. This roughly doubles the token rate, from about 45 to 80–90 tok/s on an M4 Max. The logits match the ONNX path closely (same top token at every step tested), and browsers without the needed GPU features fall back to ONNX Runtime automatically (?dec=ort forces it).

Fewer, better-placed flow steps. After the tokens, a flow-matching transformer (another 28-layer pass over the whole song) turns them into audio latents by solving an ODE. The original sampler takes 32 midpoint steps, each needing two model calls. The default here is 12 steps with the timesteps shifted toward the noisy end, where the trajectory bends most. That cuts the calls from 32 to 24 compared with the previous 16-step default, and it lands slightly closer to the 32-step reference. Euler, Heun, Adams-Bashforth, DPM++ 2M and a plain 8- or 10-step schedule were all measured against the same reference and drifted further, so they aren't the default (8 is still a button).

A hand-written flow transformer too. The flow model's forward pass is also custom WGSL (nar.js), run over every latent frame of the song at once. Its 4-bit matmuls are tiled: each workgroup unpacks a block of weights into shared memory once, reuses it for 128 frames and accumulates in f32 registers. The residual add and the SiLU gate are fused into the matmul output. Attention is a flash-attention kernel that reads the frozen prompt cache and the frames' own keys in place, without building the concatenated tensors. The ODE state stays on the GPU, so all 24 model calls are queued back to back and the latent is read back once at the end. On an M4 Max (on battery) this runs about 3 TFLOPS against ONNX Runtime's 1.6, so the flow stage for a 20-second song drops from 23 s to 12 s. Velocities match the ONNX path to within about 1%. The 12-step result is 0.73% from this kernel's own 32-step reference (ONNX: 0.83%). ?nar=ort forces ONNX Runtime, which is also the automatic fallback.

Nothing leaves the GPU that doesn't have to. The key/value cache is a set of fixed GPU buffers, shared by the prefill pass, the decoder and the flow transformer, so the prompt is never re-encoded. The VAE decodes in overlapping 128-frame windows and streams each one to the player as soon as it's ready. Weights download in parallel and are kept in the browser's Cache Storage, so a reload starts in about 2 seconds.