# Vocal & Music Splitter dependencies All audio processing is local to the browser. There is no upload or inference endpoint. - ONNX Runtime Web 1.22.0, MIT license (`ORT-LICENSE.txt`, `ORT-ThirdPartyNotices.txt`). Unmodified runtime files from https://registry.npmjs.org/onnxruntime-web/-/onnxruntime-web-1.22.0.tgz Source: https://github.com/microsoft/onnxruntime/tree/v1.22.0 - HTDemucs FT vocals specialist, fp16 weights. MIT, per the model publisher's card; original Demucs license retained in `DEMUCS-LICENSE.txt`. Pinned model: https://huggingface.co/StemSplitio/htdemucs-ft-vocals-onnx/tree/2ef0d757d3e226d0da85fb8c71514f464fcabdd0 File: `htdemucs_ft_vocals_fp16weights.onnx` (165,612,636 bytes). Original SHA-256: `0cbe651f535415c9d26a7bb614f7d322dd5a080fa0298f2e50f478030a994dce`. Original model source: https://github.com/facebookresearch/demucs - ONNX model export and reference inference: https://github.com/StemSplit/demucs-onnx/tree/85db5c80aba33f0f2bdf88034a4be6539feec85b . See `STEMSPLIT-LICENSE.txt`. The ONNX file is stored as four sequential binary parts below GitHub's single-file limit. The worker verifies each part against `model.json` before joining them. Deploy **all four** `vocals-model-*.bin` files and all runtime files alongside this document; there is no external model download service at runtime. Total additional assets are approximately 177 MB. The app decodes/resamples to 44.1 kHz stereo, then runs the vocals specialist in 7.8-second overlapping chunks in a dedicated worker. It uses only the vocal output row. Instrumental audio is the source mix minus the estimated vocals. When necessary, the same attenuation is applied to both outputs to avoid clipping. Results are stereo 16-bit PCM WAV, and the worker is terminated after completion, cancellation, error, or page exit. Graph optimization and the CPU memory arena are disabled to avoid excessive memory use during browser session creation. Up to four WASM CPU threads are used (half the available cores, capped at four) when cross-origin isolation and shared memory are available; other browsers fall back to one thread. The splitter page sends COOP `same-origin` and COEP `require-corp`; its worker/runtime assets carry COEP and same-origin resource headers. Preserve these headers when deploying behind a proxy. WebGPU is not required. A local 9-second, two-chunk benchmark on an 8-core browser measured 28.5 seconds of inference with one thread and 11.0 seconds with four; this is not a guarantee for other devices or longer recordings. Download progress is streamed, and remaining processing time is estimated from completed chunks. Duration is limited to 30 minutes (350 MB input). Decoded channels are transferred to the worker without retaining a second recording in the UI, and overlap ramps sum to one without a recording-length weights array. Retries decode the original local file again. Speed and separation quality depend on the device and recording. The model can leave bleed or other artifacts; the UI provides previews rather than promising perfect separation. Results appear as separate vocal and music waveform layers with independent playback and downloads. Shared zoom and synchronized horizontal scrolling keep their time axes aligned. Waveform peaks are sampled from the completed WAV files without decoding another full copy of either stem. Verified model parts are stored explicitly in Cache Storage under `rapidcaption-vocal-model-v1`, keyed by each part’s SHA-256. Subsequent runs read and verify saved parts without fetching their binary URLs. Missing or corrupt parts are downloaded again; complete parts survive cancellation. The small manifest is fetched each run to select the correct version. UI progress distinguishes saved-model loading from downloading. Storage failure does not block separation. Clearing site data, browser eviction, private browsing, or changing the site origin can require another download. No user audio is stored in this model cache.