# Plan: Apply & Install `md/34_VST_LINUX.md` (Revised per `md/35_VST_PLAN.md` + `md/35.1_VST_FIX.md`) ## Context Summary - **Dockerfile** already has `libgl1`, `libasound2`, `libjack-jackd2-0`, `libfreetype6`, Xvfb, fluidsynth. Missing `libcurl4`. - `requirements.txt` has `mido` but missing `sf2utils`. - `app/core/vst_engine.py` has `PluginManager` with FluidSynth C-API-based `list_soundfont_instruments()`. - `app/core/render_engine.py` FluidSynth path hardcodes `program_select(0, fid, 0, 0)`. - `app/static/js/services/aiGateway.js` — `generate_multitrack_midi` tool has NO `soundfont_id`/`soundfont_bank`/`soundfont_program`. - `soundfontPlayer.js` — oscillator-based, uses `program` only for ADSR/osc-type selection; no `controllerChange()` or `programChange()` channel-state methods. - No `soundfont_catalog.json` generation or catalog API endpoint. - `docker-compose.yml` mounts `.` to `/app`. --- ## 5 Mandatory Refinements (from 35_VST_PLAN.md) 1. **Condensed Catalog for AI**: `get_condensed_catalog_summary()` → max 40–50 instruments categorized by GM groups (Piano, Organ, Guitar, Bass, Strings, Ensemble, Brass, Reed, Pipe, Synth Lead, Synth Pad, Drum Kit). Avoids token overflow. 2. **Non-Blocking Catalog Generation**: `generate_full_catalog()` runs on first API access, not on startup. Cache in memory; refresh on SF2 upload via background task. 3. **Dual Directory Scanning**: system `/opt/daw_engine/soundfonts/` AND user uploads `app/storage/uploads/soundfonts/`. 4. **DecentSampler CWD Fix**: `os.chdir()` to `.dspreset` parent directory before `load_preset()`, so relative `samples/*.wav` paths resolve. 5. **Robust Error Handling**: per-file `try/except` in `SoundFontInspector` — skip corrupted SF2 files with warning instead of crashing. --- ## Task A: SoundFont Inspection Engine (sf2utils) ### A1 — Add dependency - Add `sf2utils>=0.9.0` to `requirements.txt`. ### A2 — Create `app/core/soundfont_inspector.py` - `inspect_sf2_file(filepath)` — wrapped in `try/except` per Refinement 5. Returns `{soundfont_id, filename, total_instruments, instruments: [{bank, program, name, is_percussion}]}`. - `generate_full_catalog(output_json_path)` — scans system dir `/opt/daw_engine/soundfonts/` + user upload dir (Refinement 3). Writes `soundfont_catalog.json`. - `get_condensed_catalog_summary()` — returns categorized dict with ≤50 entries per Refinement 1. Groups instruments by GM category (Piano=0-7, Chromatic Perc=8-15, Organ=16-23, Guitar=24-31, Bass=32-39, Strings=40-47, Ensemble=48-55, Brass=56-63, Reed=64-71, Pipe=72-79, Synth Lead=80-89, Synth Pad=90-103, Drum Kit=128). - `invalidate_catalog_cache()` — resets in-memory cache; called after SF2 upload. ### A3 — API Endpoint `GET /api/v1/plugins/soundfonts/catalog` - In `app/api/v1/plugins.py`: - Response shape: `{ full_catalog: {...}, condensed_catalog: {...} }`. - Lazily generate on first call, cache in memory (Refinement 2). - `POST /upload-soundfont` success handler: calls `invalidate_catalog_cache()` + triggers a background task (FastAPI `BackgroundTasks`) to re-scan. Does NOT block HTTP response. ### A4 — JS API wrapper - In `app/static/js/services/api.js`, add `SonicAPI.getSoundfontCatalog()` → `GET /api/v1/plugins/soundfonts/catalog`. --- ## Task B: DecentSampler + Pianobook Support ### B1 — Dockerfile updates - Add `libcurl4` to `apt-get install`. - Pre-create `/opt/daw_engine/vst3/` and `/opt/daw_engine/samples/pianobook/` with `mkdir -p`. ### B2 — Host dirs - Create `vst_plugins/` and `samples/pianobook/` at repo root. Add both to `.gitignore`. ### B3 — DecentSamplerManager in `app/core/vst_engine.py` - `create_decent_sampler_instance(dspreset_path)`: - Resolve to absolute path with `os.path.abspath()`. - Save original CWD with `os.getcwd()`, then `os.chdir()` to `.dspreset` parent dir before `load_preset()` (Refinement 4). - Restore original CWD in `finally` block. - Return `VST3Plugin` instance ready for rendering. ### B4 — Wire into `app/core/render_engine.py` - If track selects a Pianobook instrument (e.g. `instrument_source: "pianobook"`), route through `DecentSamplerManager` instead of FluidSynth or synth fallback. - The Pianobook path uses `pedalboard.Pedalboard([vst])` with MIDI messages, same as other VST3 paths. --- ## Task C: AI Tool Schema & Prompt Injection ### C1 — Update `generate_multitrack_midi` tool in `aiGateway.js` - Add to `parameters.properties.tracks.items.properties`: - `soundfont_id`: `{ type: "string", description: "ID of the SoundFont file (e.g. 'generaluser_gs')" }` - `soundfont_bank`: `{ type: "integer", default: 0, description: "MIDI Bank. 0 = melodic, 128 = drums/percussion" }` - `soundfont_program`: `{ type: "integer", description: "MIDI Program number 0-127 from instrument catalog" }` - Add all 3 to `required` array. ### C2 — Inject condensed catalog into system instruction - Modify `buildUserMessage()` in `aiGateway.js`: - When `systemInstruction` is empty and global `window.__soundfontCatalog` exists, prepend a `system` message block containing the condensed catalog text. - Format: one line per GM category with bank/program examples. - **Enforce bank rule** (per 35.1_VST_FIX.md §3.B): Add explicit instruction — _`soundfont_bank: 0` for all melodic instruments, `soundfont_bank: 128` for Drum Kits._ ### C3 — Fetch catalog on frontend startup - In `app/static/js/app.jsx`, after auth check, call `SonicAPI.getSoundfontCatalog()`. - Store result in `window.__soundfontCatalog = { condensed_catalog, full_catalog }`. - Re-fetch after any SF2 upload succeeds. --- ## Task D: Server Render — Program Change & Channel Mapping ### D1 — Read bank/program from track metadata - In `render_engine.py` MIDI rendering block, extract `soundfont_bank` and `soundfont_program` from track dict. - Default: bank=0, program=0. ### D2 — MIDI channel routing + FluidSynth update - **Channel rules** (per 35.1_VST_FIX.md §5.B): - `bank == 128` or track has `is_percussion: true` → `midi_channel = 9` (GM channel 10, percussion). - Otherwise → assign channels sequentially from 0–8, one per unique percussion-group track. - Replace hardcoded `fl.program_select(0, fid, 0, 0)` with: ```python midi_channel = 9 if (bank == 128 or track.get("is_percussion")) else channel_counter fl.program_select(midi_channel, fid, bank, prog) ``` - All note_on/note_off events for that track must use the same `midi_channel`. ### D3 — VST3/Pedalboard path: CC + PC insertion - Extend `midi_events_to_messages()` or add a wrapper that inserts two MIDI messages at sample_offset=0 before note messages: - `MidiMessage(control_change=0, value=bank, sample_offset=0)` — CONTROL_CHANGE CC 0 (Bank Select MSB) - `MidiMessage(program_change=program, sample_offset=0)` — PROGRAM_CHANGE - These are prepended to the message list before `Pedalboard([vst])` processes the buffer. --- ## Task E: Client SoundFont Player — Program Change & Channel Allocation ### E1 — Add channel-state tracking to `soundfontPlayer.js` - Add internal `_channels` array (size 16), each entry: `{ bank: 0, program: 0 }`. - `controllerChange(channel, controller, value)`: - If `controller === 0` (Bank Select MSB), store `bank` for that channel. - `programChange(channel, program)`: - Store `program` for that channel. - Modify `playNote()` to accept an optional `channel` parameter and use the stored bank/program for ADSR/osc-type selection. ### E2 — Add `applyAITrackInstrument(trackId, bank, program)` - New function in `soundfontPlayer.js`: - Determine MIDI channel: `bank === 128 || isPercussion ? 9 : track_index % 9`. - Call `controllerChange(channel, 0, bank)`. - Call `programChange(channel, program)`. - Called from `app.jsx` after AI returns `generate_multitrack_midi` with track instrument data. ### E3 — Wire into post-AI pipeline in `app.jsx` - In the DAW command dispatch loop (around line 12853), after processing `generate_multitrack_midi` function call: - For each returned track with `soundfont_bank`/`soundfont_program`, call `applyAITrackInstrument()`. - Log the action to `aiActionLog`. --- ## Task F: Background Cache Invalidation on Upload ### F1 — Update `POST /upload-soundfont` in `plugins.py` - After saving the uploaded SF2 file: 1. Call `SoundFontInspector.invalidate_catalog_cache()`. 2. Use FastAPI `BackgroundTasks` to queue a re-scan: `background_tasks.add_task(generate_full_catalog)`. 3. Return HTTP 200 immediately (not block on scan). ### F2 — Frontend catalog re-fetch after upload - In `app.jsx` upload handler, after `SonicAPI.uploadSoundFont()` succeeds, call `SonicAPI.getSoundfontCatalog()` and update `window.__soundfontCatalog`. --- ## Task G: Validation ### G1 — Catalog API - `GET /api/v1/plugins/soundfonts/catalog` → valid JSON with `{ full_catalog: {...}, condensed_catalog: {...} }`. - Condensed catalog contains ≤50 entries, grouped by GM category. ### G2 — AI generation - Input: _"Compose 8 bars of Brass horns and a drum kit"_ - Verify AI returns `generate_multitrack_midi` call with: - Brass track: `program: 56`, `bank: 0`, `soundfont_id: "generaluser_gs"`. - Drums track: `program: 0`, `bank: 128`, `soundfont_id: "generaluser_gs"`. ### G3 — Client instrument switching - After AI response, verify `applyAITrackInstrument` is called with correct bank/program per track. - Verify MIDI channel allocation: melodic → ch0-8, drums → ch9. - Verify `controllerChange(CC0)` + `programChange()` dispatched per channel. ### G4 — Server render - Export WAV, verify correct Brass horn and Drum sound. - For FluidSynth path: confirm `program_select` uses correct channel, bank, program. - For VST3 path: confirm CC0 + PC inserted before notes. ### G5 — Upload cache invalidation - Upload a new `.sf2` file → verify `catalog` endpoint updates without manual restart. - Upload a corrupted `.sf2` file → verify it is skipped gracefully (Refinement 5). ### G6 — Regression - `pytest tests/` passes with no regressions. --- ## Implementation Order 1. **A1–A4** (sf2utils + soundfont_inspector + catalog API + JS wrapper) — foundational. 2. **C1–C3** (AI tool schema + condensed prompt injection + startup fetch) — depends on A3/A4. 3. **D1–D3** (server render program change + channel mapping + CC/PC insertion) — depends on C1 for field names. 4. **F1–F2** (background cache invalidation on upload) — depends on A3. 5. **E1–E3** (client program change + channel allocation + post-AI wiring) — independent of D, but shares channel routing logic. 6. **B1–B4** (DecentSampler) — last, requires manual VST3 binary download. 7. **G1–G6** (validation).