# TECHNICAL PLAN: DECENT SAMPLER + PIANOBOOK INSTALLATION AND SOUNDFONT MAPPING EXTRACTION FOR AI AGENT --- ## 1. Executive Overview The system needs to fulfill two core requirements: * **Integrate DecentSampler & Pianobook on Linux Server:** * Install `DecentSampler.vst3` (Linux 64-bit) into the Docker Server environment. * Structure the Pianobook sample library directories (`.dspreset` + `.wav` files). * Integrate Preset Loading into Python `pedalboard` for offline rendering steps. * **Resolve the AI Agent's "Instrument Information Blindness" regarding SoundFont (`.sf2`):** * *Current State:* The AI Agent and the system only load raw `.sf2` files without knowing what instruments are contained within (which Bank, which Program/Patch number, or what the instrument names are). * *Solution:* * Build a **SoundFont Inspection Engine** using Python (`sf2utils`) to scan `.sf2` files upon upload/scan and extract the instrument catalog table (Bank, Program/Preset ID, Instrument Name). * Export a Catalog table (`soundfont_catalog.json`) and pass this context to the AI Agent. * Update the AI Tool Schema so the AI accurately passes `soundfont_id`, `bank`, and `program` (MIDI Program Change) when creating a Track. --- ## 2. DecentSampler + Pianobook Installation Plan (Server Backend) ### 2.1 Installing DecentSampler Linux Native VST3 in Docker 1. **Download DecentSampler Linux VST3:** * Download the official Linux 64-bit build from the DecentSampler website (`DecentSampler_Linux_x64.tar.gz` or `.vst3` file). 2. **Server Directory Structure:** ```text /opt/daw_engine/ ├── vst3/ │ └── DecentSampler.vst3/ <-- Native Linux VST3 Binary ├── soundfonts/ │ ├── GeneralUser_GS.sf2 │ └── SGM-V2.01.sf2 └── samples/ └── pianobook/ <-- Pianobook Sample Libraries ├── salamander_grand_piano/ │ ├── salamander_piano.dspreset │ └── samples/ (*.wav) └── acoustic_guitar/ ├── guitar.dspreset └── samples/ (*.wav) ``` 3. **Additions to `Dockerfile`:** ```dockerfile # Install audio system dependencies RUN apt-get update && apt-get install -y \ libgl1-mesa-glx \ libfreetype6 \ libcurl4 \ && rm -rf /var/lib/apt/lists/* # Copy DecentSampler VST3 to Server COPY ./vst_plugins/DecentSampler.vst3 /opt/daw_engine/vst3/DecentSampler.vst3 ``` ### 2.2 Integrating DecentSampler into the Python Engine (`app/core/vst_engine.py`) The `pedalboard` library supports loading VST3 plugins and preset files for DecentSampler: ```python import os from pedalboard import VST3Plugin class DecentSamplerManager: def __init__(self, vst_path="/opt/daw_engine/vst3/DecentSampler.vst3"): self.vst_path = vst_path def create_decent_sampler_instance(self, dspreset_path: str) -> VST3Plugin: """ Instantiates VST3 DecentSampler and loads the Pianobook sample preset (.dspreset) file. """ if not os.path.exists(self.vst_path): raise FileNotFoundError(f"DecentSampler VST3 not found at {self.vst_path}") plugin = VST3Plugin(self.vst_path) # Load the Pianobook preset file into DecentSampler VST3 if os.path.exists(dspreset_path): plugin.load_preset(dspreset_path) return plugin ``` --- ## 3. Designing the SoundFont Inspection Engine (Bank/Program Extraction) Every SoundFont (`.sf2`) file is a collection of Presets (or Programs). For the AI Agent to know what sound presets exist inside the `.sf2` file, the Backend must scan and parse the `.sf2` file. ### 3.1 Adding Python SoundFont Inspection Libraries Add to `requirements.txt`: ```text sf2utils>=0.9.0 mido>=1.3.0 ``` ### 3.2 Building the SoundFont Metadata Inspection Service (`app/core/soundfont_inspector.py`) ```python import os import json from sf2utils.sf2parse import Sf2File class SoundFontInspector: def __init__(self, sf_dir="/opt/daw_engine/soundfonts"): self.sf_dir = sf_dir def inspect_sf2_file(self, filepath: str) -> dict: """ Parses an .sf2 file and returns a complete instrument catalog (Bank, Program, Instrument Name). """ if not os.path.exists(filepath): return {} sf_name = os.path.basename(filepath) sf_id = os.path.splitext(sf_name)[0].lower() instruments = [] with open(filepath, 'rb') as f: sf2 = Sf2File(f) for preset in sf2.presets: # Ignore EOP (End of Header) preset if preset.name.strip() == "EOP" or (preset.bank == 128 and preset.preset == 127): continue instruments.append({ "bank": preset.bank, # Bank number (0 = General MIDI Standard, 128 = Percussion/Drums) "program": preset.preset, # Program/Patch number (0-127) "name": preset.name.strip(), # Instrument name (e.g. "Stereo Grand", "Violin", "Brass Section") "is_percussion": (preset.bank == 128) }) return { "soundfont_id": sf_id, "filename": sf_name, "total_instruments": len(instruments), "instruments": instruments } def generate_full_catalog(self, output_json_path="/opt/daw_engine/soundfont_catalog.json"): """ Scans all .sf2 files in the directory and builds a Catalog JSON for the AI Agent. """ catalog = {} for root, dirs, files in os.walk(self.sf_dir): for file in files: if file.endswith(('.sf2', '.SF2')): full_path = os.path.join(root, file) sf_info = self.inspect_sf2_file(full_path) catalog[sf_info["soundfont_id"]] = sf_info with open(output_json_path, 'w', encoding='utf-8') as f: json.dump(catalog, f, ensure_ascii=False, indent=2) return catalog ``` ### 3.3 Catalog File Output Structure (`soundfont_catalog.json`) This JSON file serves as an Instrument Dictionary for the AI Agent: ```json { "generaluser_gs": { "soundfont_id": "generaluser_gs", "filename": "GeneralUser_GS.sf2", "total_instruments": 128, "instruments": [ { "bank": 0, "program": 0, "name": "Stereo Grand Piano", "is_percussion": false }, { "bank": 0, "program": 19, "name": "Church Organ", "is_percussion": false }, { "bank": 0, "program": 40, "name": "Violin Ensemble", "is_percussion": false }, { "bank": 0, "program": 56, "name": "Trumpet", "is_percussion": false }, { "bank": 128, "program": 0, "name": "Standard Drum Kit", "is_percussion": true } ] } } ``` --- ## 4. Guiding the AI Agent in Instrument Selection & Accurate Note Loading When the user types: *"Create a Piano track and a Strings section track for 8 bars"*, the AI Agent needs to know precisely which soundfont, bank, and program numbers to assign to the Tracks. ### 4.1 Updating the AI Tool Function Spec (`tools_spec.json`) Add `soundfont_bank` and `soundfont_program` fields to the Tool Schema sent to the AI: ```json { "type": "function", "function": { "name": "generate_multitrack_midi", "description": "Generates multi-track MIDI data along with appropriate SoundFont Program configurations.", "parameters": { "type": "object", "properties": { "tracks": { "type": "array", "items": { "type": "object", "properties": { "track_name": { "type": "string" }, "instrument_type": { "type": "string", "enum": ["PIANO", "STRINGS", "BRASS", "SYNTH", "DRUMS"] }, "soundfont_id": { "type": "string", "description": "ID of the SoundFont to use (e.g. 'generaluser_gs')" }, "soundfont_bank": { "type": "integer", "default": 0, "description": "MIDI Bank code (0 for melodic instruments, 128 for Drums)" }, "soundfont_program": { "type": "integer", "description": "MIDI Program Number (0-127) corresponding to the instrument name in the Catalog" }, "notes": { "type": "array", "items": { "type": "object" } } }, "required": ["track_name", "soundfont_id", "soundfont_bank", "soundfont_program", "notes"] } } } } } } ``` ### 4.2 Injecting the Catalog into the AI Agent's Context Prompt (Prompt Template) Before sending the user query to the LLM, the system reads `soundfont_catalog.json` and injects a condensed catalog table into the System Instruction: ```python # System Context Prompt Injector def build_ai_system_instruction(catalog_data: dict) -> str: sf_summary = [] for sf_id, sf_info in catalog_data.items(): sf_summary.append(f"SoundFont ID: '{sf_id}' (File: {sf_info['filename']}):") for inst in sf_info['instruments'][:20]: # Inject primary instrument lists sf_summary.append( f" - [{inst['name']}]: bank={inst['bank']}, program={inst['program']}" ) catalog_context = "\n".join(sf_summary) system_instruction = f""" You are an AI Copilot for a DAW. Below is the Catalog of available SoundFonts on the system: {catalog_context} MANDATORY RULES WHEN CREATING TRACKS: 1. When creating any track, you MUST look up the catalog above and fill in the correct `soundfont_id`, `soundfont_bank`, and `soundfont_program`. 2. Example: If the user requests "Piano", select soundfont_id="generaluser_gs", soundfont_bank=0, soundfont_program=0 ("Stereo Grand Piano"). 3. If the user requests "Violin/Strings", select soundfont_bank=0, soundfont_program=40 ("Violin Ensemble"). 4. If the user requests "Drums", select soundfont_bank=128, soundfont_program=0 ("Standard Drum Kit"). """ return system_instruction ``` ### 4.3 Applying Program Change on Client & Server Render Layers #### A. Client Browser Side (FluidSynth Wasm / SoundFont Player) Upon receiving JSON from the AI, the Frontend invokes the Bank and Program selection function to trigger the correct sound: ```javascript // Client-side Javascript (spessasynth / fluidsynth.wasm) function applyAITrackInstrument(trackId, soundfontBank, soundfontProgram) { const channel = getTrackMIDIChannel(trackId); // Send MIDI Bank Select (CC 0) synthInstance.controllerChange(channel, 0, soundfontBank); // Send MIDI Program Change synthInstance.programChange(channel, soundfontProgram); } ``` #### B. Server Offline Render Side (`app/core/render_engine.py`) When rendering to a WAV file, Python inserts a MIDI Program Change event ahead of the Track's note sequence: ```python import mido def create_midi_track_with_program(notes_data, bank=0, program=0): midi_track = mido.MidiTrack() # 1. Insert Bank Select (Control Change 0) midi_track.append(mido.Message('control_change', channel=0, control=0, value=bank, time=0)) # 2. Insert Program Change (Instrument Sound Selection) midi_track.append(mido.Message('program_change', channel=0, program=program, time=0)) # 3. Insert MIDI notes generated by AI for note in notes_data: start_tick = int(note['start_beat'] * 480) # 480 ticks per beat dur_tick = int(note['duration_beats'] * 480) pitch = int(note['pitch']) vel = int(note['velocity'] * 127) midi_track.append(mido.Message('note_on', note=pitch, velocity=vel, time=start_tick)) midi_track.append(mido.Message('note_off', note=pitch, velocity=0, time=dur_tick)) return midi_track ``` --- ## 5. Action Checklist * [ ] **Step 1:** Download Linux 64-bit `DecentSampler.vst3` and copy it into `/opt/daw_engine/vst3/`. * [ ] **Step 2:** Download Pianobook sound libraries (e.g. Salamander Grand Piano) and extract them to `/opt/daw_engine/samples/pianobook/`. * [ ] **Step 3:** Add `sf2utils` to `requirements.txt` and install it in Docker. * [ ] **Step 4:** Create `app/core/soundfont_inspector.py` to automatically scan all `.sf2` files in the project and generate `soundfont_catalog.json`. * [ ] **Step 5:** Build API Endpoint `GET /api/v1/plugins/soundfonts/catalog` returning the extracted instrument catalog. * [ ] **Step 6:** Update the Prompt Template and AI Tool Schema to support `soundfont_bank` & `soundfont_program`. * [ ] **Step 7:** End-to-End Verification: Type Prompt *"Generate 8 bars of Brass horn music"* $\rightarrow$ AI reads Catalog and selects Program `56` $\rightarrow$ Web Audio & Server Render produce the correct Brass horn instrument sound.