Files

7.0 KiB
Raw Permalink Blame History

ASSESSMENT REPORT & OPTIMIZATION PLAN: LINUX VST3 & SOUNDFONT MAPPING ENGINE


1. Executive Review

Your plan closely aligns with the current codebase status and correctly identifies key bottlenecks (such as missing libcurl4, missing sf2utils, hardcoded program_select(0, fid, 0, 0) calls in render_engine.py, and missing schema fields in the AI Tool Schema).

However, to guarantee stable system operation within the Docker Linux environment and prevent failures during End-to-End execution, the plan requires the 5 critical technical refinements detailed below.


2. 5 Mandatory Technical Refinements

💡 Refinement 1: AI Prompt Context Size Control (Avoiding Token Overflow)

  • Problem in Previous Plan: Injecting the entire soundfont_catalog.json file into the AI System Instruction. A full SoundFont file (like GeneralUser GS or SGM-V2.01) can contain hundreds to thousands of presets/notes, causing LLM Token Limit overflows, inflating costs, and degrading response latency.
  • Solution:
  • Implement get_condensed_catalog_summary() within SoundFontInspector to extract only a condensed catalog (categorized into core instrument groups: Piano, Organ, Guitar, Bass, Strings, Ensemble, Brass, Reed, Pipe, Synth Lead, Synth Pad, Drum Kit).
  • Inject a maximum of 4050 of the most common instruments along with their representative bank and program codes into the AI Prompt.

💡 Refinement 2: MIDI Channel Handling for Percussion Kits (Bank 128 / Percussion)

  • Problem in Previous Plan: Defaulting to channel=0 for all tracks when invoking program_select(0, fid, bank, prog). In General MIDI and SoundFont (.sf2) standards, Drum/Percussion sounds (Bank 128) must reside on MIDI Channel 9 (the 10th channel, 0-based index 9).
  • Solution:
  • In render_engine.py, if soundfont_bank == 128 or is_percussion == True, automatically assign that track's MIDI Channel to channel = 9 for both FluidSynth rendering and mido message generation.

3. Refinement 3: Catalog Refresh on User SoundFont Upload (Cache Invalidation)

  • Problem in Previous Plan: The GET /api/v1/plugins/soundfonts/catalog endpoint scans the catalog only once or upon application startup. When a user uploads a new .sf2 file via POST /api/v1/audio/upload-soundfont, the AI remains unaware of the newly added file.
  • Solution:
  • Implement a Cache Invalidation mechanism: Upon successfully saving an uploaded .sf2 file, automatically invoke SoundFontInspector.generate_full_catalog() to update the soundfont_catalog.json file.

💡 Refinement 4: Handling Relative Sample Paths for Pianobook .dspreset Files

  • Problem in Previous Plan: Pianobook .dspreset files contain relative path links pointing to subfolder samples/*.wav files. When DecentSampler VST3 loads a .dspreset file via pedalboard, if the Working Directory is not set to the folder containing the .dspreset file, the VST3 engine triggers a "Sample Not Found" error (resulting in silence).
  • Solution:
  • Before invoking plugin.load_preset(dspreset_path), ensure an absolute path (os.path.abspath(dspreset_path)) is passed and temporarily switch the Working Directory or properly configure the Root Sample Directory for DecentSampler.

💡 Refinement 5: Robust Error Handling in SoundFontInspector

  • Problem in Previous Plan: If a user uploads a corrupted or malformed .sf2 file, the sf2utils library may throw an exception, crashing the entire Catalog scanning workflow.
  • Solution:
  • Wrap each .sf2 file processing block inside a try...except block in SoundFontInspector. If a file is corrupted, log a warning and skip that specific file instead of interrupting the complete scan process.

3. Updated Execution Plan

Task A: SoundFont Inspection Engine (sf2utils)

  • A1: Add sf2utils>=0.9.0 to requirements.txt.

  • A2: Create app/core/soundfont_inspector.py:

  • Add inspect_sf2_file(filepath) wrapped in a try...except block.

  • Add generate_full_catalog(output_json_path) scanning both /opt/daw_engine/soundfonts and app/storage/uploads/soundfonts.

  • Add get_condensed_catalog_summary() to build a condensed summary for the AI Context Prompt.

  • A3: Create API Endpoint GET /api/v1/plugins/soundfonts/catalog in app/api/v1/plugins.py:

  • Return Full Catalog for Frontend UI and Condensed Catalog for AI Agent.

  • Integrate cache refresh functionality triggered upon new .sf2 file uploads.

Task B: DecentSampler + Pianobook Support

  • B1: Update Dockerfile:

  • Add libcurl4 to the apt-get install package list.

  • Pre-create directory structures /opt/daw_engine/vst3/ and /opt/daw_engine/samples/pianobook/.

  • B2: Create local host directory structure vst_plugins/ and samples/pianobook/ (Update .gitignore).

  • B3: Update app/core/vst_engine.py:

  • Add DecentSamplerManager supporting .dspreset loading using absolute paths.

  • B4: Integrate Pianobook rendering into app/core/render_engine.py when a Track selects a Pianobook instrument.

Task C: AI Tool Schema & Prompt Injection

  • C1: Update DEFAULT_TOOLS in app/static/js/services/aiGateway.js:

  • Add 3 properties to the generate_multitrack_midi schema: soundfont_id (string), soundfont_bank (integer), soundfont_program (integer).

  • C2: Inject condensed instrument catalog into System Instruction within aiGateway.js.

  • C3: Load Catalog automatically upon Frontend application startup (app.jsx).

Task D: Server Render — Program Change & Channel Mapping

  • D1: Update render_engine.py:

  • Read soundfont_bank and soundfont_program from Track metadata.

  • MIDI channel rules: If soundfont_bank == 128 (Drums), automatically assign channel = 9 (GM Standard Channel 10). Otherwise, assign channels from 0 through 8.

  • D2: Update FluidSynth render path:

midi_channel = 9 if (bank == 128 or track.get("is_percussion")) else target_channel
fl.program_select(midi_channel, fid, bank, prog)

  • D3: Update VST3/Pedalboard render path:
  • Insert CONTROL_CHANGE (CC 0 for Bank) and PROGRAM_CHANGE events into the note sequence prior to rendering the audio buffer.

Task E: Client SoundFont Player — Program Change

  • E1: Update app/static/js/services/soundfontPlayer.js:

  • Add programChange(channel, program) and controllerChange(channel, controller, value) methods.

  • E2: Add applyAITrackInstrument(trackId, bank, program) function to dynamically switch instrument sounds in real time when AI generates new Tracks on the UI.

Task F: Validation & Testing

  • F1: Test Catalog API: GET /api/v1/plugins/soundfonts/catalog.
  • F2: Test AI Generation: Input prompt "Compose 8 bars of Brass horns and a drum kit" \rightarrow Verify AI returns JSON with program=56 (Brass) and bank=128 (Drums).
  • F3: Test Server Render: Export WAV \rightarrow Listen to output audio file to verify correct Brass horn and Drum sound execution.