fix: đã sửa lỗi MIDI note bị flicker
This commit is contained in:
Binary file not shown.
@@ -0,0 +1,344 @@
|
|||||||
|
# TECHNICAL PLAN: DECENT SAMPLER + PIANOBOOK INSTALLATION AND SOUNDFONT MAPPING EXTRACTION FOR AI AGENT
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Executive Overview
|
||||||
|
|
||||||
|
The system needs to fulfill two core requirements:
|
||||||
|
|
||||||
|
* **Integrate DecentSampler & Pianobook on Linux Server:**
|
||||||
|
* Install `DecentSampler.vst3` (Linux 64-bit) into the Docker Server environment.
|
||||||
|
* Structure the Pianobook sample library directories (`.dspreset` + `.wav` files).
|
||||||
|
* Integrate Preset Loading into Python `pedalboard` for offline rendering steps.
|
||||||
|
|
||||||
|
|
||||||
|
* **Resolve the AI Agent's "Instrument Information Blindness" regarding SoundFont (`.sf2`):**
|
||||||
|
* *Current State:* The AI Agent and the system only load raw `.sf2` files without knowing what instruments are contained within (which Bank, which Program/Patch number, or what the instrument names are).
|
||||||
|
* *Solution:*
|
||||||
|
* Build a **SoundFont Inspection Engine** using Python (`sf2utils`) to scan `.sf2` files upon upload/scan and extract the instrument catalog table (Bank, Program/Preset ID, Instrument Name).
|
||||||
|
* Export a Catalog table (`soundfont_catalog.json`) and pass this context to the AI Agent.
|
||||||
|
* Update the AI Tool Schema so the AI accurately passes `soundfont_id`, `bank`, and `program` (MIDI Program Change) when creating a Track.
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. DecentSampler + Pianobook Installation Plan (Server Backend)
|
||||||
|
|
||||||
|
### 2.1 Installing DecentSampler Linux Native VST3 in Docker
|
||||||
|
|
||||||
|
1. **Download DecentSampler Linux VST3:**
|
||||||
|
* Download the official Linux 64-bit build from the DecentSampler website (`DecentSampler_Linux_x64.tar.gz` or `.vst3` file).
|
||||||
|
|
||||||
|
|
||||||
|
2. **Server Directory Structure:**
|
||||||
|
```text
|
||||||
|
/opt/daw_engine/
|
||||||
|
├── vst3/
|
||||||
|
│ └── DecentSampler.vst3/ <-- Native Linux VST3 Binary
|
||||||
|
├── soundfonts/
|
||||||
|
│ ├── GeneralUser_GS.sf2
|
||||||
|
│ └── SGM-V2.01.sf2
|
||||||
|
└── samples/
|
||||||
|
└── pianobook/ <-- Pianobook Sample Libraries
|
||||||
|
├── salamander_grand_piano/
|
||||||
|
│ ├── salamander_piano.dspreset
|
||||||
|
│ └── samples/ (*.wav)
|
||||||
|
└── acoustic_guitar/
|
||||||
|
├── guitar.dspreset
|
||||||
|
└── samples/ (*.wav)
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
|
3. **Additions to `Dockerfile`:**
|
||||||
|
```dockerfile
|
||||||
|
# Install audio system dependencies
|
||||||
|
RUN apt-get update && apt-get install -y \
|
||||||
|
libgl1-mesa-glx \
|
||||||
|
libfreetype6 \
|
||||||
|
libcurl4 \
|
||||||
|
&& rm -rf /var/lib/apt/lists/*
|
||||||
|
|
||||||
|
# Copy DecentSampler VST3 to Server
|
||||||
|
COPY ./vst_plugins/DecentSampler.vst3 /opt/daw_engine/vst3/DecentSampler.vst3
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
### 2.2 Integrating DecentSampler into the Python Engine (`app/core/vst_engine.py`)
|
||||||
|
|
||||||
|
The `pedalboard` library supports loading VST3 plugins and preset files for DecentSampler:
|
||||||
|
|
||||||
|
```python
|
||||||
|
import os
|
||||||
|
from pedalboard import VST3Plugin
|
||||||
|
|
||||||
|
class DecentSamplerManager:
|
||||||
|
def __init__(self, vst_path="/opt/daw_engine/vst3/DecentSampler.vst3"):
|
||||||
|
self.vst_path = vst_path
|
||||||
|
|
||||||
|
def create_decent_sampler_instance(self, dspreset_path: str) -> VST3Plugin:
|
||||||
|
"""
|
||||||
|
Instantiates VST3 DecentSampler and loads the Pianobook sample preset (.dspreset) file.
|
||||||
|
"""
|
||||||
|
if not os.path.exists(self.vst_path):
|
||||||
|
raise FileNotFoundError(f"DecentSampler VST3 not found at {self.vst_path}")
|
||||||
|
|
||||||
|
plugin = VST3Plugin(self.vst_path)
|
||||||
|
|
||||||
|
# Load the Pianobook preset file into DecentSampler VST3
|
||||||
|
if os.path.exists(dspreset_path):
|
||||||
|
plugin.load_preset(dspreset_path)
|
||||||
|
|
||||||
|
return plugin
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Designing the SoundFont Inspection Engine (Bank/Program Extraction)
|
||||||
|
|
||||||
|
Every SoundFont (`.sf2`) file is a collection of Presets (or Programs). For the AI Agent to know what sound presets exist inside the `.sf2` file, the Backend must scan and parse the `.sf2` file.
|
||||||
|
|
||||||
|
### 3.1 Adding Python SoundFont Inspection Libraries
|
||||||
|
|
||||||
|
Add to `requirements.txt`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
sf2utils>=0.9.0
|
||||||
|
mido>=1.3.0
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.2 Building the SoundFont Metadata Inspection Service (`app/core/soundfont_inspector.py`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
import os
|
||||||
|
import json
|
||||||
|
from sf2utils.sf2parse import Sf2File
|
||||||
|
|
||||||
|
class SoundFontInspector:
|
||||||
|
def __init__(self, sf_dir="/opt/daw_engine/soundfonts"):
|
||||||
|
self.sf_dir = sf_dir
|
||||||
|
|
||||||
|
def inspect_sf2_file(self, filepath: str) -> dict:
|
||||||
|
"""
|
||||||
|
Parses an .sf2 file and returns a complete instrument catalog (Bank, Program, Instrument Name).
|
||||||
|
"""
|
||||||
|
if not os.path.exists(filepath):
|
||||||
|
return {}
|
||||||
|
|
||||||
|
sf_name = os.path.basename(filepath)
|
||||||
|
sf_id = os.path.splitext(sf_name)[0].lower()
|
||||||
|
|
||||||
|
instruments = []
|
||||||
|
with open(filepath, 'rb') as f:
|
||||||
|
sf2 = Sf2File(f)
|
||||||
|
for preset in sf2.presets:
|
||||||
|
# Ignore EOP (End of Header) preset
|
||||||
|
if preset.name.strip() == "EOP" or (preset.bank == 128 and preset.preset == 127):
|
||||||
|
continue
|
||||||
|
|
||||||
|
instruments.append({
|
||||||
|
"bank": preset.bank, # Bank number (0 = General MIDI Standard, 128 = Percussion/Drums)
|
||||||
|
"program": preset.preset, # Program/Patch number (0-127)
|
||||||
|
"name": preset.name.strip(), # Instrument name (e.g. "Stereo Grand", "Violin", "Brass Section")
|
||||||
|
"is_percussion": (preset.bank == 128)
|
||||||
|
})
|
||||||
|
|
||||||
|
return {
|
||||||
|
"soundfont_id": sf_id,
|
||||||
|
"filename": sf_name,
|
||||||
|
"total_instruments": len(instruments),
|
||||||
|
"instruments": instruments
|
||||||
|
}
|
||||||
|
|
||||||
|
def generate_full_catalog(self, output_json_path="/opt/daw_engine/soundfont_catalog.json"):
|
||||||
|
"""
|
||||||
|
Scans all .sf2 files in the directory and builds a Catalog JSON for the AI Agent.
|
||||||
|
"""
|
||||||
|
catalog = {}
|
||||||
|
for root, dirs, files in os.walk(self.sf_dir):
|
||||||
|
for file in files:
|
||||||
|
if file.endswith(('.sf2', '.SF2')):
|
||||||
|
full_path = os.path.join(root, file)
|
||||||
|
sf_info = self.inspect_sf2_file(full_path)
|
||||||
|
catalog[sf_info["soundfont_id"]] = sf_info
|
||||||
|
|
||||||
|
with open(output_json_path, 'w', encoding='utf-8') as f:
|
||||||
|
json.dump(catalog, f, ensure_ascii=False, indent=2)
|
||||||
|
|
||||||
|
return catalog
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.3 Catalog File Output Structure (`soundfont_catalog.json`)
|
||||||
|
|
||||||
|
This JSON file serves as an Instrument Dictionary for the AI Agent:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"generaluser_gs": {
|
||||||
|
"soundfont_id": "generaluser_gs",
|
||||||
|
"filename": "GeneralUser_GS.sf2",
|
||||||
|
"total_instruments": 128,
|
||||||
|
"instruments": [
|
||||||
|
{ "bank": 0, "program": 0, "name": "Stereo Grand Piano", "is_percussion": false },
|
||||||
|
{ "bank": 0, "program": 19, "name": "Church Organ", "is_percussion": false },
|
||||||
|
{ "bank": 0, "program": 40, "name": "Violin Ensemble", "is_percussion": false },
|
||||||
|
{ "bank": 0, "program": 56, "name": "Trumpet", "is_percussion": false },
|
||||||
|
{ "bank": 128, "program": 0, "name": "Standard Drum Kit", "is_percussion": true }
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Guiding the AI Agent in Instrument Selection & Accurate Note Loading
|
||||||
|
|
||||||
|
When the user types: *"Create a Piano track and a Strings section track for 8 bars"*, the AI Agent needs to know precisely which soundfont, bank, and program numbers to assign to the Tracks.
|
||||||
|
|
||||||
|
### 4.1 Updating the AI Tool Function Spec (`tools_spec.json`)
|
||||||
|
|
||||||
|
Add `soundfont_bank` and `soundfont_program` fields to the Tool Schema sent to the AI:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"type": "function",
|
||||||
|
"function": {
|
||||||
|
"name": "generate_multitrack_midi",
|
||||||
|
"description": "Generates multi-track MIDI data along with appropriate SoundFont Program configurations.",
|
||||||
|
"parameters": {
|
||||||
|
"type": "object",
|
||||||
|
"properties": {
|
||||||
|
"tracks": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "object",
|
||||||
|
"properties": {
|
||||||
|
"track_name": { "type": "string" },
|
||||||
|
"instrument_type": { "type": "string", "enum": ["PIANO", "STRINGS", "BRASS", "SYNTH", "DRUMS"] },
|
||||||
|
"soundfont_id": {
|
||||||
|
"type": "string",
|
||||||
|
"description": "ID of the SoundFont to use (e.g. 'generaluser_gs')"
|
||||||
|
},
|
||||||
|
"soundfont_bank": {
|
||||||
|
"type": "integer",
|
||||||
|
"default": 0,
|
||||||
|
"description": "MIDI Bank code (0 for melodic instruments, 128 for Drums)"
|
||||||
|
},
|
||||||
|
"soundfont_program": {
|
||||||
|
"type": "integer",
|
||||||
|
"description": "MIDI Program Number (0-127) corresponding to the instrument name in the Catalog"
|
||||||
|
},
|
||||||
|
"notes": { "type": "array", "items": { "type": "object" } }
|
||||||
|
},
|
||||||
|
"required": ["track_name", "soundfont_id", "soundfont_bank", "soundfont_program", "notes"]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4.2 Injecting the Catalog into the AI Agent's Context Prompt (Prompt Template)
|
||||||
|
|
||||||
|
Before sending the user query to the LLM, the system reads `soundfont_catalog.json` and injects a condensed catalog table into the System Instruction:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# System Context Prompt Injector
|
||||||
|
def build_ai_system_instruction(catalog_data: dict) -> str:
|
||||||
|
sf_summary = []
|
||||||
|
for sf_id, sf_info in catalog_data.items():
|
||||||
|
sf_summary.append(f"SoundFont ID: '{sf_id}' (File: {sf_info['filename']}):")
|
||||||
|
for inst in sf_info['instruments'][:20]: # Inject primary instrument lists
|
||||||
|
sf_summary.append(
|
||||||
|
f" - [{inst['name']}]: bank={inst['bank']}, program={inst['program']}"
|
||||||
|
)
|
||||||
|
|
||||||
|
catalog_context = "\n".join(sf_summary)
|
||||||
|
|
||||||
|
system_instruction = f"""
|
||||||
|
You are an AI Copilot for a DAW. Below is the Catalog of available SoundFonts on the system:
|
||||||
|
|
||||||
|
{catalog_context}
|
||||||
|
|
||||||
|
MANDATORY RULES WHEN CREATING TRACKS:
|
||||||
|
1. When creating any track, you MUST look up the catalog above and fill in the correct `soundfont_id`, `soundfont_bank`, and `soundfont_program`.
|
||||||
|
2. Example: If the user requests "Piano", select soundfont_id="generaluser_gs", soundfont_bank=0, soundfont_program=0 ("Stereo Grand Piano").
|
||||||
|
3. If the user requests "Violin/Strings", select soundfont_bank=0, soundfont_program=40 ("Violin Ensemble").
|
||||||
|
4. If the user requests "Drums", select soundfont_bank=128, soundfont_program=0 ("Standard Drum Kit").
|
||||||
|
"""
|
||||||
|
return system_instruction
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4.3 Applying Program Change on Client & Server Render Layers
|
||||||
|
|
||||||
|
#### A. Client Browser Side (FluidSynth Wasm / SoundFont Player)
|
||||||
|
|
||||||
|
Upon receiving JSON from the AI, the Frontend invokes the Bank and Program selection function to trigger the correct sound:
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
// Client-side Javascript (spessasynth / fluidsynth.wasm)
|
||||||
|
function applyAITrackInstrument(trackId, soundfontBank, soundfontProgram) {
|
||||||
|
const channel = getTrackMIDIChannel(trackId);
|
||||||
|
|
||||||
|
// Send MIDI Bank Select (CC 0)
|
||||||
|
synthInstance.controllerChange(channel, 0, soundfontBank);
|
||||||
|
|
||||||
|
// Send MIDI Program Change
|
||||||
|
synthInstance.programChange(channel, soundfontProgram);
|
||||||
|
}
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
#### B. Server Offline Render Side (`app/core/render_engine.py`)
|
||||||
|
|
||||||
|
When rendering to a WAV file, Python inserts a MIDI Program Change event ahead of the Track's note sequence:
|
||||||
|
|
||||||
|
```python
|
||||||
|
import mido
|
||||||
|
|
||||||
|
def create_midi_track_with_program(notes_data, bank=0, program=0):
|
||||||
|
midi_track = mido.MidiTrack()
|
||||||
|
|
||||||
|
# 1. Insert Bank Select (Control Change 0)
|
||||||
|
midi_track.append(mido.Message('control_change', channel=0, control=0, value=bank, time=0))
|
||||||
|
|
||||||
|
# 2. Insert Program Change (Instrument Sound Selection)
|
||||||
|
midi_track.append(mido.Message('program_change', channel=0, program=program, time=0))
|
||||||
|
|
||||||
|
# 3. Insert MIDI notes generated by AI
|
||||||
|
for note in notes_data:
|
||||||
|
start_tick = int(note['start_beat'] * 480) # 480 ticks per beat
|
||||||
|
dur_tick = int(note['duration_beats'] * 480)
|
||||||
|
pitch = int(note['pitch'])
|
||||||
|
vel = int(note['velocity'] * 127)
|
||||||
|
|
||||||
|
midi_track.append(mido.Message('note_on', note=pitch, velocity=vel, time=start_tick))
|
||||||
|
midi_track.append(mido.Message('note_off', note=pitch, velocity=0, time=dur_tick))
|
||||||
|
|
||||||
|
return midi_track
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Action Checklist
|
||||||
|
|
||||||
|
* [ ] **Step 1:** Download Linux 64-bit `DecentSampler.vst3` and copy it into `/opt/daw_engine/vst3/`.
|
||||||
|
* [ ] **Step 2:** Download Pianobook sound libraries (e.g. Salamander Grand Piano) and extract them to `/opt/daw_engine/samples/pianobook/`.
|
||||||
|
* [ ] **Step 3:** Add `sf2utils` to `requirements.txt` and install it in Docker.
|
||||||
|
* [ ] **Step 4:** Create `app/core/soundfont_inspector.py` to automatically scan all `.sf2` files in the project and generate `soundfont_catalog.json`.
|
||||||
|
* [ ] **Step 5:** Build API Endpoint `GET /api/v1/plugins/soundfonts/catalog` returning the extracted instrument catalog.
|
||||||
|
* [ ] **Step 6:** Update the Prompt Template and AI Tool Schema to support `soundfont_bank` & `soundfont_program`.
|
||||||
|
* [ ] **Step 7:** End-to-End Verification: Type Prompt *"Generate 8 bars of Brass horn music"* $\rightarrow$ AI reads Catalog and selects Program `56` $\rightarrow$ Web Audio & Server Render produce the correct Brass horn instrument sound.
|
||||||
Reference in New Issue
Block a user