fix: đã sửa lỗi MIDI note bị flicker

This commit is contained in:
2026-07-25 20:46:53 +07:00
parent f2d02a358f
commit fde9a91c9b
2 changed files with 344 additions and 0 deletions
Binary file not shown.
+344
View File
@@ -0,0 +1,344 @@
# TECHNICAL PLAN: DECENT SAMPLER + PIANOBOOK INSTALLATION AND SOUNDFONT MAPPING EXTRACTION FOR AI AGENT
---
## 1. Executive Overview
The system needs to fulfill two core requirements:
* **Integrate DecentSampler & Pianobook on Linux Server:**
* Install `DecentSampler.vst3` (Linux 64-bit) into the Docker Server environment.
* Structure the Pianobook sample library directories (`.dspreset` + `.wav` files).
* Integrate Preset Loading into Python `pedalboard` for offline rendering steps.
* **Resolve the AI Agent's "Instrument Information Blindness" regarding SoundFont (`.sf2`):**
* *Current State:* The AI Agent and the system only load raw `.sf2` files without knowing what instruments are contained within (which Bank, which Program/Patch number, or what the instrument names are).
* *Solution:*
* Build a **SoundFont Inspection Engine** using Python (`sf2utils`) to scan `.sf2` files upon upload/scan and extract the instrument catalog table (Bank, Program/Preset ID, Instrument Name).
* Export a Catalog table (`soundfont_catalog.json`) and pass this context to the AI Agent.
* Update the AI Tool Schema so the AI accurately passes `soundfont_id`, `bank`, and `program` (MIDI Program Change) when creating a Track.
---
## 2. DecentSampler + Pianobook Installation Plan (Server Backend)
### 2.1 Installing DecentSampler Linux Native VST3 in Docker
1. **Download DecentSampler Linux VST3:**
* Download the official Linux 64-bit build from the DecentSampler website (`DecentSampler_Linux_x64.tar.gz` or `.vst3` file).
2. **Server Directory Structure:**
```text
/opt/daw_engine/
├── vst3/
│ └── DecentSampler.vst3/ <-- Native Linux VST3 Binary
├── soundfonts/
│ ├── GeneralUser_GS.sf2
│ └── SGM-V2.01.sf2
└── samples/
└── pianobook/ <-- Pianobook Sample Libraries
├── salamander_grand_piano/
│ ├── salamander_piano.dspreset
│ └── samples/ (*.wav)
└── acoustic_guitar/
├── guitar.dspreset
└── samples/ (*.wav)
```
3. **Additions to `Dockerfile`:**
```dockerfile
# Install audio system dependencies
RUN apt-get update && apt-get install -y \
libgl1-mesa-glx \
libfreetype6 \
libcurl4 \
&& rm -rf /var/lib/apt/lists/*
# Copy DecentSampler VST3 to Server
COPY ./vst_plugins/DecentSampler.vst3 /opt/daw_engine/vst3/DecentSampler.vst3
```
### 2.2 Integrating DecentSampler into the Python Engine (`app/core/vst_engine.py`)
The `pedalboard` library supports loading VST3 plugins and preset files for DecentSampler:
```python
import os
from pedalboard import VST3Plugin
class DecentSamplerManager:
def __init__(self, vst_path="/opt/daw_engine/vst3/DecentSampler.vst3"):
self.vst_path = vst_path
def create_decent_sampler_instance(self, dspreset_path: str) -> VST3Plugin:
"""
Instantiates VST3 DecentSampler and loads the Pianobook sample preset (.dspreset) file.
"""
if not os.path.exists(self.vst_path):
raise FileNotFoundError(f"DecentSampler VST3 not found at {self.vst_path}")
plugin = VST3Plugin(self.vst_path)
# Load the Pianobook preset file into DecentSampler VST3
if os.path.exists(dspreset_path):
plugin.load_preset(dspreset_path)
return plugin
```
---
## 3. Designing the SoundFont Inspection Engine (Bank/Program Extraction)
Every SoundFont (`.sf2`) file is a collection of Presets (or Programs). For the AI Agent to know what sound presets exist inside the `.sf2` file, the Backend must scan and parse the `.sf2` file.
### 3.1 Adding Python SoundFont Inspection Libraries
Add to `requirements.txt`:
```text
sf2utils>=0.9.0
mido>=1.3.0
```
### 3.2 Building the SoundFont Metadata Inspection Service (`app/core/soundfont_inspector.py`)
```python
import os
import json
from sf2utils.sf2parse import Sf2File
class SoundFontInspector:
def __init__(self, sf_dir="/opt/daw_engine/soundfonts"):
self.sf_dir = sf_dir
def inspect_sf2_file(self, filepath: str) -> dict:
"""
Parses an .sf2 file and returns a complete instrument catalog (Bank, Program, Instrument Name).
"""
if not os.path.exists(filepath):
return {}
sf_name = os.path.basename(filepath)
sf_id = os.path.splitext(sf_name)[0].lower()
instruments = []
with open(filepath, 'rb') as f:
sf2 = Sf2File(f)
for preset in sf2.presets:
# Ignore EOP (End of Header) preset
if preset.name.strip() == "EOP" or (preset.bank == 128 and preset.preset == 127):
continue
instruments.append({
"bank": preset.bank, # Bank number (0 = General MIDI Standard, 128 = Percussion/Drums)
"program": preset.preset, # Program/Patch number (0-127)
"name": preset.name.strip(), # Instrument name (e.g. "Stereo Grand", "Violin", "Brass Section")
"is_percussion": (preset.bank == 128)
})
return {
"soundfont_id": sf_id,
"filename": sf_name,
"total_instruments": len(instruments),
"instruments": instruments
}
def generate_full_catalog(self, output_json_path="/opt/daw_engine/soundfont_catalog.json"):
"""
Scans all .sf2 files in the directory and builds a Catalog JSON for the AI Agent.
"""
catalog = {}
for root, dirs, files in os.walk(self.sf_dir):
for file in files:
if file.endswith(('.sf2', '.SF2')):
full_path = os.path.join(root, file)
sf_info = self.inspect_sf2_file(full_path)
catalog[sf_info["soundfont_id"]] = sf_info
with open(output_json_path, 'w', encoding='utf-8') as f:
json.dump(catalog, f, ensure_ascii=False, indent=2)
return catalog
```
### 3.3 Catalog File Output Structure (`soundfont_catalog.json`)
This JSON file serves as an Instrument Dictionary for the AI Agent:
```json
{
"generaluser_gs": {
"soundfont_id": "generaluser_gs",
"filename": "GeneralUser_GS.sf2",
"total_instruments": 128,
"instruments": [
{ "bank": 0, "program": 0, "name": "Stereo Grand Piano", "is_percussion": false },
{ "bank": 0, "program": 19, "name": "Church Organ", "is_percussion": false },
{ "bank": 0, "program": 40, "name": "Violin Ensemble", "is_percussion": false },
{ "bank": 0, "program": 56, "name": "Trumpet", "is_percussion": false },
{ "bank": 128, "program": 0, "name": "Standard Drum Kit", "is_percussion": true }
]
}
}
```
---
## 4. Guiding the AI Agent in Instrument Selection & Accurate Note Loading
When the user types: *"Create a Piano track and a Strings section track for 8 bars"*, the AI Agent needs to know precisely which soundfont, bank, and program numbers to assign to the Tracks.
### 4.1 Updating the AI Tool Function Spec (`tools_spec.json`)
Add `soundfont_bank` and `soundfont_program` fields to the Tool Schema sent to the AI:
```json
{
"type": "function",
"function": {
"name": "generate_multitrack_midi",
"description": "Generates multi-track MIDI data along with appropriate SoundFont Program configurations.",
"parameters": {
"type": "object",
"properties": {
"tracks": {
"type": "array",
"items": {
"type": "object",
"properties": {
"track_name": { "type": "string" },
"instrument_type": { "type": "string", "enum": ["PIANO", "STRINGS", "BRASS", "SYNTH", "DRUMS"] },
"soundfont_id": {
"type": "string",
"description": "ID of the SoundFont to use (e.g. 'generaluser_gs')"
},
"soundfont_bank": {
"type": "integer",
"default": 0,
"description": "MIDI Bank code (0 for melodic instruments, 128 for Drums)"
},
"soundfont_program": {
"type": "integer",
"description": "MIDI Program Number (0-127) corresponding to the instrument name in the Catalog"
},
"notes": { "type": "array", "items": { "type": "object" } }
},
"required": ["track_name", "soundfont_id", "soundfont_bank", "soundfont_program", "notes"]
}
}
}
}
}
}
```
### 4.2 Injecting the Catalog into the AI Agent's Context Prompt (Prompt Template)
Before sending the user query to the LLM, the system reads `soundfont_catalog.json` and injects a condensed catalog table into the System Instruction:
```python
# System Context Prompt Injector
def build_ai_system_instruction(catalog_data: dict) -> str:
sf_summary = []
for sf_id, sf_info in catalog_data.items():
sf_summary.append(f"SoundFont ID: '{sf_id}' (File: {sf_info['filename']}):")
for inst in sf_info['instruments'][:20]: # Inject primary instrument lists
sf_summary.append(
f" - [{inst['name']}]: bank={inst['bank']}, program={inst['program']}"
)
catalog_context = "\n".join(sf_summary)
system_instruction = f"""
You are an AI Copilot for a DAW. Below is the Catalog of available SoundFonts on the system:
{catalog_context}
MANDATORY RULES WHEN CREATING TRACKS:
1. When creating any track, you MUST look up the catalog above and fill in the correct `soundfont_id`, `soundfont_bank`, and `soundfont_program`.
2. Example: If the user requests "Piano", select soundfont_id="generaluser_gs", soundfont_bank=0, soundfont_program=0 ("Stereo Grand Piano").
3. If the user requests "Violin/Strings", select soundfont_bank=0, soundfont_program=40 ("Violin Ensemble").
4. If the user requests "Drums", select soundfont_bank=128, soundfont_program=0 ("Standard Drum Kit").
"""
return system_instruction
```
### 4.3 Applying Program Change on Client & Server Render Layers
#### A. Client Browser Side (FluidSynth Wasm / SoundFont Player)
Upon receiving JSON from the AI, the Frontend invokes the Bank and Program selection function to trigger the correct sound:
```javascript
// Client-side Javascript (spessasynth / fluidsynth.wasm)
function applyAITrackInstrument(trackId, soundfontBank, soundfontProgram) {
const channel = getTrackMIDIChannel(trackId);
// Send MIDI Bank Select (CC 0)
synthInstance.controllerChange(channel, 0, soundfontBank);
// Send MIDI Program Change
synthInstance.programChange(channel, soundfontProgram);
}
```
#### B. Server Offline Render Side (`app/core/render_engine.py`)
When rendering to a WAV file, Python inserts a MIDI Program Change event ahead of the Track's note sequence:
```python
import mido
def create_midi_track_with_program(notes_data, bank=0, program=0):
midi_track = mido.MidiTrack()
# 1. Insert Bank Select (Control Change 0)
midi_track.append(mido.Message('control_change', channel=0, control=0, value=bank, time=0))
# 2. Insert Program Change (Instrument Sound Selection)
midi_track.append(mido.Message('program_change', channel=0, program=program, time=0))
# 3. Insert MIDI notes generated by AI
for note in notes_data:
start_tick = int(note['start_beat'] * 480) # 480 ticks per beat
dur_tick = int(note['duration_beats'] * 480)
pitch = int(note['pitch'])
vel = int(note['velocity'] * 127)
midi_track.append(mido.Message('note_on', note=pitch, velocity=vel, time=start_tick))
midi_track.append(mido.Message('note_off', note=pitch, velocity=0, time=dur_tick))
return midi_track
```
---
## 5. Action Checklist
* [ ] **Step 1:** Download Linux 64-bit `DecentSampler.vst3` and copy it into `/opt/daw_engine/vst3/`.
* [ ] **Step 2:** Download Pianobook sound libraries (e.g. Salamander Grand Piano) and extract them to `/opt/daw_engine/samples/pianobook/`.
* [ ] **Step 3:** Add `sf2utils` to `requirements.txt` and install it in Docker.
* [ ] **Step 4:** Create `app/core/soundfont_inspector.py` to automatically scan all `.sf2` files in the project and generate `soundfont_catalog.json`.
* [ ] **Step 5:** Build API Endpoint `GET /api/v1/plugins/soundfonts/catalog` returning the extracted instrument catalog.
* [ ] **Step 6:** Update the Prompt Template and AI Tool Schema to support `soundfont_bank` & `soundfont_program`.
* [ ] **Step 7:** End-to-End Verification: Type Prompt *"Generate 8 bars of Brass horn music"* $\rightarrow$ AI reads Catalog and selects Program `56` $\rightarrow$ Web Audio & Server Render produce the correct Brass horn instrument sound.