feat: bổ sung MIDI

This commit is contained in:
2026-07-23 08:32:23 +07:00
parent 6545e1746e
commit 0e44e44ecb
15 changed files with 3304 additions and 12680 deletions
+760
View File
@@ -0,0 +1,760 @@
Here is the complete translation and conversion of the document into a clean, professionally formatted Markdown layout:
# ARCHITECTURAL, TECHNICAL, AND ALGORITHMIC SPECIFICATION
## Hybrid Web-Based Digital Audio Workstation (DAW) with Nested Section Architecture and Non-Destructive Timeline Mechanics
---
### 1. System Overview & Architecture Design
#### 1.1 High-Level Architecture Topology
The system follows a hybrid Client-Server architecture designed for real-time Web-based audio production, composition, and high-performance offline DSP rendering.
* **Frontend Client (HTML5 / Vanilla JS / Web Audio API / HTML5 Canvas)**
* **UI Layer:** HTML5 Canvas / Web Components for high-FPS multi-lane timeline rendering, Piano Roll canvas, Sample Editor, and Sub-Tab navigation.
* **Audio Engine Layer:** Web Audio API `AudioContext` graph, Custom `AudioWorklet` Processors (WebAssembly/JS) for real-time synthesis, playback scheduling, sample playback, and latency-compensated signal routing.
* **State Management Engine:** Immutable/Reactive Central State Store handling Session tree hierarchy, Section Store registries, Undo/Redo stack, and view-state context isolation.
* **Backend Server (Python Engine)**
* **RESTful / WebSocket API:** Event-driven client communication layer (FastAPI or AIOHTTP).
* **DSP / Rendering Engine:** Python-based audio processing (`numpy`, `scipy`, `pyo`, `pedalboard`) for offline stem bouncing, high-fidelity export, sample processing, and optional VST/VSTi hosting/bridging.
```text
+-----------------------------------------------------------------------------------+
| FRONTEND (HTML5/JS) |
| |
| +-----------------------------------------------------------------------------+ |
| | UI & View State System | |
| | +---------------------+ +----------------------+ +--------------------+ | |
| | | Main Session Canvas | | Section-Tab View | | Piano Roll View | | |
| | +---------------------+ +----------------------+ +--------------------+ | |
| +-----------------------------------------------------------------------------+ |
| | |
| +-----------------------------------------------------------------------------+ |
| | Central Data State Store | |
| | [Project Model] ---> [Section Store] ---> [Item Clip Metadata] | |
| +-----------------------------------------------------------------------------+ |
| | |
| +-----------------------------------------------------------------------------+ |
| | Audio & Clock Engine | |
| | +------------------------+ +------------------+ +---------------------+ | |
| | | Precision Scheduler | | Web Audio Graph | | AudioWorklet Synth | | |
| | | (Lookahead Timer) | | AudioNode Router | | / WebAssembly Core | | |
| | +------------------------+ +------------------+ +---------------------+ | |
| +-----------------------------------------------------------------------------+ |
+------------------------------------------^----------------------------------------+
| WebSocket / REST API
+------------------------------------------v----------------------------------------+
| BACKEND SERVER (PYTHON) |
| +-----------------------------------------------------------------------------+ |
| | FastAPI / WebSocket Handler | |
| +-----------------------------------------------------------------------------+ |
| | DSP Engine (Pedalboard / Numpy / Scipy) - Offline Render, Audio Export | |
| +-----------------------------------------------------------------------------+ |
| | VST / VSTi Hosting Bridge & Plugin State Persistence | |
| +-----------------------------------------------------------------------------+ |
+-----------------------------------------------------------------------------------+
```
---
### 2. Detailed Data Schemas (JSON Specification)
#### 2.1 Project Root Schema (`project_schema.json`)
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"title": "DAWProject",
"type": "object",
"properties": {
"project_id": { "type": "string", "format": "uuid" },
"metadata": {
"type": "object",
"properties": {
"title": { "type": "string" },
"bpm": { "type": "number", "minimum": 20.0, "maximum": 999.0, "default": 120.0 },
"time_signature_numerator": { "type": "integer", "default": 4 },
"time_signature_denominator": { "type": "integer", "default": 4 },
"sample_rate": { "type": "integer", "default": 44100 }
},
"required": ["title", "bpm", "time_signature_numerator", "time_signature_denominator", "sample_rate"]
},
"main_session": { "$ref": "#/definitions/SessionContainer" },
"section_store": {
"type": "object",
"description": "Auxiliary registry mapping section_id to sub-session containers",
"additionalProperties": { "$ref": "#/definitions/SessionContainer" }
}
},
"required": ["project_id", "metadata", "main_session", "section_store"],
"definitions": {
"SessionContainer": {
"type": "object",
"properties": {
"id": { "type": "string" },
"name": { "type": "string" },
"is_root": { "type": "boolean" },
"length_bars": { "type": "number", "description": "Computed or manually set total length in bars" },
"auto_compute_length": { "type": "boolean", "default": true },
"tracks": {
"type": "array",
"items": { "$ref": "#/definitions/Track" }
}
},
"required": ["id", "is_root", "tracks"]
},
"Track": {
"type": "object",
"properties": {
"id": { "type": "string" },
"name": { "type": "string" },
"type": { "type": "string", "enum": ["AUDIO", "MIDI", "SECTION"] },
"volume_db": { "type": "number", "default": 0.0 },
"pan": { "type": "number", "minimum": -1.0, "maximum": 1.0, "default": 0.0 },
"mute": { "type": "boolean", "default": false },
"solo": { "type": "boolean", "default": false },
"fx_chain": {
"type": "array",
"items": { "$ref": "#/definitions/FXPlugin" }
},
"synth_engine": { "$ref": "#/definitions/SynthPlugin" },
"items": {
"type": "array",
"items": { "$ref": "#/definitions/TimelineItem" }
}
},
"required": ["id", "name", "type", "items"]
},
"TimelineItem": {
"type": "object",
"properties": {
"id": { "type": "string" },
"name": { "type": "string" },
"type": { "type": "string", "enum": ["AUDIO_ITEM", "MIDI_ITEM", "SECTION_ITEM"] },
"start_bar": { "type": "number", "description": "Global timeline position where the item starts" },
"duration_bars": { "type": "number", "description": "Visible duration on the track timeline in bars" },
"clip_start_offset_bars": { "type": "number", "description": "Internal start offset inside the source buffer/item" },
"source_data": {
"type": "object",
"oneOf": [
{ "$ref": "#/definitions/AudioSourceData" },
{ "$ref": "#/definitions/MIDISourceData" },
{ "$ref": "#/definitions/SectionSourceData" }
]
}
},
"required": ["id", "type", "start_bar", "duration_bars", "clip_start_offset_bars", "source_data"]
},
"AudioSourceData": {
"type": "object",
"properties": {
"audio_file_url": { "type": "string" },
"sample_rate": { "type": "integer" },
"channels": { "type": "integer" },
"gain": { "type": "number", "default": 1.0 }
},
"required": ["audio_file_url"]
},
"MIDISourceData": {
"type": "object",
"properties": {
"total_buffer_bars": { "type": "number", "default": 8.0 },
"notes": {
"type": "array",
"items": { "$ref": "#/definitions/MIDINote" }
}
},
"required": ["total_buffer_bars", "notes"]
},
"SectionSourceData": {
"type": "object",
"properties": {
"referenced_section_id": { "type": "string", "description": "Pointer to section_store key" }
},
"required": ["referenced_section_id"]
},
"MIDINote": {
"type": "object",
"properties": {
"id": { "type": "string" },
"pitch": { "type": "integer", "minimum": 0, "maximum": 127 },
"start_beat": { "type": "number", "description": "Beat offset relative to the start of the source buffer (bar 0)" },
"duration_beats": { "type": "number" },
"velocity": { "type": "number", "minimum": 0.0, "maximum": 1.0, "default": 0.8 },
"pan": { "type": "number", "minimum": -1.0, "maximum": 1.0, "default": 0.0 }
},
"required": ["id", "pitch", "start_beat", "duration_beats", "velocity"]
},
"FXPlugin": {
"type": "object",
"properties": {
"plugin_id": { "type": "string" },
"name": { "type": "string" },
"bypass": { "type": "boolean", "default": false },
"parameters": { "type": "object" }
}
},
"SynthPlugin": {
"type": "object",
"properties": {
"plugin_id": { "type": "string" },
"preset_id": { "type": "string" },
"parameters": { "type": "object" }
}
}
}
}
```
---
### 3. UI, Tab Navigation & View State Management
#### 3.1 Tab Context Model, Pinning Rules & Close Prevention Hierarchy
The application manages view tabs dynamically while maintaining strict lifecycle integrity:
* **Main Session Tab (Fixed / Pinned):** Always pinned at index 0 (`is_closeable: false`). It cannot be closed under any circumstances.
* **Sub-Tabs (Section-Tab, Piano Roll Tab, Audio Sample Editor Sub-Tab):** Dynamic views (`is_closeable: true`).
* **Parent-Child Tab Dependency Rules:**
* A Section-Tab represents an intermediate sub-session.
* When a user opens a child item (e.g., a `MIDIItem` or `AudioItem` inside a Section-Tab) into a Piano Roll Tab or Audio Sample Editor Sub-Tab, a parent-child context lineage is registered.
* **Close Block Rule:** A Section-Tab cannot be closed while any of its child items are currently open in active sub-tabs. Attempting to close the parent Section-Tab displays a block notice highlighting open child editors.
```text
+---------------------------------------+
| Tab Navigation Controller |
+-------------------+-------------------+
|
+--------------------------------+--------------------------------+
| (Pinned / Uncloseable) | (Dynamic / Closable) | (Dynamic / Closable)
+--------v--------+ +--------v--------+ +--------v--------+
| MAIN SESSION | | SECTION TAB | | PIANO ROLL TAB |
| (Root Context) | | (Sub-Session) | | (Item Context) |
| | | [Parent Context] | [Child Context]|
+-----------------+ +--------+--------+ +--------+--------+
| |
+---- Depends on child closure ---+
```
##### State Object Schema with Tab Dependency Tracking:
```json
{
"active_tab_id": "tab_pr_1",
"open_tabs": [
{
"tab_id": "tab_root",
"title": "MAIN SESSION",
"type": "MAIN_SESSION",
"target_id": "main",
"is_closeable": false,
"parent_tab_id": null
},
{
"tab_id": "tab_sec_1",
"title": "Section: Verse 1",
"type": "SECTION_TAB",
"target_id": "Section_01",
"is_closeable": true,
"parent_tab_id": "tab_root"
},
{
"tab_id": "tab_pr_1",
"title": "Piano Roll: Bassline",
"type": "PIANO_ROLL",
"target_id": "ItemMIDI_Bassline",
"is_closeable": true,
"parent_tab_id": "tab_sec_1"
}
],
"piano_roll_state": {
"target_item_id": "ItemMIDI_Bassline",
"viewport_start_bar": 0.0,
"viewport_bar_width": 8.0,
"scroll_y_pitch": 60,
"snap_resolution": "1/16",
"note_selection": []
}
}
```
#### 3.2 Piano Roll View Canvas Layout & Interaction Spec
* **Top Navigation Rule Pane (Bars/Beats Bar):**
* Displays bars from $0$ to $N$ (where $N = \text{total\_buffer\_bars}$, e.g., 8 bars).
* Highlights active clip visibility bounds (e.g., Bar 4.0 to Bar 6.0 shaded with active overlay, exterior bars dimmed).
* **Left Piano Keybed:**
* Anchored vertically, spans pitches $0$ (C-1) through $127$ (G9).
* Draws standard 88 key / 128 key pattern with distinct black key visually offset bars and pitch labeling ($C3$, $C4$, etc.).
* **Note Grid Canvas (Right Pane):**
* Synced to vertical pitch scroll and horizontal beat zoom.
* **Row Background Rendering:** Black key rows are assigned darker background fill color `#1A1A1E`, white key rows use `#25252A`.
* **Snap Grid Lines:** Rendered dynamically based on selected snap mode: Free, 1/1 Bar, 1/2 Beat, 1/4 Beat, 1/8 Beat, 1/16 Beat, 1/32 Beat.
* **Bottom Controller Pane (CC / Velocity / Pan Lane):**
* Synchronized horizontally with note grid.
* Displays vertical stem bars per note representing properties (Velocity, Pan). Allows click-and-drag line shaping or direct stem adjustment.
---
### 4. Audio & Synth Engine Routing Architecture (Web Audio API)
#### 4.1 Real-Time Signal Flow Graph
```text
[MIDI Scheduler] ---> [AudioWorklet / Virtual Synth Engine]
|
v (Audio Buffer / Stream)
[Audio Sample Playback Node] ----> [Track Channel FX Chain]
|
v
[Track Gain / Pan Node]
|
v
+---------------------+---------------------+
| |
v (If inside Section) v (If Direct Track)
[Section Sub-Mix Bus] [Main Master Mixer Bus]
| |
+-------------------->----------------------+
|
v
[Web Audio Destination]
```
#### 4.2 Web Audio Node Architecture Specifications
* **AudioTrack Node Structure:**
```javascript
TrackAudioGraph = {
inputNode: GainNode,
fxChain: [ BiquadFilterNode, DelayNode, ConvolverNode ],
panNode: StereoPannerNode,
outputGainNode: GainNode,
connect(destination) { ... }
}
```
* **Section Bus Graph Routing:**
* Each Section in Section-tab Store instantiates an intermediate `GainNode` sub-mixer (`SectionBus`).
* Tracks within the Section connect their final outputs to `SectionBus`.
* When a `SectionItem` is placed on a Main Session track, the `SectionBus` output is routed into the Main Session track's input node, preserving non-destructive DSP processing hierarchies.
---
### 5. Core Mathematical & Technical Algorithms
#### 5.1 Algorithm 1: Non-Destructive Item Slicing & Offset Playback Math
##### Mathematical Formulation
Let:
* $T_{\text{global}}$ = Current global playback time in seconds on the main timeline.
* $\text{BPM}$ = Beats Per Minute of the project.
* $\text{TS}_{\text{num}}$ = Time Signature Numerator (e.g., 4 beats per bar).
* $S_{\text{item}}$ = Item start position in global bars ($\text{start\_bar}$).
* $L_{\text{item}}$ = Item visible length on timeline in bars ($\text{duration\_bars}$).
* $O_{\text{item}}$ = Source internal start offset in bars ($\text{clip\_start\_offset\_bars}$).
Bar to Time Conversion Factor:
$$\text{SecondsPerBeat} = \frac{60.0}{\text{BPM}}$$
$$\text{SecondsPerBar} = \text{SecondsPerBeat} \times \text{TS}_{\text{num}}$$
Item Global Time Bounds:
$$T_{\text{start}} = S_{\text{item}} \times \text{SecondsPerBar}$$
$$T_{\text{end}} = (S_{\text{item}} + L_{\text{item}}) \times \text{SecondsPerBar}$$
Active Playback Slicing Condition: An item is active if and only if:
$$T_{\text{start}} \le T_{\text{global}} < T_{\text{end}}$$
Local Item Buffer Time Mapping ($T_{\text{local}}$): When $T_{\text{global}}$ falls within $[T_{\text{start}}, T_{\text{end}}]$, the corresponding time $T_{\text{local\_bars}}$ relative to the internal source clip buffer (0 to $\text{BufferLength}$) is:
$$T_{\text{local\_bars}} = \frac{T_{\text{global}} - T_{\text{start}}}{\text{SecondsPerBar}} + O_{\text{item}}$$
MIDI Note Slicing & Filtering Rule: For a MIDI note $N$ inside the item source with start beat $N_{\text{start\_beat}}$ and length $N_{\text{dur\_beat}}$ (converted to internal bar metric $N_{\text{bar\_start}} = \frac{N_{\text{start\_beat}}}{\text{TS}_{\text{num}}}$, $N_{\text{bar\_dur}} = \frac{N_{\text{dur\_beat}}}{\text{TS}_{\text{num}}}$):
The note is triggered during main playback if and only if:
$$N_{\text{bar\_start}} \ge O_{\text{item}} \quad \text{AND} \quad N_{\text{bar\_start}} < (O_{\text{item}} + L_{\text{item}})$$
##### Pseudocode Implementation
```javascript
function getActiveMIDINotesForPlayback(item, currentGlobalBar, timeSigNum) {
const itemStartBar = item.start_bar;
const itemEndBar = item.start_bar + item.duration_bars;
const offsetBar = item.clip_start_offset_bars;
// Check if playback cursor is inside visible item clip
if (currentGlobalBar < itemStartBar || currentGlobalBar >= itemEndBar) {
return []; // Item inactive
}
const activeNotes = [];
const internalWindowStartBar = offsetBar;
const internalWindowEndBar = offsetBar + item.duration_bars;
for (const note of item.source_data.notes) {
const noteStartBar = note.start_beat / timeSigNum;
const noteEndBar = noteStartBar + (note.duration_beats / timeSigNum);
// Filter notes outside the non-destructive visible window
if (noteStartBar >= internalWindowStartBar && noteStartBar < internalWindowEndBar) {
// Calculate playback time relative to global session
const relativeBarInItem = noteStartBar - internalWindowStartBar;
const targetGlobalBar = itemStartBar + relativeBarInItem;
activeNotes.push({
note: note,
scheduledGlobalBar: targetGlobalBar
});
}
}
return activeNotes;
}
```
#### 5.2 Algorithm 2: Dynamic Section Length Calculation Algorithm
When `auto_compute_length` is enabled for a Section, its total duration in bars $L_{\text{section}}$ is dynamically evaluated from the boundary bounds of all child items across all tracks inside that Section.
##### Mathematical Formulation
Let $T$ be the set of tracks in the section, and $I(t)$ be the set of items in track $t$.
$$L_{\text{section}} = \max_{t \in T} \left( \max_{i \in I(t)} \left( i.\text{start\_bar} + i.\text{duration\_bars} \right) \right)$$
If $I(t)$ is empty for all $t$, then $L_{\text{section}} = 4.0$ (default baseline minimum).
##### Implementation Architecture
```javascript
function recomputeSectionLength(sectionContainer) {
if (!sectionContainer.auto_compute_length) {
return sectionContainer.length_bars;
}
let maxEndBar = 0.0;
for (const track of sectionContainer.tracks) {
for (const item of track.items) {
const itemEndBar = item.start_bar + item.duration_bars;
if (itemEndBar > maxEndBar) {
maxEndBar = itemEndBar;
}
}
}
// Enforce baseline grid quantization rounding (e.g. minimum 1 bar)
const computedLength = Math.max(1.0, Math.ceil(maxEndBar));
sectionContainer.length_bars = computedLength;
return computedLength;
}
```
#### 5.3 Algorithm 3: Piano Roll Grid Mapping & Quantization Math
##### Grid Coordinate Transformation Formulae
Let:
* $X_{\text{px}}$ = Pixel X-coordinate on Piano Roll Canvas.
* $Y_{\text{px}}$ = Pixel Y-coordinate on Piano Roll Canvas.
* $\text{Zoom}_x$ = Pixels per Beat.
* $\text{NoteHeight}$ = Height in pixels per pitch key row (e.g., 18px).
* $\text{Scroll}_x$ = Horizontal scroll offset in beats.
* $\text{Scroll}_y$ = Vertical scroll top note pitch (e.g., pitch 127 down to 0).
Beat to Canvas Pixel Conversion:
$$X_{\text{px}} = (\text{Beat} - \text{Scroll}_x) \times \text{Zoom}_x$$
$$\text{Beat} = \frac{X_{\text{px}}}{\text{Zoom}_x} + \text{Scroll}_x$$
Pitch to Canvas Pixel Conversion:
$$Y_{\text{px}} = (127 - \text{Pitch} - \text{Scroll}_y) \times \text{NoteHeight}$$
$$\text{Pitch} = 127 - \left\lfloor \frac{Y_{\text{px}}}{\text{NoteHeight}} \right\rfloor - \text{Scroll}_y$$
##### Quantization (Snap To Grid) Math
Let $Q$ be the snap unit in beats (e.g., $1/4 \text{ bar} = 1.0 \text{ beat}$, $1/16 \text{ note} = 0.25 \text{ beat}$). Given raw unquantized beat $B_{\text{raw}}$:
$$B_{\text{quantized}} = \text{round}\left(\frac{B_{\text{raw}}}{Q}\right) \times Q$$
#### 5.4 Algorithm 4: Tab Close Dependency & Lifecycle Validation Algorithm
This algorithm validates whether a tab close request can be fulfilled, enforcing the fixed Main Session constraint and preventing parent Section tab closures while child editor sub-tabs remain active.
```javascript
function requestCloseTab(tabIdToClose, stateStore) {
const targetTab = stateStore.open_tabs.find(tab => tab.tab_id === tabIdToClose);
if (!targetTab) {
return { success: false, reason: "TAB_NOT_FOUND" };
}
// 1. Rule: Main Session cannot be closed
if (!targetTab.is_closeable || targetTab.type === 'MAIN_SESSION') {
return { success: false, reason: "CANNOT_CLOSE_MAIN_SESSION" };
}
// 2. Rule: Section Tab cannot be closed if child tabs are active
if (targetTab.type === 'SECTION_TAB') {
const activeChildTabs = stateStore.open_tabs.filter(
tab => tab.parent_tab_id === targetTab.tab_id
);
if (activeChildTabs.length > 0) {
return {
success: false,
reason: "SECTION_HAS_ACTIVE_CHILD_EDITORS",
activeChildTabs: activeChildTabs.map(t => ({ id: t.tab_id, title: t.title }))
};
}
}
// 3. Execution: Perform clean tab shutdown and update active context
const updatedTabs = stateStore.open_tabs.filter(tab => tab.tab_id !== tabIdToClose);
// Fallback active tab selection if current active tab is being closed
let nextActiveTabId = stateStore.active_tab_id;
if (stateStore.active_tab_id === tabIdToClose) {
// Fallback to parent tab, or default to main session (index 0)
nextActiveTabId = targetTab.parent_tab_id || updatedTabs[0].tab_id;
}
stateStore.open_tabs = updatedTabs;
stateStore.active_tab_id = nextActiveTabId;
return { success: true, nextActiveTabId: nextActiveTabId };
}
```
#### 5.5 Algorithm 5: Sample-Accurate Lookahead MIDI & Audio Scheduler
Web Audio API timing operates on a high-precision hardware audio clock (`audioContext.currentTime`). JavaScript timers (`setTimeout`/`setInterval`) lack frame accuracy. The Lookahead Scheduler combines JS interval ticks with Web Audio precision scheduling.
```text
Lookahead Window (e.g. 100ms)
|-------------------------------------------|
| AudioContext Time: 10.0s |
| Schedule horizon: 10.1s |
| |
| [Event 1 @ 10.02s] -> Scheduled in WebAudio
| [Event 2 @ 10.08s] -> Scheduled in WebAudio
|___________________________________________|
```
##### Scheduler Specification
```javascript
class PrecisionAudioScheduler {
constructor(audioCtx, lookaheadMs = 25.0, scheduleAheadTimeSec = 0.1) {
this.audioCtx = audioCtx;
this.lookaheadMs = lookaheadMs; // Frequency of timer evaluation
this.scheduleAheadTime = scheduleAheadTimeSec; // How far ahead to queue WebAudio events
this.nextNoteBeat = 0.0;
this.currentBeat = 0.0;
this.bpm = 120.0;
this.timerId = null;
}
beatToTime(beat) {
const secondsPerBeat = 60.0 / this.bpm;
return beat * secondsPerBeat;
}
timeToBeat(timeSec) {
const secondsPerBeat = 60.0 / this.bpm;
return timeSec / secondsPerBeat;
}
schedulerTick(activeSession) {
const currentTime = this.audioCtx.currentTime;
const horizonTime = currentTime + this.scheduleAheadTime;
// Traverse session items and find notes falling within [currentTime, horizonTime]
const pendingEvents = activeSession.getEventsInTimeRange(
this.timeToBeat(currentTime),
this.timeToBeat(horizonTime)
);
for (const evt of pendingEvents) {
if (!evt.scheduled) {
const preciseAudioTime = currentTime + this.beatToTime(evt.targetBeat - this.currentBeat);
this.triggerWebAudioEvent(evt, preciseAudioTime);
evt.scheduled = true;
}
}
}
triggerWebAudioEvent(evt, exactAudioTime) {
if (evt.type === 'MIDI_NOTE_ON') {
const synthNode = evt.trackSynthNode;
synthNode.noteOn(evt.note.pitch, evt.note.velocity, exactAudioTime);
synthNode.noteOff(evt.note.pitch, exactAudioTime + this.beatToTime(evt.note.duration_beats));
} else if (evt.type === 'AUDIO_CLIP') {
const sourceNode = this.audioCtx.createBufferSource();
sourceNode.buffer = evt.audioBuffer;
sourceNode.connect(evt.trackGainNode);
sourceNode.start(exactAudioTime, evt.offsetSec, evt.durationSec);
}
}
start(session) {
this.timerId = setInterval(() => this.schedulerTick(session), this.lookaheadMs);
}
stop() {
if (this.timerId) clearInterval(this.timerId);
}
}
```
#### 5.6 Algorithm 6: Playhead UI Rendering Sync Loop
UI Playhead rendering uses `requestAnimationFrame` and queries `audioContext.currentTime` directly to prevent visual jitter or lag.
$$\text{Current Beat UI} = \frac{\text{audioCtx.currentTime} - \text{PlaybackStartTimeSec}}{\text{SecondsPerBeat}}$$
$$\text{Pixel Position X} = (\text{Current Beat UI} - \text{ViewportStartBeat}) \times \text{Zoom}_x$$
---
### 6. Backend Python Server Architecture & Offline Render Spec
#### 6.1 Server Architecture Framework
* **Framework:** FastAPI with Async WebSocket endpoints for real-time state synchronization.
* **DSP Engine:** `pedalboard` (Spotify's Python Audio Processing Library) and `numpy` for multi-track mixing, high-quality audio resampling, and plugin hosting.
#### 6.2 Python Offline Stem Bouncing Engine Specification (`render_engine.py`)
```python
import numpy as np
from pedalboard import Pedalboard, Gain, Reverb, Compressor
import soundfile as sf
class PythonRenderEngine:
def __init__(self, sample_rate=44100):
self.sample_rate = sample_rate
def bars_to_samples(self, bars: float, bpm: float, time_sig_num: int) -> int:
seconds_per_beat = 60.0 / bpm
seconds_per_bar = seconds_per_beat * time_sig_num
return int(bars * seconds_per_bar * self.sample_rate)
def render_project(self, project_json: dict, output_filepath: str):
bpm = project_json["metadata"]["bpm"]
time_sig_num = project_json["metadata"]["time_signature_numerator"]
main_session = project_json["main_session"]
# 1. Compute total project samples
total_bars = main_session.get("length_bars", 16.0)
total_samples = self.bars_to_samples(total_bars, bpm, time_sig_num)
# Stereo Master Buffer
master_buffer = np.zeros((2, total_samples), dtype=np.float32)
# 2. Iterate and process main tracks
for track in main_session["tracks"]:
track_type = track["type"]
track_buffer = np.zeros((2, total_samples), dtype=np.float32)
for item in track["items"]:
start_sample = self.bars_to_samples(item["start_bar"], bpm, time_sig_num)
dur_samples = self.bars_to_samples(item["duration_bars"], bpm, time_sig_num)
offset_sample = self.bars_to_samples(item["clip_start_offset_bars"], bpm, time_sig_num)
if item["type"] == "AUDIO_ITEM":
# Load audio source sample array
audio_data, sr = sf.read(item["source_data"]["audio_file_url"], dtype='float32')
audio_data = audio_data.T # Shape: (channels, samples)
# Apply non-destructive trimming offset
sliced_audio = audio_data[:, offset_sample : offset_sample + dur_samples]
# Accumulate into track buffer with bounds checks
end_sample = min(start_sample + sliced_audio.shape[1], total_samples)
actual_len = end_sample - start_sample
track_buffer[:, start_sample:end_sample] += sliced_audio[:, :actual_len]
# Apply Track Gain and FX Chain via Pedalboard
board = Pedalboard([Gain(gain_db=track.get("volume_db", 0.0))])
processed_track = board(track_buffer, sample_rate=self.sample_rate)
# Mix down to Master
master_buffer += processed_track
# 3. Write final output file
sf.write(output_filepath, master_buffer.T, self.sample_rate)
return output_filepath
```
---
### 7. Execution Context & Sub-Tab Lifecycle Matrix
| Context Tab Type | Scope Identifier | View Boundaries | Is Closeable | Close Dependency Conditions | Audio Routing Target |
| --- | --- | --- | --- | --- | --- |
| **MAIN SESSION** | Root | Full Master Timeline ($0 \to N$ Bars) | No | Pinned permanently; cannot be closed | WebAudio Hardware Destination |
| **SECTION TAB** | Section_ID | Dynamic Section Bounds ($0 \to L_{\text{section}}$) | Yes | Blocked if any child editor sub-tabs are open | Target Section Bus Gain Node |
| **PIANO ROLL** | MIDIItem_ID | Item Source Length Bounds ($0 \to N_{\text{buffer}}$) | Yes | Can close freely; notifies parent Section tab | Track Instrument Synth Engine |
| **SAMPLE EDITOR** | AudioItem_ID | Sample Buffer Waveform ($0 \to T_{\text{sample}}$) | Yes | Can close freely; notifies parent Section tab | Track Audio Node Router |
---
### 8. Summary of Non-Destructive Slice & Tab Lifecycle Validation
* **Tab Close Prevention Test:**
1. `MAIN SESSION` close request is rejected immediately (`CANNOT_CLOSE_MAIN_SESSION`).
2. `Section_01` tab has an active child editor tab (`Piano Roll: Bassline`).
3. Request to close `Section_01` tab returns `SECTION_HAS_ACTIVE_CHILD_EDITORS`.
4. User closes `Piano Roll: Bassline` tab first.
5. Subsequent close request for `Section_01` succeeds and cleans up UI context.
* **8-Bar Source with 2-Bar Visible Crop Test:**
1. Given `MIDIItem` length = 8 bars ($0 \dots 8$).
2. User drags left boundary to Bar 4 and right boundary to Bar 6.
3. `start_bar = 4.0` (Global Session Placement), `duration_bars = 2.0`, `clip_start_offset_bars = 4.0`.
4. Transport reaches global Bar 4.0 $\to$ scheduler evaluates internal bounds $[4.0, 6.0)$ and triggers only visible notes while preserving complete 8-bar non-destructive source.
-361
View File
@@ -1,361 +0,0 @@
Here is the conversion of the document into a professional English Markdown format:
# ARCHITECTURAL, TECHNICAL, AND ALGORITHMIC SPECIFICATION
## Sub-Session System, Section Arrangement & Piano Roll Tab (Hybrid DAW)
This document details the technical solution for building a Hierarchical DAW Engine. This architecture enables nesting Sub-Sessions (Sections) inside the Main Session, alongside a Sub-Tab Editor system (including Piano Roll and Audio Sample Editor) to precisely edit MIDI and Audio Items.
---
### 0. Non-Breaking Modular Principles (Integration & Backward Compatibility)
To guarantee that new features do not disrupt the DAW's existing core logic and codebase, the entire extension architecture is designed according to these principles:
* **Extensibility & Encapsulation:**
* The current Session architecture serves directly as the **Project Root / Main Session**.
* `SectionItem`, `ItemMIDI`, and `ItemAudio` operate as **Polymorphic Item Types** inheriting from the existing base `Item` class/interface. Existing Item logic (e.g., drag-and-drop, timeline trimming) remains $100\%$ untouched.
* **Decoupled State Pipeline:**
* The logic governing the Playhead, Transport controls (Play/Pause/Stop), and the global Audio Context of the Main Session remains unmodified.
* **Nested Time Mapping** acts solely as an intermediate Transformation Layer when passing time coordinates down into Sub-Sessions. It does not overwrite or mutate the beat synchronization loop of the Main Timeline.
* **Plugin Style Architecture (Audio & MIDI Engine):**
* Synth Tracks, Audio Clip Processors, and Sub-Session Sub-Mix Buses plug into the existing AudioNode Graph as auxiliary nodes. They route directly back to the current Master Node without breaking pre-established Gain/Pan/FX pipelines.
---
### 1. Hierarchical Data Model
To support embedding Sessions within Sessions as well as isolated Clip/Sample-level editing, the data state model expands into an encapsulated Tree Graph structure.
```text
Project Root
├── Main Session (Root Session - Current Session Structure)
│ ├── Track 01 (Audio Track)
│ │ └── ItemAudio: "Vocals.wav" ──► [Opens Audio Sample Editor Sub-Tab]
│ ├── Track 02 (MIDI Track + Synth Engine)
│ │ └── ItemMIDI: "Melody_Main" ──► [Opens Piano Roll Sub-Tab]
│ └── Track 03 (Section Track - New Track Type)
│ └── Item: Section_A (Referencing SubSession_01)
├── Sub-Sessions Store (Auxiliary Memory Registry)
│ ├── SubSession_01 ("Verse 1")
│ │ ├── Computed Length: Dynamic Bars (Auto-calculated from longest Item)
│ │ ├── Track 1.1 (Audio Track)
│ │ │ └── ItemAudio: "Guitar_Riff.wav" ──► [Opens Audio Sample Editor Sub-Tab]
│ │ └── Track 1.2 (MIDI Track)
│ │ └── ItemMIDI: "Bassline" ─────────► [Opens Piano Roll Sub-Tab]
│ └── SubSession_02 ("Chorus")
└── Active Editor Views / Sub-Tabs (Isolated Editing Contexts)
├── Audio Sample Editor Sub-Tab (Edits Audio Clips from Main Session or Sub-Session)
└── Piano Roll Sub-Tab (Edits MIDI Items from Main Session or Sub-Session)
```
#### Detailed Data Schemas (JSON Specs)
**a. Schema: `NoteMIDI**`
```typescript
interface NoteMIDI {
id: string;
pitch: number; // 0 - 127 (Midi Note Number, e.g., 60 = C4)
startTick: number; // Time coordinate based on Pulses Per Quarter note (PPQ, e.g., 960 PPQ)
durationTicks: number;
velocity: number; // 0 - 127
selected?: boolean;
}
```
**b. Schema: `ItemMIDI` (Belongs to MIDI Track - Inherits from Base Item)**
```typescript
interface ItemMIDI {
id: string;
type: 'MIDI';
name: string;
parentSessionId: string; // Target Session ID (Main or Sub-Session)
startBar: number; // Start position on the Timeline (Bar)
lengthBars: number; // Item duration in Bars
offsetTick: number; // Internal trim offset
notes: NoteMIDI[]; // Array tracking MIDI Notes
}
```
**c. Schema: `ItemAudio` (Belongs to Audio Track - Inherits from Base Item)**
```typescript
interface ItemAudio {
id: string;
type: 'AUDIO';
name: string;
parentSessionId: string; // Target Session ID (Main or Sub-Session)
startBar: number;
lengthBars: number;
samplePath: string; // Audio file path or Buffer Key
sampleOffsetSec: number; // Playback start point offset (Trim In)
gain: number; // Clip Gain
pitchShiftSemi: number; // Pitch Shift (Semitones)
}
```
**d. Schema: `SectionItem` (Represents a Sub-Session inside the Main Session)**
```typescript
interface SectionItem {
id: string;
type: 'SECTION';
subSessionId: string; // Reference ID pointing to SubSession inside Memory Store
name: string;
startBar: number;
lengthBars: number; // Defaults to SubSession.computedLengthBars unless trimmed/cropped
loop: boolean; // Enables repetition if lengthBars > SubSession.computedLengthBars
}
```
**e. Schema: `Session` (Unified structure for both Main Session and Sub-Session)**
```typescript
interface Session {
id: string;
name: string;
isMain: boolean;
timeSignature: [number, number]; // e.g., [4, 4]
bpm: number;
tracks: Track[];
// Dynamically calculated derived state; never assigned manually
get computedLengthBars(): number;
}
```
---
### 2. Audio & Synth Engine Routing Architecture (Web Audio API)
For MIDI tracks to output audio, each is bound to an Instrument/Synth Instance. When a Section is placed onto the Main Session, all audio generated by its child tracks is bussed directly into the existing Gain/Pan matrix.
#### Audio Node Graph Diagram
```text
[MIDI Items] ──(Triggers)──► [Synth Engine / Soundfont / WebAssembly VSTi]
[Audio Items] ──(Buffer Source)──────────┤
[Track Gain / Pan Node]
[Sub-Session Sub-Mix Bus Node]
┌──────────────────────┴──────────────────────┐
▼ ▼
[Main Session Audio Graph] [Solo / Mute Logic]
(Current Audio Processing Logic)
[Master Destination]
```
**Instrument Engine Processing Logic for MIDI Tracks:**
* **Virtual Instrument Binding:** Every MIDI Track instantiates a synthesis `AudioNode` (e.g., Web Audio API Soundfont Player, WebSynth JS, or WASM Synthesizer).
* **Dynamic Polyphony Engine:** As playback scans across MIDI Notes, the system triggers `noteOn(pitch, velocity, time)` and `noteOff(pitch, time)` events. These are scheduled ahead of time ($100\text{ms} - 200\text{ms}$ Lookahead) via the `AudioContext.currentTime` clock.
---
### 3. Tab UI Management & Event Processing Flow (Tab Navigation Stack)
The graphical interface expands on a Tab Manager & Navigation Stack model to handle isolated views (Views/Sub-tabs) for specific data entities.
```text
[ Tabs Bar ] ── [ Main Session ] │ [ Sub-Session: Verse 1 ] │ [ Piano Roll: Bassline ] │ [ Sample Edit: Vocals.wav ]
```
#### Interaction & Navigation Mechanics:
* **Opening a Sub-Session Tab:**
* *Action:* User double-clicks a `SectionItem` on a Main Track.
* *Result:*
* Instantiates a new Tab using `ID = SubSession.id`.
* Maps the Timeline Viewport rendering context to the SubSession.
* Enables adding, editing, or deleting child tracks (Audio & MIDI) within the Sub-Session boundary.
* **Opening the Piano Roll Sub-Tab:**
* *Action:* User double-clicks an `ItemMIDI` inside the Main Session OR a Sub-Session.
* *Result:*
* Instantiates a Sub-tab labeled: `Piano Roll - [Item Name]`.
* Caches context references: `{ itemId, parentSessionId }`.
* Passes the `ItemMIDI.notes` array directly into the Canvas/Piano Roll Grid.
* Any add/edit/delete actions executed on notes inside the Piano Roll instantly update the native `ItemMIDI` in the target Session via Mutable/Immutable References.
* **Opening the Audio Sample Editor Sub-Tab (Session Edit Audio Sample):**
* *Action:* User double-clicks OR right-clicks and selects "Edit" on an `ItemAudio` inside the Main Session or a Sub-Session.
* *Result:*
* Instantiates a Sub-tab labeled: `Audio Editor - [Clip Name]`.
* Loads the high-resolution Waveform of the target `ItemAudio` onto the sample editing Viewport.
* Provides access to tools: Trim start/end, Normalized Peak, Pitch Shift, Reverse, Fade In/Out, or DSP slicing.
* When clicking *Save / Apply Changes*: The system updates the `ItemAudio` attributes (or dispatches a DSP processing request to the Python Server for heavy tasks) and forces a visual refresh of the Clip on the Main Session / Sub-Session timeline.
#### Data Persistence & Dynamic Sub-Session Length Updates:
* Because JavaScript handles array/object data passing by **Reference**, modifications made to Notes in the Piano Roll Tab or Clips in the Audio Editor directly update the origin State of the corresponding Session.
* Any add/remove/move/stretch operation targeting an Item inside a Sub-session will immediately trigger the **Dynamic Length Recalculation** algorithm to update the temporal boundary of the Sub-Session.
---
### 4. Core Algorithms
#### Algorithm 1: Dynamic Sub-Session Length Calculation
Sub-sessions do not enforce rigid length constraints. Instead, they dynamically map their duration ($L_{\text{bars}}$) to match the furthest end-point of all encapsulated Items.
**Formula:**
Given a Sub-Session containing a list of $T$ tracks, where each track $t$ holds a list of $I_t$ items (Audio, MIDI, etc.):
$$\text{ItemEndBar}(item) = item.\text{startBar} + item.\text{lengthBars}$$
$$L_{\text{bars}} = \max_{t \in T} \left( \max_{i \in I_t} (\text{ItemEndBar}(i)) \right)$$
*If the Sub-session is entirely empty (contains no Items), $L_{\text{bars}}$ defaults to $1$ Bar (or the default duration of a single grid bar).*
```javascript
function calculateSubSessionLength(subSession) {
let maxEndBar = 1; // Minimum duration fallback for empty sub-sessions
for (const track of subSession.tracks) {
for (const item of track.items) {
const itemEndBar = item.startBar + item.lengthBars;
if (itemEndBar > maxEndBar) {
maxEndBar = itemEndBar;
}
}
}
return maxEndBar;
}
```
#### Algorithm 2: Nested Time Mapping
When the Main Session Playhead tracks time $T_{\text{main}}$ (seconds), the engine must calculate the relative time coordinate $T_{\text{sub}}$ inside the active Sub-Session.
**Formula:**
Assume:
* $S_{\text{bar}}$: The starting Bar of the Section Item on the Main Timeline.
* $L_{\text{bars}}$: The dynamically evaluated length of the root Sub-Session ($L_{\text{bars}} = \text{calculateSubSessionLength}(\text{SubSession})$).
* $BPM$: Beats Per Minute.
* $TimeSig$: Beats per Bar (e.g., 4 beats).
$$\text{SecondsPerBar} = \frac{60}{\text{BPM}} \times \text{TimeSig}$$
$$\text{OffsetSeconds} = (T_{\text{main}} - (S_{\text{bar}} - 1) \times \text{SecondsPerBar})$$
If `SectionItem.loop = true`:
$$T_{\text{sub}} = \text{OffsetSeconds} \pmod{L_{\text{bars}} \times \text{SecondsPerBar}}$$
If `SectionItem.loop = false`:
$$T_{\text{sub}} = \begin{cases} \text{OffsetSeconds} & \text{if } 0 \le \text{OffsetSeconds} \le (L_{\text{bars}} \times \text{SecondsPerBar}) \\ \text{undefined} & \text{if out of bounds} \end{cases}$$
#### Algorithm 3: Lookahead MIDI Scheduler
JavaScript's `setInterval` function lacks the temporal precision required for audio playback. We employ the **Web Audio Lookahead Scheduler** algorithm combined with Ticks $\rightarrow$ Seconds translation.
```javascript
const PPQ = 960; // 960 Pulses Per Quarter note (Standard MIDI resolution)
let nextNoteIndex = 0;
const scheduleAheadTime = 0.2; // 200ms Lookahead buffer
const lookaheadMs = 25; // Polling interval interval block (25ms)
function ticksToSeconds(ticks, bpm) {
const secondsPerQuarterNote = 60.0 / bpm;
return (ticks / PPQ) * secondsPerQuarterNote;
}
function scheduler(midiItem, audioCtx, currentPlayheadTime) {
// Extract notes mapped within [currentPlayheadTime, currentPlayheadTime + scheduleAheadTime]
while (nextNoteIndex < midiItem.notes.length) {
const note = midiItem.notes[nextNoteIndex];
const noteStartTimeSec = ticksToSeconds(note.startTick, currentBpm);
if (noteStartTimeSec >= currentPlayheadTime + scheduleAheadTime) {
break; // Note start bounds exceed the active Lookahead window
}
if (noteStartTimeSec >= currentPlayheadTime) {
// Calculate absolute scheduling time against the AudioContext Clock
const audioCtxStartTime = audioCtx.currentTime + (noteStartTimeSec - currentPlayheadTime);
const durationSec = ticksToSeconds(note.durationTicks, currentBpm);
// Fire the VSTi/Synth Engine
trackSynthEngine.playNote(note.pitch, note.velocity, audioCtxStartTime, durationSec);
}
nextNoteIndex++;
}
}
```
#### Algorithm 4: Grid Snapping & Quantization (Piano Roll)
When adding or dragging a MIDI Note in the Piano Roll Tab, the $X$ coordinate of the mouse cursor must snap to the nearest rhythmic grid boundary (1/4, 1/8, 1/16, 1/32 Note).
```javascript
function snapTickToGrid(rawTick, gridFraction, ppq) {
// gridFraction: 0.25 (1/4 note), 0.125 (1/8 note), 0.0625 (1/16 note)
const ticksPerGridStep = ppq * (gridFraction * 4);
// Snap rounding formula targeting the nearest grid boundary
const snappedTick = Math.round(rawTick / ticksPerGridStep) * ticksPerGridStep;
return Math.max(0, snappedTick);
}
```
---
### 5. Performance Optimization
* **Virtual Rendering for Piano Roll, Audio Sample Editor & Main Session:**
* Never render the entire array of MIDI Notes or total Audio Waveforms simultaneously into the HTML DOM.
* Mandatory use of HTML5 Canvas 2D / WebGL paired with **Virtual Viewport Rendering** (only drawing Notes/Samples situated within the active Viewport Rect boundary).
* **Audio Bouncing / Freezing (For Heavy Sections):**
* If a Sub-Session houses too many Tracks and VSTi plugins, causing CPU bottlenecks during Main Session playback:
* Enable the **"Freeze Section"** action: The Python backend processes the request, rendering that entire Sub-Session block into a single temporary Audio WAV file (Bounce to Disk).
* The Main Session then only processes one discrete Audio file instead of simultaneously calculating dozens of child tracks.
* **Immutable State & Undo/Redo Engine:**
* Project State management is handled via the Redux/Zustand pattern model.
* Every add/edit/delete operation applied to Notes on the Piano Roll or edits made to Audio Clips generates an Action that pushes to the `UndoStack`, supporting seamless `Ctrl + Z` shortcuts across every Sub-tab context.
+426
View File
@@ -0,0 +1,426 @@
# TECHNICAL SPECIFICATION: CLIENT-SIDE REAL-TIME RECORDING ENGINE
## Browser-Based Microphone & Hardware MIDI Keyboard Recording Module
---
## 1. System Overview
The Client-side recording module enables the DAW to capture live audio signals directly from Microphone/Line-in interfaces (via the Web MediaDevices API) and keypress events from Hardware MIDI Keyboards/Controllers (via the Web MIDI API) in real time. The module operates with low latency and includes hardware latency compensation.
```text
+-----------------------------------------------------------------------------------+
| CLIENT BROWSER |
| |
| +-------------------------+ +-----------------------------+ |
| | Hardware MIDI Keyboard | | Live Microphone / Line-In | |
| +------------+------------+ +--------------+--------------+ |
| | | |
| v (Web MIDI API) v (getUserMedia) |
| +------------+------------+ +--------------+--------------+ |
| | Web MIDI Input Handler | | MediaStreamAudioSourceNode | |
| +------------+------------+ +--------------+--------------+ |
| | | |
| +------------------+ | |
| | | v |
| v v +--------------+--------------+ |
| +------------+-----+ +---------+-----------+ | Track Input Gain Node | |
| | Event Clock / | | WebAudio Virtual | +--------------+--------------+ |
| | Latency Engine | | Synth Engine | | |
| +------------+-----+ +---------+-----------+ +--------+--------+ |
| | | | | |
| v v v v |
| +------------+-----+ (Live Sound) +-------+-------+ +-------+------+ |
| | Recorded MIDI | | AudioWorklet | | Monitoring | |
| | Buffer | | Ring-Buffer | | Switch | |
| +------------+-----+ | Recorder | +-------+------+ |
| | +-------+-------+ | |
| v | v |
| [ Timeline MIDI ] v [ Master Mix ] |
| [ Item Creation ] +-------+-------+ |
| | Float32 PCM | |
| | Audio Buffer | |
| +-------+-------+ |
| | |
| v |
| [ Timeline Audio] |
| [ Item Creation ] |
+-----------------------------------------------------------------------------------+
```
---
## 2. Hardware I/O & API Contracts
### 2.1 MediaDevices (Microphone Capture)
**Permission Request:** Uses `navigator.mediaDevices.getUserMedia` configured to disable automatic browser processing DSP algorithms to capture pure, unprocessed audio signals:
```javascript
const audioConstraints = {
audio: {
deviceId: selectedDeviceId ? { exact: selectedDeviceId } : undefined,
echoCancellation: false, // Disables echo cancellation to prevent instrument sound distortion
noiseSuppression: false, // Disables automatic noise suppression to preserve full frequency range
autoGainControl: false, // Disables Automatic Gain Control (AGC)
latency: 0 // Requests minimal latency from OS audio driver
}
};
```
### 2.2 Web MIDI API Integration
**Device Enumeration & Listener Assignment:**
* Uses `navigator.requestMIDIAccess({ sysex: false })` to scan for USB-connected keyboard devices.
* **Timestamp Precision:** Obtains event timestamps from `MIDIMessageEvent.timeStamp` (as a `DOMHighResTimeStamp` in microseconds) and synchronizes them with `AudioContext.currentTime`.
---
## 3. Recording Lifecycle & State Machine
```text
[IDLE] ───► (User Arms Track) ───► [ARMED] ───► (Press Rec + Play) ───► [COUNT-IN / PRE-ROLL]
|
[STOP & COMMIT] ◄─── (Press Stop) ◄─── [RECORDING IN PROGRESS] ◄───────────────+
```
* **Arming Phase (Record Enable):**
* The user selects an input source and activates the Arm (R) button on the target track.
* Initializes the input level meter (VU Meter Canvas) to display input volume levels in real time.
* **Pre-Roll / Count-In Phase:**
* Transport triggers the metronome count-in (e.g., 1 Bar = 4 beats). The metronome plays click sounds based on project BPM.
* The engine does not write data to the Timeline yet, but begins reading the input buffer to prepare memory buffers.
* **Recording Phase:**
* Once the transport passes the Start Bar boundary, incoming MIDI key events or PCM Float32 audio samples are written into the active recording buffer memory.
* Canvas UI displays real-time visual feedback, rendering waveforms or MIDI note blocks dynamically.
* **Stop & Commit Phase:**
* Pressing Stop halts the recording process.
* Converts temporary memory buffers into a structured `MIDIItem` or `AudioItem`.
* Inserts the new Item onto the target track within the Main Session or Section tab.
---
## 4. Data Structures
### 4.1 Live MIDI Event Buffer Element Schema
```json
{
"type": "object",
"properties": {
"pitch": { "type": "integer", "minimum": 0, "maximum": 127 },
"start_beat": { "type": "number", "description": "Start position in beats on the timeline" },
"duration_beats": { "type": "number", "description": "Keypress duration in beats" },
"velocity": { "type": "number", "minimum": 0.0, "maximum": 1.0 },
"channel": { "type": "integer", "default": 0 }
}
}
```
### 4.2 Recording Track Input Configuration State
```json
{
"track_id": "track_midi_01",
"is_armed": true,
"monitoring_enabled": true,
"input_source": {
"device_type": "MIDI_KEYBOARD",
"device_id": "midi_input_usb_keyboard_0",
"channel": 1
},
"input_gain_db": 0.0,
"latency_offset_ms": 12.5
}
```
---
## 5. Core Algorithms & Latency Compensation
### 5.1 Algorithm 1: Hardware Latency Compensation Formula
When recording, the physical moment a key is pressed or sound enters the microphone is inherently delayed relative to speaker output due to input buffers ($L_{\text{input}}$) and output buffers ($L_{\text{output}}$).
#### Mathematical Formulation
Let:
* $T_{\text{audio\_ctx}}$ = Current timestamp in seconds on the `AudioContext` clock (`audioCtx.currentTime`).
* $T_{\text{rec\_start}}$ = Recording start timestamp in seconds.
* $\text{BPM}$ = Song tempo (Beats Per Minute).
* $\text{TS}_{\text{num}}$ = Time Signature Numerator (beats per bar).
* $\text{Bar}_{\text{start}}$ = Target timeline start bar for recording.
* $L_{\text{comp}}$ = Total hardware latency offset ($L_{\text{input}} + L_{\text{output}} + L_{\text{user\_offset}}$) in seconds.
**Actual Elapsed Audio Time ($T_{\text{elapsed}}$):**
$$T_{\text{elapsed}} = \max\left(0, T_{\text{audio\_ctx}} - T_{\text{rec\_start}} - L_{\text{comp}}\right)$$
**Audio Time to Beat Conversion ($\text{Beat}_{\text{current}}$):**
$$\text{SecondsPerBeat} = \frac{60.0}{\text{BPM}}$$
$$\text{Beat}_{\text{current}} = \frac{T_{\text{elapsed}}}{\text{SecondsPerBeat}} + \left(\text{Bar}_{\text{start}} \times \text{TS}_{\text{num}}\right)$$
**Timeline Placement Mapping:**
$$\text{StartBeat}_{\text{item}} = \text{Beat}_{\text{current}}$$
---
### 5.2 Algorithm 2: AudioWorklet PCM Ring-Buffer Processor
To prevent audio glitches or missing PCM frames when the browser's main thread is processing heavy UI renders, microphone recording runs inside an `AudioWorkletProcessor`:
```javascript
// public/processors/pcm-recorder-processor.js
class PCMRecorderProcessor extends AudioWorkletProcessor {
constructor() {
super();
this.bufferSize = 4096;
this.buffer = new Float32Array(this.bufferSize);
this.bufferIndex = 0;
}
process(inputs, outputs, parameters) {
const input = inputs[0];
if (input && input.length > 0) {
const inputChannel = input[0]; // Mono Channel 0
for (let i = 0; i < inputChannel.length; i++) {
this.buffer[this.bufferIndex++] = inputChannel[i];
// When Ring-Buffer fills, send Float32Array to Main Thread
if (this.bufferIndex >= this.bufferSize) {
this.port.postMessage({
type: 'PCM_DATA',
buffer: this.buffer.slice(0, this.bufferSize)
});
this.bufferIndex = 0;
}
}
}
return true; // Keep worklet active
}
}
registerProcessor('pcm-recorder-processor', PCMRecorderProcessor);
```
---
### 5.3 Algorithm 3: Client MIDIRecorder Class Implementation
```javascript
class ClientMIDIRecorder {
constructor(audioContext, bpm = 120, timeSigNumerator = 4) {
this.audioCtx = audioContext;
this.bpm = bpm;
this.timeSigNum = timeSigNumerator;
this.isRecording = false;
this.activeNotes = new Map(); // Store pitch -> { noteId, startBeat, velocity }
this.recordedNotes = [];
this.recStartAudioTime = 0.0;
this.recStartBar = 0.0;
// Compute round-trip browser latency
this.latencyCompSec = (this.audioCtx.baseLatency || 0) + (this.audioCtx.outputLatency || 0);
}
start(startBar = 0.0) {
this.isRecording = true;
this.recordedNotes = [];
this.activeNotes.clear();
this.recStartBar = startBar;
this.recStartAudioTime = this.audioCtx.currentTime;
this.bindMIDIInputs();
}
bindMIDIInputs() {
if (navigator.requestMIDIAccess) {
navigator.requestMIDIAccess().then(midiAccess => {
for (let input of midiAccess.inputs.values()) {
input.onmidimessage = (event) => this.handleMIDIMessage(event);
}
});
}
}
handleMIDIMessage(event) {
if (!this.isRecording) return;
const [status, pitch, velocity] = event.data;
const command = status >> 4;
// Apply latency compensation formula
const currentTimeSec = Math.max(0, this.audioCtx.currentTime - this.recStartAudioTime - this.latencyCompSec);
const secondsPerBeat = 60.0 / this.bpm;
const currentBeat = (currentTimeSec / secondsPerBeat) + (this.recStartBar * this.timeSigNum);
// Command 0x9: Note On
if (command === 0x9 && velocity > 0) {
const noteId = `rec_${Date.now()}_${pitch}`;
this.activeNotes.set(pitch, {
id: noteId,
pitch: pitch,
start_beat: currentBeat,
velocity: velocity / 127.0
});
}
// Command 0x8: Note Off (or Note On with velocity = 0)
else if (command === 0x8 || (command === 0x9 && velocity === 0)) {
if (this.activeNotes.has(pitch)) {
const note = this.activeNotes.get(pitch);
const durationBeats = Math.max(0.125, currentBeat - note.start_beat); // Min 1/32 note
this.recordedNotes.push({
id: note.id,
pitch: note.pitch,
start_beat: note.start_beat,
duration_beats: durationBeats,
velocity: note.velocity,
pan: 0.0
});
this.activeNotes.delete(pitch);
}
}
}
stop() {
this.isRecording = false;
// Flush remaining active keypresses when stop is triggered
const currentTimeSec = Math.max(0, this.audioCtx.currentTime - this.recStartAudioTime - this.latencyCompSec);
const currentBeat = (currentTimeSec / (60.0 / this.bpm)) + (this.recStartBar * this.timeSigNum);
for (let [pitch, note] of this.activeNotes.entries()) {
this.recordedNotes.push({
id: note.id,
pitch: note.pitch,
start_beat: note.start_beat,
duration_beats: Math.max(0.25, currentBeat - note.start_beat),
velocity: note.velocity,
pan: 0.0
});
}
this.activeNotes.clear();
return this.recordedNotes;
}
}
```
---
### 5.4 Algorithm 4: Client AudioRecorder & AudioBuffer Splicing Class Implementation
```javascript
class ClientAudioRecorder {
constructor(audioContext) {
this.audioCtx = audioContext;
this.mediaStream = null;
this.sourceNode = null;
this.workletNode = null;
this.pcmChunks = [];
this.isRecording = false;
}
async initializeInput(deviceId = null) {
const constraints = {
audio: {
deviceId: deviceId ? { exact: deviceId } : undefined,
echoCancellation: false,
noiseSuppression: false,
autoGainControl: false
}
};
this.mediaStream = await navigator.mediaDevices.getUserMedia(constraints);
this.sourceNode = this.audioCtx.createMediaStreamSource(this.mediaStream);
}
async start(destinationTrackGainNode, enableMonitoring = true) {
this.pcmChunks = [];
this.isRecording = true;
// Load Worklet Processor Module
await this.audioCtx.audioWorklet.addModule('/processors/pcm-recorder-processor.js');
this.workletNode = new AudioWorkletNode(this.audioCtx, 'pcm-recorder-processor');
// Receive PCM data streams from AudioWorklet
this.workletNode.port.onmessage = (event) => {
if (this.isRecording && event.data.type === 'PCM_DATA') {
this.pcmChunks.push(new Float32Array(event.data.buffer));
}
};
// Route Audio Nodes
this.sourceNode.connect(this.workletNode);
// Enable Live Input Monitoring if requested
if (enableMonitoring) {
this.sourceNode.connect(destinationTrackGainNode);
}
}
async stop() {
this.isRecording = false;
if (this.sourceNode && this.workletNode) {
this.sourceNode.disconnect(this.workletNode);
}
// Concatenate PCM Float32Array chunks into a single AudioBuffer
const totalSamples = this.pcmChunks.reduce((sum, chunk) => sum + chunk.length, 0);
if (totalSamples === 0) return null;
const audioBuffer = this.audioCtx.createBuffer(1, totalSamples, this.audioCtx.sampleRate);
const channelData = audioBuffer.getChannelData(0);
let offset = 0;
for (const chunk of this.pcmChunks) {
channelData.set(chunk, offset);
offset += chunk.length;
}
return audioBuffer; // Return compiled AudioBuffer for timeline insertion
}
}
```
---
## 6. UI Components & User Interactions
* **Track Header Arming Controls:**
* **[R] Button (Arm Track):** Highlights red when armed for recording on the target track.
* **[I] Button (Input Monitor):** Toggles live monitoring for incoming Microphone or Synth audio during performance.
* **Input Selector Dropdown:** Allows selection of available Microphone devices or USB Hardware MIDI Keyboards.
* **Real-time VU Meter Component:**
* Displays input signal gain level from $-60\text{ dB}$ to $0\text{ dB}$. Displays red clipping indicators when signal levels exceed $0\text{ dBFS}$.
* **Live Waveform & MIDI Preview Rendering:**
* **Microphone Recording:** The canvas UI renders incoming waveform signals progressing along the Playhead position in real time.
* **MIDI Performance:** Rectangular note blocks (green/orange) appear at note-on trigger events and extend until key release (note-off).
+237
View File
@@ -0,0 +1,237 @@
Here is the clean, nicely formatted Markdown version of the technical specification document:
# TECHNICAL INSTALLATION & INTEGRATION GUIDE FOR SOUNDFONT / VSTI IN DAW
This document provides a detailed technical architecture model for integrating SoundFonts, WebAssembly Plugins (Client), and Native VSTi/AU (Server). It clearly delineates components pre-installed by the Developer (Coder) versus those open for User uploads and additions.
---
## 1. Architectural Distribution Overview (Developer vs. User)
| Plugin / Asset Category | Processing Location | Installed By | Storage & Management Method | Security & Safety Profile |
| --- | --- | --- | --- | --- |
| **Default SoundFont (`.sf2`)** | Client (Wasm) | Coder | Static Assets hosted on Web Server / CDN | Extremely High |
| **User Custom SoundFont (`.sf2`)** | Client (Wasm) | User | Browser `IndexedDB` or User Cloud Storage | Extremely High (Runs inside Wasm Sandbox) |
| **WebAssembly Synths (WAMs)** | Client (JS/Wasm) | Coder | Bundled within Frontend Source Code | Extremely High |
| **Core Server VSTi (Vital, Surge...)** | Server (Python) | Coder | System Directory inside Docker/Linux Container | High (Controlled binary footprint) |
| **User Custom VST3 / Preset** | Server (Python) | User (Restricted) | Stores `.vst3` files or `.fxp`/`.json` on Container | High Security Risk (Requires Sandboxing) |
---
## 2. Client-Side Integration Tech (Browser / WebAssembly)
The Client-Side handles zero-latency real-time composition and audio previews.
### 2.1 Coder Pre-bundled Assets
* **Static SoundFont Hosting:**
* The developer places standard `.sf2` files (such as `GeneralUser_GS.sf2`) into the `public/soundfonts/` directory or hosts them via CDN.
* Upon application startup, default SoundFonts are queried via REST API:
```http
GET /api/v1/assets/default-soundfonts
```
```json
[
{ "id": "sf_generaluser", "name": "GeneralUser GS v1.471", "size_mb": 31.2, "url": "/soundfonts/GeneralUser.sf2" },
{ "id": "sf_sso", "name": "Sonatina Symphonic Orchestra", "size_mb": 95.0, "url": "/soundfonts/SSO.sf2" }
]
```
* **FluidSynth WebAssembly Engine Integration:**
* Compiles FluidSynth C/C++ code into WebAssembly (`fluidsynth.wasm` + `fluidsynth.js`) using Emscripten.
* Alternatively, leverages open JavaScript wrappers such as `@soundfont/player` or `SpessaSynth`.
### 2.2 Allowing User Custom SoundFont (`.sf2`) Uploads
Delivers a flexible user experience without overloading server storage:
* **Upload Mechanism & Local Cache (`IndexedDB`):**
* Users drag and drop `.sf2` files directly into the DAW interface.
* JavaScript reads the file as an `ArrayBuffer` via the `FileReader` API.
* The file persists directly within the browser's local `IndexedDB` cache for immediate reuse across sessions without re-uploading to the server.
* **Dynamic Injection into WebAssembly Memory:**
```javascript
// Client-side JavaScript snippet
async function loadUserSoundFont(fileBuffer) {
const uint8Array = new Uint8Array(fileBuffer);
// Write buffer straight into Emscripten FluidSynth Virtual File System (MEMFS)
Module.FS.writeFile('/user_font.sf2', uint8Array);
// Call Wasm C-function to load bank
const sfont_id = Module._fluid_synth_sfload(synthInstance, '/user_font.sf2', 1);
console.log(`User SoundFont loaded successfully with ID: ${sfont_id}`);
}
```
---
## 3. Server-Side Integration Tech (Python Backend Engine)
The Server-Side executes high-resolution offline WAV rendering when an operator triggers the Export / Bounce workflow.
### 3.1 Server Environment Installed by Coder
The developer configures the Server environment (or Docker Container) with pre-installed Native C++ libraries and Python utilities.
1. **Server Base `Dockerfile` Configuration:**
```dockerfile
FROM python:3.10-slim
# Install Linux audio libraries
RUN apt-get update && apt-get install -y \
fluidsynth \
libfluidsynth-dev \
libasound2-dev \
libjack-jackd2-dev \
build-essential \
&& rm -rf /var/lib/apt/lists/*
# Initialize directories for Native VST3 and system SoundFonts
RUN mkdir -p /opt/daw_engine/vst3 \
&& mkdir -p /opt/daw_engine/soundfonts
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
```
2. **Pre-installing Native VST3 Plugins:**
Places 64-bit Linux `.vst3` binary builds of open-source synths inside `/opt/daw_engine/vst3/`:
* `/opt/daw_engine/vst3/Vital.vst3`
* `/opt/daw_engine/vst3/Surge XT.vst3`
* `/opt/daw_engine/vst3/Dexed.vst3`
3. **Python Backend Integration via Spotify `pedalboard`:**
```python
# render_engine/vst_loader.py
import os
from pedalboard import VST3Plugin, Pedalboard
class PluginManager:
def __init__(self, vst_dir="/opt/daw_engine/vst3"):
self.vst_dir = vst_dir
self.available_plugins = self._scan_plugins()
def _scan_plugins(self):
plugins = {}
for root, dirs, files in os.walk(self.vst_dir):
for file in files:
if file.endswith(".vst3") or file.endswith(".so"):
plugin_path = os.path.join(root, file)
plugin_name = os.path.splitext(file)[0]
plugins[plugin_name] = plugin_path
return plugins
def load_vst(self, plugin_name: str, preset_data: dict = None) -> VST3Plugin:
if plugin_name not in self.available_plugins:
raise FileNotFoundError(f"VST3 Plugin '{plugin_name}' not found on server.")
path = self.available_plugins[plugin_name]
vst_instance = VST3Plugin(path)
# Inject parameters if provided
if preset_data:
for param_name, param_value in preset_data.items():
setattr(vst_instance, param_name, param_value)
return vst_instance
```
### 3.2 Handling User Custom Plugins / Presets
#### Option 1: User Presets / Patches Uploads (**RECOMMENDED - Safe**)
* **Implementation:** The backend locks native VST3 installations to common open engines (Vital, Dexed, Surge XT). Users upload lightweight preset patches like `.vitalbank`, `.syx` (DX7 patches), `.fxp`, or JSON parameter states.
* **Workflow:**
1. User selects the Vital Synth on the Client UI.
2. User clicks "Import Preset" $\rightarrow$ Uploads a `.vital` file or JSON parameter bundle.
3. Server parses JSON parameters and injects them directly into the VST3 instance via `pedalboard` during render execution.
* **Benefits:** Absolutely safe, minimal footprint, zero security vulnerabilities to the host infrastructure.
#### Option 2: User Native Binary VST3 Uploads (**HIGH RISK - Requires Isolation**)
* **Risk:** A `.vst3` file contains executable machine code (`.so` Shared Object on Linux). Accepting arbitrary uploads grants 100% vector exposure to Remote Code Execution (RCE) attacks.
* **Technical Mitigation (If Mandatory):**
* **Sandboxing Isolation:** Every user Export/Render request runs inside an isolated, short-lived container (Ephemeral Docker / Firejail / gVisor) stripped of `root` privileges and completely isolated from external internet interfaces.
* **Time-To-Live (TTL):** User `.vst3` binaries persist inside temporary directories `/tmp/user_sessions/{user_id}/` and purge automatically upon render job completion.
---
## 4. API Specification for SoundFonts & Plugins
### 4.1 OpenAPI Endpoint Spec for Frontend
```yaml
/api/v1/plugins/available:
get:
summary: Query available VSTi engines and SoundFont resources on the Server
responses:
200:
content:
application/json:
example:
vst_instruments:
- id: "vst_vital"
name: "Vital Wavetable Synth"
type: "VST3"
has_native_support: true
- id: "vst_dexed"
name: "Dexed FM Synth"
type: "VST3"
has_native_support: true
soundfonts:
- id: "sf_generaluser"
name: "GeneralUser GS"
file: "GeneralUser.sf2"
/api/v1/projects/render:
post:
summary: Trigger offline DAW Project rendering to WAV on the Server
requestBody:
required: true
content:
application/json:
schema:
$ref: '#/components/schemas/ProjectSchema'
responses:
200:
description: Returns the URL pointing to the rendered WAV file
```
---
## 5. Development Team Best Practices Summary
* **SoundFont (`.sf2`):**
* **For Users:** Encourage unrestricted local uploads on the Client (Browser). Store assets in `IndexedDB` to ensure optimal real-time performance without straining server resources.
* **For Developers:** Supply 12 default General MIDI (GM) SoundFont banks (`GeneralUser_GS.sf2`) bundled on both Client and Server.
* **VSTi Instruments:**
* **For Developers:** Pre-install top open-source Linux-native synths on the Server (Vital, Surge XT, Dexed, OB-Xd).
* **For Users:** Do **not** allow direct `.vst3` binary uploads to the production server. Instead, permit users to upload Presets / Patches / JSON parameters for the supported synth models. This guarantees 100% security while saving storage and network bandwidth.