13 KiB
Technical Specification: Sandbox Isolation & Sub-Tab DSP Editing Algorithms
This document defines the processing workflow design and digital signal processing (DSP) algorithms dedicated to localized clip editing within a temporary isolated document workspace (Sub-tab).
1. Sandbox Splicing Workflow
When a user highlights a time region on the Main Tab and triggers "Edit in Sub-tab" or presses the edit keyboard shortcut:
[ MAIN TAB - MULTITRACK ]
Track 01: ───[█████ Selected Segment █████]───
│
▼ (Copy to Clipboard Buffer)
[ KHỔI TẠO TAB TẠM THỜI (SUB-TAB) ]
- Instantiates a single Track (Height bounds: 48px - 200px via ns-resize)
- Timeline Ruler axis resets to t = 0.0s
│
▼ (Automated Insertion - Auto-Paste)
Track 01 (Sub-tab): [█████ Isolated Segment █████] at t = 0s
-
Extract Buffer: The underlying engine extracts the binary sample array (
Float32Array) of the highlighted region from the active track, caching it securely into the application's clipboard buffer memory. -
Sandbox Environment Initialization:
-
Appends a temporary document window onto the global Tab Bar (e.g.,
Tab: sẤit tiá...norSub_Edit_1). -
Focuses the viewport down into the sandboxed tab. Here, a single standalone track lane is drawn, mapping the timeline ruler scale to start at
t = 0.0\text{ s}up to the absolute duration limit (T_{\text{clip}}) of the extracted audio asset. -
Track Height Resizing:
-
Hovering the cursor over the lower layout bounding path of the track lane changes the style configuration to
ns-resize. -
Dragging downward expands the vertical height ceiling (up to an upper bound of
200\text{ px}), maximizing the waveform amplitude drawing path for precision clip editing. Dragging upward compresses the physical row dimensions (down to a lower constraint of48\text{ px}) to protect screen space. -
Auto-Paste Routine: The framework automates the insertion sequence, dropping the cached array block onto the root index milestone (
t = 0.0\text{ s}) inside the isolated single-track layer.
2. Apply & Sync-Back Workflow
When an editor finishes processing steps inside the sandbox workspace and engages the Apply action:
[ SUB-TAB - AUDIO SANDBOX ]
y_sub = [█████ Edited Waveform █████]
│
▼ (Click "Apply" - Trigger Overwrite)
[ MAIN TAB - ORIGINAL TRACK ]
Track 01: ───[█████ Overwritten Segment █████]─── at t = t_start
│
(Sub-tab remains open, Undo Stack kept)
▼
[User presses Undo (Ctrl+Z) inside Sub-tab to iterate]
y_sub = [█████ Rollbacked Waveform █████]
│
(Click "Apply" again)
▼
Track 01: ───[█████ Corrected Segment █████]──── at t = t_start
2.1. Target Mapping & Metadata Linkage
Throughout its lifecycle, each sub-tab persistently locks standard metadata records linking back to the origin source elements:
parent_track_id: Unique identifier referencing the primary source track on the Main Tab.parent_clip_id: Unique identifier tracking the original source audio clip.t_start(seconds): The exact historical start time position of the sliced block on the Main Tab timeline view.original_duration(seconds): The baseline temporal duration of the region prior to modification.
2.2. In-place Overwrite & Splicing
- Edited Buffer Extraction: The system reads the active sample sequence from the sub-tab (
y_{\text{sub}}) along with its updated duration boundaryT_{\text{sub}}(which fluctuates if time-stretching or rate scaling actions have occurred). - Main Session Integration:
- The core route mapper checks for the matching
parent_track_idparameter on the Main Tab. - Purges the legacy audio segment stretching from
t_{\text{start}}throught_{\text{start}} + T_{\text{original}}. - Splices the updated signal array
y_{\text{sub}}precisely at the historical insertion indext_{\text{start}}. - Micro-crossfade: Executes a ultra-fast crossfade envelope (
10\text{ ms}) across both the initial and terminating splice boundaries. Blending adjacent files prevents phase cancellation or signal breakage that manifests as transient clicks/pops.
- Visual Update Tracking: Commands the canvas engine to redraw the waveform visualization grid for the origin track lane inside the Main Tab view.
2.3. Persistence for Iterative Editing
- Tab Lifetime: Engaging the Apply trigger propagates data back to the primary environment but does not close down the active sub-tab view.
- Undo Stack Isolation: The tracking loop containing the localized Undo/Redo History Stack inside the sub-tab sandbox remains entirely preserved.
- Iterative Loop Workflow:
- If monitoring the Main Tab arrangement uncovers an audio anomaly, the user switches focus back to the Sub-tab workspace.
- Pressing
Ctrl + Z(Undo) rollbacks the localized signal to its earlier state. - The editor runs separate DSP actions.
- Hitting Apply overwrites the updated audio slice over the same target coordinates on the Main Tab.
- Explicit Destruction Hook: The sandboxed tab structure is only unmapped when the user clicks the explicit close icon (
\times) on the horizontal tab bar.
3. Sub-Tab DSP Algorithm Specification
Editing operations executed inside the sub-tab environment calculate discrete changes over the amplitude sample arrays (x[n]). These map to Web Audio API routines on the client layer and standard NumPy/SciPy audio arrays on the Dockerized backend.
3.1. Time-Stretching & Speed Math
Alters the duration bounds of the audio clip with optional pitch-shifting linking logic:
- Pitch-preserving Time-stretching: Utilizes the Phase Vocoder method to analyze the Short-Time Fourier Transform (STFT) of the signal, shifts spectral frames across the frequency domain, and reconstructs the audio via the Inverse Short-Time Fourier Transform (ISTFT) to align with a new playback velocity ratio
S:
S = \frac{D}{D'} \times 100\%
Where: D corresponds to the legacy unscaled duration (seconds), and D' maps to the updated value post-resizing (executed by holding down the Alt key and dragging the right boundary handle).
- Resampling (Pitch-shifting Speed Scale): Runs a standard linear interpolation algorithm to resample the core data array size:
x_{\text{new}}[m] = x\left[ \frac{m \cdot D}{D'} \right]
3.2. Peak Normalization
Amplifies the signal scale uniformly across the active block until the single maximum absolute sample peak reaches a specified ceiling parameter A_{\text{target}} (typically locked at 1.0 or 0\text{ dBFS}):
- Evaluate the absolute maximum peak within the array bounds:
A_{\text{max}} = \max_{n=0}^{N-1} \vert x[n] \vert
- Compute the static gain multiplier constant
G:
G = \frac{A_{\text{target}}}{A_{\text{max}}}
- Multiply the entire audio array values by
G:
x_{\text{norm}}[n] = x[n] \cdot G
3.3. Volume Gain Adjustment (dB Scaling)
- Capture the decibel variance target (
\Delta \text{dB}). - Translate the logarithmic value into a standard linear scalar multiplier variable
G_{\text{linear}}:
G_{\text{linear}} = 10^{\frac{\Delta \text{dB}}{20}}
- Apply the gain multiplier directly into the sample values:
x_{\text{gained}}[n] = x[n] \cdot G_{\text{linear}}
3.4. Pitch Shifting
Shifts the fundamental frequencies of the signal up or down by a specific number of semitones (n) while keeping the temporal duration value completely intact.
- Frequency Transposition Ratio (
F_{\text{ratio}}):
F_{\text{ratio}} = 2^{\frac{n}{12}}
- DSP Processing Pipeline: Employs either a Pitch Synchronous Overlap and Add (PSOLA) routine or a spectral Phase Vocoder to expand/compress the frequency components, then passes the array into a time-stretching step to return the physical track length to its source metric
T_{\text{clip}}.
3.5. Linear Fade-In & Fade-Out Curves
Applies a linear fading envelope over the boundaries of the audio data block.
- Linear Fade-In Envelope (Across a duration bound of
L_{\text{fade}}samples):
x_{\text{fade}}[n] = x[n] \cdot \left( \frac{n}{L_{\text{fade}}} \right) \quad \text{for } 0 \le n < L_{\text{fade}}
- Linear Fade-Out Envelope (Across the final trailing
L_{\text{fade}}samples):
x_{\text{fade}}[N - 1 - n] = x[N - 1 - n] \cdot \left( \frac{n}{L_{\text{fade}}} \right) \quad \text{for } 0 \le n < L_{\text{fade}}
3.6. Array Splitting & Merging
- Split at Position (
n_{\text{cut}}): Unlinks a single sample blockx[n]of sizeNinto two separate independent sub-arrays:
x_1[n] = x[n] \quad (0 \le n < n_{\text{cut}})
x_2[n] = x[n + n_{\text{cut}}] \quad (0 \le n < N - n_{\text{cut}})
- Merge Segments: Concatenates separate sample sequences end-to-end. The stitching logic runs a
10\text{ ms}micro-crossfade overlay envelope at the junction to smooth out phase gaps that prompt click artifacts.
4. Python Backend Implementation Manual
This prototype Python class (core/sub_tab_dsp.py) handles the sandboxed operations and includes the crossfaded structural splicing algorithm designed to run inside the Docker engine:
import numpy as np
import scipy.signal as signal
import librosa
class SubTabDSPEngine:
@staticmethod
def change_speed(y: np.ndarray, sr: int, speed_ratio: float, preserve_pitch: bool = True) -> np.ndarray:
"""
Alters the playback velocity (Time-Stretching) of a NumPy signal array.
"""
if speed_ratio == 1.0:
return y
if preserve_pitch:
return librosa.effects.time_stretch(y, rate=speed_ratio)
else:
num_samples_new = int(len(y) / speed_ratio)
return signal.resample(y, num_samples_new)
@staticmethod
def normalize(y: np.ndarray, target_db: float = 0.0) -> np.ndarray:
"""
Performs Peak Normalization on an array to scale it to the target decibel value.
"""
target_amplitude = 10.0 ** (target_db / 20.0)
max_amplitude = np.max(np.abs(y))
if max_amplitude == 0:
return y
gain = target_amplitude / max_amplitude
return y * gain
@staticmethod
def merge_back_to_parent(
parent_track_audio: np.ndarray,
sr: int,
edited_sub_audio: np.ndarray,
start_seconds: float,
original_duration_seconds: float
) -> np.ndarray:
"""
Splices the modified audio segment from the Sub-tab back into the parent track array.
Applies a 10ms micro-crossfade at the boundaries to eliminate pop/click noise.
"""
start_sample = int(start_seconds * sr)
original_samples_len = int(original_duration_seconds * sr)
edited_samples_len = len(edited_sub_audio)
crossfade_samples = int(0.01 * sr) # 10ms crossfade window
# 1. Allocate the target output array dimension bounds
new_total_len = len(parent_track_audio) - original_samples_len + edited_samples_len
output_audio = np.zeros(new_total_len, dtype=np.float32)
# 2. Extract leading unedited block
output_audio[:start_sample] = parent_track_audio[:start_sample]
# 3. Stitch the modified audio payload
output_audio[start_sample:start_sample + edited_samples_len] = edited_sub_audio
# 4. Extract trailing unedited block
post_start_original = start_sample + original_samples_len
post_start_new = start_sample + edited_samples_len
output_audio[post_start_new:] = parent_track_audio[post_start_original:]
# 5. Execute micro-crossfade across the initial splice junction
if start_sample > crossfade_samples:
fade_in_ramp = np.linspace(0.0, 1.0, crossfade_samples)
fade_out_ramp = np.linspace(1.0, 0.0, crossfade_samples)
# Smooth 10ms interpolation overlay
output_audio[start_sample : start_sample + crossfade_samples] = (
edited_sub_audio[:crossfade_samples] * fade_in_ramp +
parent_track_audio[start_sample : start_sample + crossfade_samples] * fade_out_ramp
)
# 6. Execute micro-crossfade across the trailing splice junction
if post_start_new + crossfade_samples < len(output_audio):
fade_in_ramp = np.linspace(0.0, 1.0, crossfade_samples)
fade_out_ramp = np.linspace(1.0, 0.0, crossfade_samples)
output_audio[post_start_new : post_start_new + crossfade_samples] = (
parent_track_audio[post_start_original : post_start_original + crossfade_samples] * fade_in_ramp +
edited_sub_audio[-crossfade_samples:] * fade_out_ramp
)
return output_audio