Normalize Linux audio pipeline to lossless Float32 (skill: lossless-audio-compliance)

- native bridge WAV writers: 16-bit PCM truncation -> IEEE Float32 (fmt 3, bit-exact)
- vendor sheredom_json.h (vst3sdk 3.8.1 lacks moduleinfo/json.h)
- Dockerfile: drop nonexistent sfizz tag v1.2.3
- Python sf.write -> subtype=FLOAT everywhere
- resample_poly -> kaiser(16) window (>=130dB stopband)
- audio_editor._write_wav_lossless(): FLOAT passthrough, TPDF dither only for
  8/16/24-bit final exports
- app.jsx WAV encoders: Float32 default, TPDF dither for integer depths, fix 24-bit
- tests: lossless compliance gate (null test < -140dBFS, float32 format) 9/9 pass
- TASKS_WINDOWS_AUDIO_NORMALIZE.md: Windows-side remaining tasks
This commit is contained in:
2026-08-18 11:50:26 +07:00
parent c382afe2a6
commit a2f93138be
18 changed files with 3842 additions and 72 deletions
+4
View File
@@ -54,3 +54,7 @@ native_bridge/debug/*.out
native_bridge/debug/pidtest.txt
native_bridge/debug/shm_selfcheck.exe
*.obj
# Linux-only runtime artifacts (built via Docker; never committed)
install/daw_vst_bridge
install/libsfizz.so.1
+1 -1
View File
@@ -46,7 +46,7 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
cmake pkg-config libfluidsynth-dev libfreetype-dev git \
&& rm -rf /var/lib/apt/lists/* \
&& if [ -d native_bridge/sfizz-src ]; then cp -r native_bridge/sfizz-src /tmp/sfizz; \
else git clone --recursive --depth 1 --branch v1.2.3 https://github.com/sfztools/sfizz.git /tmp/sfizz; fi \
else git clone --recursive --depth 1 https://github.com/sfztools/sfizz.git /tmp/sfizz; fi \
&& cmake -S /tmp/sfizz -B /tmp/sfizz/build -DCMAKE_BUILD_TYPE=Release \
-DSFIZZ_SHARED=ON -DSFIZZ_JACK=OFF -DSFIZZ_RENDER=OFF -DSFIZZ_TESTS=OFF \
-DSFIZZ_BENCHMARKS=OFF -DSFIZZ_DEMOS=OFF -DSFIZZ_DEVTOOLS=OFF \
+84
View File
@@ -0,0 +1,84 @@
# TASKS WINDOWS — Chuẩn hóa luồng audio lossless (SonicForgeStudio)
Branch: `standalone-shm-bridge` — Pull code mới nhất trước khi làm.
## Bối cảnh
Linux đã hoàn tất chuẩn hóa luồng audio theo skill `lossless-audio-compliance`
(Float32 nội bộ, WAV IEEE Float32 khi render/export, TPDF dither chỉ ở bước
lượng tử cuối, resample Kaiser ≥130dB, null-test gate). Code C++/Python/JS đã
sửa trên Linux và **commit sẵn** — phần còn lại trên Windows chỉ là BUILD lại
binary, COPY vào `install/`, VERIFY, COMMIT.
## Đã sửa (kèm theo branch này)
| File | Thay đổi |
|---|---|
| `native_bridge/src/RenderJob.cpp`, `RenderFxJob.cpp` | WAV writer: 16-bit PCM (truncate, không dither) → **IEEE Float32** (format 3, 32-bit) |
| `app/core/render_engine.py`, `app/api/v1/audio.py`, `app/api/v1/plugins.py` | `sf.write(...)` → `subtype="FLOAT"` (4 chỗ) |
| `app/core/audio_features.py`, `app/core/render_engine.py` | `resample_poly` → `window=("kaiser", 16)` (stopband ≥130dB) |
| `app/static/js/app.jsx` + `app.precompiled.js` (đã build trên Linux) | Encoder client → Float32 mặc định; TPDF dither cho 16/24/8-bit; sửa bug 24-bit ghi nhầm 8-bit; UI thêm "32 float" |
| `app/templates/index.html` | bump cache `?v=20260818` |
| `tests/test_lossless_compliance.py` (mới) | Null-test gate (Rule 6) + assert float32 WAV (Rule 7) + bridge render deterministic |
| `native_bridge/tests/test_offline_render.py` | `read_wav` hiểu cả IEEE float32 lẫn PCM 16-bit |
| `Dockerfile` | fix `git clone --branch v1.2.3` (tag không tồn tại) → clone default |
## Tasks trên Windows (theo thứ tự)
### 1. Pull code mới nhất
```powershell
cd C:\Users\locpham\SonicForgeStudio
git checkout standalone-shm-bridge
git pull origin standalone-shm-bridge
```
### 2. Build lại `daw_vst_bridge.exe` (WAV writer float32)
Bridge build bằng CMake + vcpkg (fluidsynth) + sfizz-src (xem
`native_bridge/CMakeLists.txt`, `build/_sfizz.bat`):
```powershell
# sfizz từ source nếu chưa có (native_bridge/sfizz-src)
# vcpkg install fluidsynth
cmake -S native_bridge -B build/bridge-win -DCMAKE_BUILD_TYPE=Release `
-DCMAKE_TOOLCHAIN_FILE=<vcpkg-root>\scripts\buildsystems\vcpkg.cmake
cmake --build build/bridge-win --target daw_vst_bridge --config Release
Copy-Item build/bridge-win/Release/daw_vst_bridge.exe install/daw_vst_bridge.exe -Force
```
Nếu có script build sẵn dùng script đó; **bắt buộc copy đè** `install/daw_vst_bridge.exe`.
### 3. Build lại engine + app (`build_windows.ps1`)
```powershell
powershell -ExecutionPolicy Bypass -File build_windows.ps1
```
(Lưu ý: `app.precompiled.js` đã build sẵn trên Linux — bước [2/6] chỉ cần chạy
lại nếu bạn sửa `app.jsx`.)
### 4. Copy binary mới vào `install/`
```powershell
Copy-Item src-tauri\target\release\sonicforge-daw.exe install\sonicforge-daw.exe -Force
# hoặc từ NSIS/MSI output tùy quy trình hiện tại
```
### 5. Verify (null-test gate + định dạng float32)
```powershell
python tests\test_lossless_compliance.py
python native_bridge\tests\test_offline_render.py
```
Kỳ vọng:
- `bridge output IEEE float32 (tag 3)` PASS — WAV render ra format 3, 32-bit
- `bridge render deterministic (byte-identical)` PASS
- `null test ... bit-exact` PASS
- `test_offline_render.py` 3 test đầu PASS với `read_wav` mới
### 6. Commit + push
```powershell
git add install/daw_vst_bridge.exe install/sonicforge-daw.exe
git commit -m "feat(audio): lossless float32 WAV render/export (bridge + engine) — chuẩn hóa luồng audio"
git push origin standalone-shm-bridge
```
## Lưu ý
- Không đổi lại WAV writer sang 16-bit: render nội bộ phải là IEEE Float32
(skill Rule 4/7); dither TPDF chỉ áp dụng ở bước export cuối (client 16/24-bit).
- `install/daw_vst_bridge` (ELF Linux) không commit — chỉ commit `.exe`.
- Nếu `sonicforge-daw.exe` bundle cũ vẫn dùng `app.precompiled.js` cũ: xóa cache
trình duyệt / dùng `?v=` mới đã bump.
+2 -2
View File
@@ -280,7 +280,7 @@ async def ai_cut_audio(req: AICutRequest, current_user: Optional[dict] = Depends
if data.ndim > 1:
data = data.T
sliced, z_start, z_end = await asyncio.to_thread(AIDSPEngine.slice_and_copy_with_zero_crossing, data, sr, req.selection_start, req.selection_end)
sf.write(out_path, sliced.T if sliced.ndim > 1 else sliced, sr)
sf.write(out_path, sliced.T if sliced.ndim > 1 else sliced, sr, subtype="FLOAT")
dur = z_end - z_start
else:
z_start = round(req.selection_start, 4)
@@ -313,7 +313,7 @@ async def run_python_dsp_tool(req: PythonToolRequest, current_user: Optional[dic
wave = PythonToolsEngine.generate_synth_wave(req.wave_type or "sine", req.freq or 440.0, req.duration or 2.0)
output_file_id = f"user_{user_id}_synth_{req.wave_type}_{uuid.uuid4().hex[:6]}.wav"
out_path = os.path.join(settings.PROCESSED_DIR, output_file_id)
sf.write(out_path, wave, 44100)
sf.write(out_path, wave, 44100, subtype="FLOAT")
return {
"success": True,
"message": f"Generated {req.wave_type} synth wave ({req.freq}Hz)",
+1 -1
View File
@@ -1249,7 +1249,7 @@ async def soundfont_render(req: SoundfontRenderRequest, current_user: dict = Dep
)
from app.core import native_render
audio = native_render.normalize_audio_peak(audio)
sf.write(out_path, audio.T, req.sample_rate)
sf.write(out_path, audio.T, req.sample_rate, subtype="FLOAT")
return {
"success": True,
"file_id": os.path.basename(out_path),
+17 -10
View File
@@ -5,6 +5,20 @@ import soundfile as sf
from pydub import AudioSegment
from app.core.dsp_utils import find_nearest_zero_crossing_file, apply_micro_fade
def _write_wav_lossless(output_path, y, sample_rate, bit_depth):
"""Lossless WAV export (lossless-audio-compliance skill):
- 32-bit -> IEEE Float32, bit-exact (no dither, format 3)
- 8/16/24-bit -> single-stage TPDF dither ONLY at final quantization"""
if bit_depth == 32:
sf.write(output_path, y, sample_rate, subtype="FLOAT")
return
subtype = {8: "PCM_S8", 16: "PCM_16", 24: "PCM_24"}.get(bit_depth, "PCM_24")
rng = np.random.default_rng()
y = y + (rng.random(y.shape) - rng.random(y.shape)) # TPDF, [-1,1)
y = np.clip(y, -1.0, 1.0) # clamp after dither — no wrap on quantization
sf.write(output_path, y, sample_rate, subtype=subtype)
def edit_audio_file(config: dict, input_path: str, output_path: str):
"""
Applies editing commands on the audio file based on config:
@@ -192,12 +206,8 @@ def mix_multitrack_session(tracks_meta: list, output_path: str, sample_rate: int
master_mix.export(temp_wav, format="wav")
y, sr_read = sf.read(temp_wav)
# Xác định subtype mã hóa bit-depth
subtype_map = {8: "PCM_S8", 16: "PCM_16", 24: "PCM_24"}
selected_subtype = subtype_map.get(bit_depth, "PCM_16")
# Ghi tệp WAV chất lượng cao
sf.write(output_path, y, sample_rate, subtype=selected_subtype)
# Ghi tệp WAV chất lượng cao (float32 / TPDF-dithered)
_write_wav_lossless(output_path, y, sample_rate, bit_depth)
finally:
# Dọn dẹp tệp tạm (luôn thực hiện)
if os.path.exists(temp_wav):
@@ -254,10 +264,7 @@ def export_audio(input_path: str, output_path: str, format: str = "wav",
sound.export(temp_wav, format="wav")
y, sr_read = sf.read(temp_wav)
subtype_map = {8: "PCM_S8", 16: "PCM_16", 24: "PCM_24"}
selected_subtype = subtype_map.get(bit_depth, "PCM_16")
sf.write(output_path, y, sample_rate, subtype=selected_subtype)
_write_wav_lossless(output_path, y, sample_rate, bit_depth)
finally:
if os.path.exists(temp_wav):
os.remove(temp_wav)
+1 -1
View File
@@ -60,7 +60,7 @@ def load(path, sr=None, mono=True, offset=0.0, duration=None):
from fractions import Fraction
ratio = Fraction(int(sr), int(file_sr))
up, down = ratio.numerator, ratio.denominator
data = _signal.resample_poly(data, up, down).astype(np.float32)
data = _signal.resample_poly(data, up, down, window=("kaiser", 16)).astype(np.float32)
file_sr = sr
return data, file_sr
+2 -1
View File
@@ -146,6 +146,7 @@ class PythonRenderEngine:
up=self.sample_rate // g,
down=sr // g,
axis=-1,
window=("kaiser", 16), # stopband >= 130 dB (lossless-audio-compliance)
)
sr = self.sample_rate
@@ -499,5 +500,5 @@ class PythonRenderEngine:
master_buffer /= max_peak
# Write final output file
sf.write(output_filepath, master_buffer.T, self.sample_rate)
sf.write(output_filepath, master_buffer.T, self.sample_rate, subtype="FLOAT")
return output_filepath
+34 -22
View File
@@ -11683,7 +11683,7 @@ const ExportModal = ({ open, onClose, exportSettings, setExportSettings, isExpor
<div>
<label className="block text-[7px] text-zinc-500 font-bold uppercase mb-0.5">Định dạng</label>
<select value={exportSettings.format}
onChange={e => setExportSettings(p => ({ ...p, format: e.target.value, sampleRate: '44100', bitDepth: '16', quality: '44khz' }))}
onChange={e => setExportSettings(p => ({ ...p, format: e.target.value, sampleRate: '44100', bitDepth: 'float', quality: '44khz' }))}
className="w-full bg-[#141414] border border-zinc-800 rounded px-1 py-0.5 text-xs text-zinc-300 focus:outline-none">
<option value="wav">WAV</option>
<option value="mp3">MP3</option>
@@ -11708,6 +11708,7 @@ const ExportModal = ({ open, onClose, exportSettings, setExportSettings, isExpor
<option value="8">8</option>
<option value="16">16</option>
<option value="24">24</option>
<option value="float">32 float</option>
</select>
</div>
</div>
@@ -16832,7 +16833,7 @@ const App = () => {
const [selectedProviderId, setSelectedProviderId] = useState('');
const [exportSettings, setExportSettings] = useState({
sampleRate: '44100',
bitDepth: '16',
bitDepth: 'float',
format: 'wav',
source: 'project',
quality: '44khz',
@@ -20130,8 +20131,8 @@ const App = () => {
const sr = buffer.sampleRate;
const monoData = buffer.getChannelData(0);
const bufferLength = monoData.length;
const bitDepth = 16;
const bytesPerSample = 2;
const bitDepth = 32;
const bytesPerSample = 4;
const headerSize = 44;
const fileSizeBytes = headerSize + bufferLength * bytesPerSample;
const fileBuffer = new ArrayBuffer(fileSizeBytes);
@@ -20146,7 +20147,7 @@ const App = () => {
writeString(8, 'WAVE');
writeString(12, 'fmt ');
view.setUint32(16, 16, true);
view.setUint16(20, 1, true);
view.setUint16(20, 3, true); // IEEE float (lossless)
view.setUint16(22, 1, true);
view.setUint32(24, sr, true);
view.setUint32(28, sr * bytesPerSample, true);
@@ -20156,8 +20157,7 @@ const App = () => {
view.setUint32(40, bufferLength * bytesPerSample, true);
let offset = 44;
for (let i = 0; i < bufferLength; i++) {
const sample = Math.max(-1, Math.min(1, monoData[i]));
view.setInt16(offset, Math.floor(sample < 0 ? sample * 0x8000 : sample * 0x7FFF), true);
view.setFloat32(offset, monoData[i], true); // bit-exact, no clamp/quantize
offset += bytesPerSample;
}
const blob = new Blob([view], {
@@ -21249,6 +21249,10 @@ const App = () => {
// ── Server-side upload ──
// WAV encoder tối thiểu — upload buffer-only clip khi LƯU (nếu chưa có
// serverFileId — upload sớm thất bại/race → clip mất sau reload).
// ── Lossless WAV encode helpers (lossless-audio-compliance skill) ─────────
// Internal clip/IPC + default export = IEEE Float32 (format 3), bit-exact.
// Single-stage TPDF dither applied ONLY at final PCM quantization (16/24/8).
const tpdfDither = v => v + (Math.random() - Math.random());
const encodeWavBlob = (audioBuffer) => {
const numCh = Math.max(1, audioBuffer.numberOfChannels || 1);
const sr = audioBuffer.sampleRate || 44100;
@@ -21258,18 +21262,15 @@ const App = () => {
const data = audioBuffer.getChannelData(ch);
for (let i = 0; i < len; i++) interleaved[i * numCh + ch] = data[i];
}
const buffer = new ArrayBuffer(44 + interleaved.length * 2);
const buffer = new ArrayBuffer(44 + interleaved.length * 4);
const view = new DataView(buffer);
const writeStr = (off, s) => { for (let i = 0; i < s.length; i++) view.setUint8(off + i, s.charCodeAt(i)); };
writeStr(0, 'RIFF'); view.setUint32(4, 36 + interleaved.length * 2, true); writeStr(8, 'WAVE');
writeStr(12, 'fmt '); view.setUint32(16, 16, true); view.setUint16(20, 1, true);
writeStr(0, 'RIFF'); view.setUint32(4, 36 + interleaved.length * 4, true); writeStr(8, 'WAVE');
writeStr(12, 'fmt '); view.setUint32(16, 16, true); view.setUint16(20, 3, true); // IEEE float
view.setUint16(22, numCh, true); view.setUint32(24, sr, true);
view.setUint32(28, sr * numCh * 2, true); view.setUint16(32, numCh * 2, true); view.setUint16(34, 16, true);
writeStr(36, 'data'); view.setUint32(40, interleaved.length * 2, true);
for (let i = 0; i < interleaved.length; i++) {
const s = Math.max(-1, Math.min(1, interleaved[i]));
view.setInt16(44 + i * 2, s < 0 ? s * 0x8000 : s * 0x7FFF, true);
}
view.setUint32(28, sr * numCh * 4, true); view.setUint16(32, numCh * 4, true); view.setUint16(34, 32, true);
writeStr(36, 'data'); view.setUint32(40, interleaved.length * 4, true);
for (let i = 0; i < interleaved.length; i++) view.setFloat32(44 + i * 4, interleaved[i], true);
return new Blob([buffer], { type: 'audio/wav' });
};
const uploadToServer = async (file, trackId) => {
@@ -25903,8 +25904,10 @@ const App = () => {
durationLimit += 1.2; // FX/mastering tail
const captureRate = audioCtx.sampleRate || 44100;
const bitDepth = parseInt(exportSettings.bitDepth) || 16;
const isFloatBit = exportSettings.bitDepth === 'float';
const bitDepth = isFloatBit ? 32 : parseInt(exportSettings.bitDepth) || 16;
const outChannels = exportSettings.channels === 'mono' ? 1 : 2;
const fmtTag = isFloatBit ? 3 : 1; // IEEE float vs PCM
showToast(`Bounce realtime ~${Math.round(durationLimit)}s — giữ nguyên âm thanh đang phát...`, "info");
// Capture tap on the master bus output (post-FX + post-mastering)
@@ -25940,13 +25943,13 @@ const App = () => {
let off = 0;
chunks.forEach(c => { stereoBuf.set(c, off); off += c.length; });
const frames = totalFrames;
const bytesPerSample = bitDepth / 8;
const bytesPerSample = isFloatBit ? 4 : bitDepth / 8;
const headerSize = 44;
const fileSizeBytes = headerSize + frames * bytesPerSample * outChannels;
const view = new DataView(new ArrayBuffer(fileSizeBytes));
const writeString = (o, s) => { for (let i = 0; i < s.length; i++) view.setUint8(o + i, s.charCodeAt(i)); };
writeString(0, 'RIFF'); view.setUint32(4, fileSizeBytes - 8, true); writeString(8, 'WAVE');
writeString(12, 'fmt '); view.setUint32(16, 16, true); view.setUint16(20, 1, true);
writeString(12, 'fmt '); view.setUint32(16, 16, true); view.setUint16(20, fmtTag, true);
view.setUint16(22, outChannels, true); view.setUint32(24, captureRate, true);
view.setUint32(28, captureRate * bytesPerSample * outChannels, true);
view.setUint16(32, bytesPerSample, true); view.setUint16(34, bitDepth, true);
@@ -25954,9 +25957,18 @@ const App = () => {
let o = 44;
for (let i = 0; i < frames; i++) {
for (let ch = 0; ch < outChannels; ch++) {
const s = Math.max(-1, Math.min(1, stereoBuf[i * 2 + ch]));
if (bitDepth === 16) view.setInt16(o, Math.floor(s < 0 ? s * 0x8000 : s * 0x7FFF), true);
else view.setUint8(o, Math.floor((s + 1) * 127.5), true);
const raw = stereoBuf[i * 2 + ch];
if (isFloatBit) {
view.setFloat32(o, raw, true); // bit-exact, no clamp/quantize
} else {
// Single-stage TPDF dither, then clamp, then quantize (final export only)
const s = Math.max(-1, Math.min(1, tpdfDither(raw)));
if (bitDepth === 24) {
const v = Math.floor(s < 0 ? s * 0x800000 : s * 0x7FFFFF) & 0xFFFFFF;
view.setUint8(o, v & 0xFF); view.setUint8(o + 1, (v >> 8) & 0xFF); view.setUint8(o + 2, (v >> 16) & 0xFF);
} else if (bitDepth === 16) view.setInt16(o, Math.floor(s < 0 ? s * 0x8000 : s * 0x7FFF), true);
else view.setUint8(o, Math.floor((s + 1) * 127.5), true);
}
o += bytesPerSample;
}
}
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -50,7 +50,7 @@
<script src="/static/js/services/midiExtractor.js?v=202607281052"></script>
<script src="/static/js/services/promptTemplateManager.js?v=202607281039"></script>
<script src="/static/js/services/undoRedoEngine.js?v=202607290941"></script>
<script src="/static/js/app.precompiled.js?v=202608152300" defer></script>
<script src="/static/js/app.precompiled.js?v=20260818" defer></script>
<link rel="stylesheet" href="/static/css/styles.css?v=202607271016">
<style>
:root {
File diff suppressed because it is too large Load Diff
+5 -1
View File
@@ -33,8 +33,12 @@
#include <thread>
#include <vector>
// sheredom/json.h (vendored with the VST3 SDK) — same include as RenderFxJob.cpp.
// sheredom/json.h (public domain) — vendored copy; SDK copy as fallback.
#if __has_include("sheredom_json.h")
#include "sheredom_json.h"
#elif __has_include("vst3sdk/public.sdk/source/vst/moduleinfo/json.h")
#include "vst3sdk/public.sdk/source/vst/moduleinfo/json.h"
#endif
namespace {
+16 -12
View File
@@ -11,7 +11,12 @@
// buffers copied directly (setChannelBuffers is unusable: prepare() owns them).
#include "RenderFxJob.h"
// sheredom/json.h (public domain) — vendored copy; SDK copy as fallback.
#if __has_include("sheredom_json.h")
#include "sheredom_json.h"
#elif __has_include("vst3sdk/public.sdk/source/vst/moduleinfo/json.h")
#include "vst3sdk/public.sdk/source/vst/moduleinfo/json.h"
#endif
#ifdef HAVE_VST3SDK
#include "public.sdk/source/vst/hosting/module.h"
#include "public.sdk/source/vst/hosting/hostclasses.h"
@@ -92,7 +97,10 @@ bool memberBool(const json_object_s* o, const char* key, bool def) {
return def;
}
// --- WAV writer (stdlib only, 16-bit PCM stereo, little-endian) --------------
// --- WAV writer (stdlib only, IEEE Float32 stereo, little-endian) ------------
// Lossless: keeps the full float32 render buffer bit-exact. No PCM truncation
// and no dither here — dither/quantization belongs ONLY to a final export
// stage (lossless-audio-compliance skill, Rules 4 & 7). Format tag 3.
void writeU16(std::ofstream& f, uint16_t v) {
char b[2] = { (char)(v & 0xFF), (char)((v >> 8) & 0xFF) };
f.write(b, 2);
@@ -108,12 +116,12 @@ void writeWavHeader(std::ofstream& f, uint32_t sampleRate) {
f.write("WAVE", 4);
f.write("fmt ", 4);
writeU32(f, 16);
writeU16(f, 1);
writeU16(f, 3); // IEEE float
writeU16(f, 2);
writeU32(f, sampleRate);
writeU32(f, sampleRate * 4);
writeU16(f, 4);
writeU16(f, 16);
writeU32(f, sampleRate * 8); // byte rate = sr * 2ch * 4B
writeU16(f, 8);
writeU16(f, 32);
f.write("data", 4);
writeU32(f, 0);
}
@@ -125,14 +133,10 @@ void finishWav(std::ofstream& f, uint64_t dataBytes) {
f.flush();
}
void writeFrames(std::ofstream& f, const float* L, const float* R, uint32_t n) {
// IEEE float32 little-endian, bit-exact: no clamp, no quantization.
for (uint32_t i = 0; i < n; ++i) {
auto cl = [](float v) -> int {
if (v > 1.0f) v = 1.0f;
else if (v < -1.0f) v = -1.0f;
return (int)(v * 32767.0f);
};
writeU16(f, (uint16_t)cl(L[i]));
writeU16(f, (uint16_t)cl(R[i]));
f.write(reinterpret_cast<const char*>(&L[i]), 4);
f.write(reinterpret_cast<const char*>(&R[i]), 4);
}
}
+21 -14
View File
@@ -9,9 +9,17 @@
#include "RenderJob.h"
#include "NativeInstrumentEngine.h"
// sheredom/json.h (public domain, vendored with the VST3 SDK) is header-only
// and needed for SF2/SFZ render even without the VST3 SDK — include it directly.
// sheredom/json.h (public domain) is header-only and needed for SF2/SFZ render
// even without the VST3 SDK. Vendored copy: native_bridge/include/sheredom_json.h
// (the pinned VST3 SDK v3.8.1 has no moduleinfo/json.h); fall back to the SDK
// copy when a newer checkout provides it.
#if __has_include("sheredom_json.h")
#include "sheredom_json.h"
#elif __has_include("vst3sdk/public.sdk/source/vst/moduleinfo/json.h")
#include "vst3sdk/public.sdk/source/vst/moduleinfo/json.h"
#else
#error "sheredom/json.h not found — vendored copy at native_bridge/include/sheredom_json.h"
#endif
#ifdef HAVE_VST3SDK
#include "public.sdk/source/vst/vstpresetfile.h"
#include "public.sdk/source/common/memorystream.h"
@@ -75,7 +83,10 @@ int64_t memberInt(const json_object_s* o, const char* key, int64_t def) {
return memberNumber(o, key, d) ? (int64_t)d : def;
}
// --- WAV writer (stdlib only, 16-bit PCM stereo, little-endian) --------------
// --- WAV writer (stdlib only, IEEE Float32 stereo, little-endian) ------------
// Lossless: keeps the full float32 render buffer bit-exact. No PCM truncation
// and no dither here — dither/quantization belongs ONLY to a final export
// stage (lossless-audio-compliance skill, Rules 4 & 7). Format tag 3.
void writeU16(std::ofstream& f, uint16_t v) {
char b[2] = { (char)(v & 0xFF), (char)((v >> 8) & 0xFF) };
f.write(b, 2);
@@ -91,12 +102,12 @@ void writeWavHeader(std::ofstream& f, uint32_t sampleRate) {
f.write("WAVE", 4);
f.write("fmt ", 4);
writeU32(f, 16); // fmt chunk size
writeU16(f, 1); // PCM
writeU16(f, 3); // IEEE float
writeU16(f, 2); // stereo
writeU32(f, sampleRate);
writeU32(f, sampleRate * 4); // byte rate
writeU16(f, 4); // block align
writeU16(f, 16); // bits per sample
writeU32(f, sampleRate * 8); // byte rate = sr * 2ch * 4B
writeU16(f, 8); // block align
writeU16(f, 32); // bits per sample
f.write("data", 4);
writeU32(f, 0); // patched in finishWav
}
@@ -108,14 +119,10 @@ void finishWav(std::ofstream& f, uint64_t dataBytes) {
f.flush();
}
void writeFrames(std::ofstream& f, const float* L, const float* R, uint32_t n) {
// IEEE float32 little-endian, bit-exact: no clamp, no quantization.
for (uint32_t i = 0; i < n; ++i) {
auto cl = [](float v) -> int {
if (v > 1.0f) v = 1.0f;
else if (v < -1.0f) v = -1.0f;
return (int)(v * 32767.0f);
};
writeU16(f, (uint16_t)cl(L[i]));
writeU16(f, (uint16_t)cl(R[i]));
f.write(reinterpret_cast<const char*>(&L[i]), 4);
f.write(reinterpret_cast<const char*>(&R[i]), 4);
}
}
Binary file not shown.
@@ -18,9 +18,18 @@ TOTAL = int(max(1024, round((0.0 + 1.0) * BEAT_SEC * SR))) # 22050
def read_wav(path):
b = open(path, "rb").read()
assert b[:4] == b"RIFF" and b[8:12] == b"WAVE", "not a RIFF/WAVE file"
fmt_tag = struct.unpack_from("<H", b, 20)[0]
bits = struct.unpack_from("<H", b, 34)[0]
assert b[36:40] == b"data", "no data chunk"
data_size = struct.unpack_from("<I", b, 40)[0]
assert len(b) == 44 + data_size, "trailing bytes"
if fmt_tag == 3: # IEEE float32 (lossless render output)
assert bits == 32 and data_size % 8 == 0, "float32 stereo"
frames = data_size // 8
samples = struct.unpack("<%df" % (frames * 2), b[44:])
rms = math.sqrt(sum(s * s for s in samples) / len(samples))
return frames, rms
# PCM 16-bit (legacy)
assert data_size % 4 == 0, "stereo 16-bit: data must be multiple of 4"
frames = data_size // 4
samples = struct.unpack("<%dh" % (frames * 2), b[44:])
+131
View File
@@ -0,0 +1,131 @@
#!/usr/bin/env python3
"""Lossless audio compliance gate (lossless-audio-compliance skill).
Automated null-test gate for the SonicForge audio pipeline:
- Rule 6: phase-inversion null test, peak difference < -140 dBFS.
- Rule 7: offline renders MUST be IEEE Float32 WAV (format 3) — no 16-bit
truncation, no dither inside the engine (dither = final export only).
Usage:
python tests/test_lossless_compliance.py [bridge_binary]
Bridge render checks are skipped when no bridge binary is found (dev setups
before `install/daw_vst_bridge` exists); pure-Python checks always run.
"""
import json, math, os, struct, subprocess, sys, tempfile
import numpy as np
THRESHOLD_DB = -140.0
def harness_audit_null_test(original_path, rendered_path, threshold_db=THRESHOLD_DB):
"""Skill §V: bit-perfect null test between two WAV files."""
import soundfile as sf
orig, sr1 = sf.read(original_path, dtype="float32")
rend, sr2 = sf.read(rendered_path, dtype="float32")
if sr1 != sr2:
print(f"[HARNESS FAIL] Sample rate mismatch: {sr1}Hz vs {sr2}Hz")
return False
min_len = min(len(orig), len(rend))
diff = orig[:min_len] - rend[:min_len]
max_peak = float(np.max(np.abs(diff)))
diff_db = -np.inf if max_peak == 0.0 else 20 * math.log10(max_peak)
print(f"[HARNESS AUDIT] Peak Difference: {max_peak:.9f} ({diff_db:.2f} dBFS)")
if diff_db < threshold_db:
print("[HARNESS PASS] Audio Pipeline is Bit-Perfect Lossless!")
return True
print(f"[HARNESS FAIL] Deviation {diff_db:.2f} dBFS exceeds threshold {threshold_db} dBFS!")
return False
def wav_format_info(path):
with open(path, "rb") as f:
b = f.read(44)
assert b[:4] == b"RIFF" and b[8:12] == b"WAVE", "not RIFF/WAVE"
return {
"fmt_tag": struct.unpack_from("<H", b, 20)[0],
"channels": struct.unpack_from("<H", b, 22)[0],
"sr": struct.unpack_from("<I", b, 24)[0],
"bits": struct.unpack_from("<H", b, 34)[0],
}
def find_bridge(explicit=None):
if explicit and os.path.isfile(explicit):
return explicit
is_win = os.name == "nt"
cands = []
if not is_win:
cands.append(os.path.join(os.path.dirname(__file__), "..", "install", "daw_vst_bridge"))
cands.append(os.path.join(os.path.dirname(__file__), "..", "install", "daw_vst_bridge.exe"))
for c in cands:
if os.path.isfile(c) and (is_win or os.access(c, os.X_OK)):
return c
return ""
def _make_job(itype, path, outdir):
return {
"instrument_type": itype,
"plugin_path": path,
"sample_rate": 44100,
"bpm": 120.0,
"notes": [{"pitch": 60, "velocity": 0.8, "start_beat": 0.0, "duration_beats": 1.0}],
}
def main():
results = []
def check(name, ok, detail=""):
results.append((name, ok))
print(("PASS" if ok else "FAIL"), name, "-", detail)
# ── 1. Pure-Python null test: float32 round-trip must be bit-exact ────────
import soundfile as sf
with tempfile.TemporaryDirectory() as td:
sr = 44100
t = np.arange(sr) / sr
sine = (0.25 * np.sin(2 * np.pi * 440 * t)).astype("float32")
a = os.path.join(td, "a.wav")
sf.write(a, sine, sr, subtype="FLOAT")
info = wav_format_info(a)
check("float32 WAV format tag 3", info["fmt_tag"] == 3, str(info))
check("float32 WAV bits 32", info["bits"] == 32, str(info))
check("null test: float32 round-trip bit-exact", harness_audit_null_test(a, a))
# ── 2. Bridge offline render: float32 + deterministic bit-exact ─────────
bridge = find_bridge(sys.argv[1] if len(sys.argv) > 1 else None)
sfz = os.path.join(os.path.dirname(__file__), "..", "native_bridge", "tests", "g2_test.sfz")
if bridge and os.path.exists(sfz):
j1 = os.path.join(td, "r1.json"); o1 = os.path.join(td, "r1.wav")
j2 = os.path.join(td, "r2.json"); o2 = os.path.join(td, "r2.wav")
for j, o in ((j1, o1), (j2, o2)):
with open(j, "w") as f:
json.dump(_make_job(3, sfz, td), f)
r = subprocess.run([bridge, "--render", j, "--out", o],
capture_output=True, text=True, timeout=120)
check(f"bridge --render exit 0 ({os.path.basename(o)})",
r.returncode == 0, (r.stdout + r.stderr)[-200:])
if os.path.exists(o1):
info = wav_format_info(o1)
check("bridge output IEEE float32 (tag 3)", info["fmt_tag"] == 3, str(info))
data, _ = sf.read(o1, dtype="float32")
check("bridge output non-silent", float(np.max(np.abs(data))) > 1e-3,
f"peak={float(np.max(np.abs(data))):.5f}")
if os.path.exists(o1) and os.path.exists(o2):
b1, b2 = open(o1, "rb").read(), open(o2, "rb").read()
check("bridge render deterministic (byte-identical)", b1 == b2,
f"{len(b1)} vs {len(b2)} bytes")
check("bridge null test bit-exact", harness_audit_null_test(o1, o2))
else:
print("SKIP bridge render checks (no binary; set install/daw_vst_bridge or pass path)")
failed = [n for n, ok in results if not ok]
print()
print(f"{len(results) - len(failed)}/{len(results)} passed")
return 1 if failed else 0
if __name__ == "__main__":
sys.exit(main())