diff --git a/README-fr.md b/README-fr.md index 3fd830d..b095faa 100644 --- a/README-fr.md +++ b/README-fr.md @@ -19,7 +19,7 @@ K.G.Studio est une DAW légère et moderne qui fonctionne entièrement dans le navigateur, avec **K.G.Studio Musician Assistant** en son cœur. Il propose une lecture réaliste des instruments via les samplers Tone.js, un éditeur piano roll, une gestion des pistes et des régions avec annulation/rétablissement complets, la persistance des projets dans OPFS (Origin Private File System), un panneau de configuration personnalisable, et un assistant IA intégré avec exécution d'outils. -**K.G.Studio Musician Assistant** est un agent d'assistance IA pour l'harmonie, l'arrangement et l'édition de notes, mais pas pour la composition entièrement automatique. +**K.G.Studio Musician Assistant** est un co-créateur IA conscient du projet, conçu pour enrichir votre flux créatif. Plutôt que de générer directement des fichiers audio bruts, il opère au niveau structuré des pistes et des notes — vous aidant à composer des mélodies, construire des progressions harmoniques et éditer des notes MIDI, tout en vous laissant le contrôle total pour ajuster, affiner et perfectionner chaque détail.
K.G.One Logo @@ -29,6 +29,15 @@ K.G.Studio est une DAW légère et moderne qui fonctionne entièrement dans le n ## Dernières mises à jour +- **2026.06.05** : développement majeur du **K.G.Studio Musician Assistant**, transformé en agent IA de niveau projet complet : + - **Outils de gestion de pistes** — l'agent peut désormais lister, créer, mettre à jour et supprimer des pistes, et parcourir tous les instruments disponibles, sans nécessiter de sélectionner une région au préalable. + - **Outils de pistes globales** — accès complet en lecture/écriture/suppression aux quatre pistes globales : **Chord Progression**, **Tempo (BPM)**, **Key Signature** et **Markers**. L'agent peut restructurer l'ensemble du squelette harmonique et rythmique d'un arrangement en une seule conversation. + - **Confirmation des outils** — les opérations d'écriture affichent une étape de confirmation dans le chat avant d'être exécutées, vous donnant la possibilité de vérifier avant que quoi que ce soit ne change. + - **Liste de tâches de l'agent** — l'assistant maintient désormais une liste de tâches en ligne, affichée sous forme de cartes de snapshot en direct directement dans le chat, rendant les plans en plusieurs étapes transparents et traçables. + - **Historique des conversations** — les sessions de chat sont persistées par projet et peuvent être reprises après rechargement de la page. + - **Compactage automatique du contexte** — les longues conversations sont résumées automatiquement lorsque la limite de contexte approche, permettant aux sessions de continuer sans intervention manuelle. + - **Mode agent efficace** — un prompt simplifié et un ensemble d'outils réduit pour les modèles de langage plus petits ou locaux, activé automatiquement lors de l'utilisation du Local Browser LLM. + - **2026.05.30** : ajout de la **prise en charge de l'internationalisation (i18n)** — K.G.Studio est désormais disponible en quatre langues : **English**, **Simplified Chinese (简体中文)**, **Traditional Chinese (繁體中文)** et **Français**. La langue active se configure dans **Réglages ⚙️ → Général → Langue**, avec une option `Auto` qui détecte automatiquement la langue de votre navigateur. - **2026.05.27** : @@ -49,13 +58,8 @@ K.G.Studio est une DAW légère et moderne qui fonctionne entièrement dans le n
- **2026.05.02** : ajout de la **visualisation spectrogramme des pistes audio**. Les régions audio affichent désormais une superposition de spectrogramme en temps réel dans la grille des pistes. Ajout du **Piano Roll hybrid mode** : ouvrez le piano roll sur une région MIDI pendant que le spectrogramme d'une région audio voisine est affiché comme couche de référence, afin d'éditer les notes MIDI en fonction de la forme visuelle de l'audio. Ajout du **zoom avant/arrière dans le piano roll** avec conservation de la position de la vue pour qu'elle reste ancrée près de la tête de lecture actuelle. Ajout également du **réglage fin de la position des régions** avec petits incréments pour un placement précis. Enfin, ajout de la synchronisation du défilement de la tête de lecture entre composants afin que la grille principale et le piano roll restent synchronisés pendant la lecture. -- **2026.04.29** : ajout de **Remix** et **Repaint** au panneau K.G.One Music Generator (alimenté par ACE-Step 1.5). **Remix** vous permet de refaire une région audio existante dans un nouveau style : sélectionnez une région audio, décrivez le style cible et fournissez éventuellement de nouvelles paroles, et ACE-Step réinterprétera le morceau avec l'instrumentation et l'ambiance demandées. **Repaint** vous permet de régénérer de façon chirurgicale une section précise d'un morceau : définissez une plage de boucle sur la timeline pour créer la fenêtre de repaint, puis décrivez le résultat souhaité pour cette section ; le reste du morceau reste inchangé. Les deux outils prennent en charge le même flux d'import que les autres onglets K.G.One : prévisualisation dans le lecteur intégré, glisser-déposer sur une piste, ou clic sur **Import Aligned to Source** pour placer automatiquement le résultat sous la région d'origine dans une nouvelle piste. -- **2026.04.24** : ajout de l'intégration [**K.G.One Music Studio**](https://github.com/KGAudioLab/K.G.One) ! Lorsque K.G.Studio se connecte à un serveur K.G.One local, le panneau **K.G.One Music Generator** (bouton baguette magique ✦ dans la barre d'outils) devient disponible avec trois outils alimentés par l'IA : **Full Song Generation** (alimenté par ACE-Step 1.5 pour générer des chansons complètes à partir de prompts textuels), **Clip Generation** (alimenté par Foundation-1 pour générer des clips instrumentaux et des boucles MIDI à partir de texte), et **Stem Separation** (alimenté par python-audio-separator pour séparer n'importe quel audio en voix, instrumentaux, etc.). Les audios et MIDI générés peuvent être prévisualisés instantanément et glissés directement sur vos pistes. K.G.One fonctionne entièrement sur votre propre machine (Windows/Linux, GPU CUDA requis) ; consultez le [dépôt K.G.One](https://github.com/KGAudioLab/K.G.One) pour les instructions d'installation. -- **2026.04.11** : migration du stockage des projets depuis IndexedDB vers OPFS (Origin Private File System) avec une structure basée sur des dossiers, afin de mieux gérer les fichiers médias. Ajout de la prise en charge des pistes audio avec import WAV/MP3, lecture, boucle et découpe non destructive des régions. Ajout également de l'export bounce vers WAV/MP3 via rendu hors ligne. -- **2026.04.05** : migration de l'agent IA, passant d'un appel d'outils basé sur XML au function calling natif du SDK OpenAI afin d'améliorer la fiabilité et la compatibilité. Ajout de nouvelles options de modèles LLM, y compris la série GPT-5.4. -- **2026.01.23** : implémentation de la lecture en boucle sans couture ! Faites glisser sur les numéros de mesure pour définir une plage de boucle, ou activez le mode boucle avec le bouton Loop dans la barre d'outils. La lecture en boucle utilise la fonctionnalité native de boucle de `Tone.js` pour une boucle précise au niveau de l'échantillon et sans coupure. -- **2025.12.21** : implémentation de la prise en charge du clavier MIDI ! Vous pouvez désormais connecter un clavier MIDI et l'utiliser pour jouer des sons. Veuillez noter que cette fonction peut ne pas fonctionner de manière optimale dans Safari et certains autres navigateurs ne prenant pas entièrement en charge l'interface Web MIDI. -- **2025.12.15** : ajout de l'assistant d'accords intelligent avec guidage harmonique fonctionnel (T/S/D). Survolez les touches du piano pour voir des suggestions d'accords adaptées au contexte et créer des accords complets en un clic ! + +Pour consulter l'historique complet des versions, voir les [**Notes de version**](./docs/RELEASE_NOTES.md). ## État du projet @@ -120,7 +124,7 @@ K.G.Studio peut exécuter **Gemma 4 E4B** directement dans votre navigateur grâ - Dans **OpenAI Compatible Server → Base URL**, saisissez `https://openrouter.ai/api/v1`. **Conseils :** -- Vous pouvez également utiliser l'API officielle d'OpenAI, d'autres services compatibles OpenAI, ou un serveur LLM auto-hébergé (par exemple Ollama, vLLM). Notez que la qualité des modèles varie : tous les modèles ne sont pas aussi performants pour les tâches d'édition musicale. Pour l'hébergement local, nous recommandons `qwen3.5-35b-a3b` comme bon équilibre entre qualité de génération et exigences matérielles. +- Vous pouvez également utiliser l'API officielle d'OpenAI, d'autres services compatibles OpenAI, ou un serveur LLM auto-hébergé (par exemple Ollama, vLLM). Notez que la qualité des modèles varie : tous les modèles ne sont pas aussi performants pour les tâches d'édition musicale. Pour l'auto-déploiement (nécessite ~24 Go de VRAM ou 24-32 Go de mémoire unifiée), nous recommandons : `qwen/qwen3.6-35b-a3b`, `google/gemma-4-26b-a4b-it` ou `google/gemma-4-31b-it`. - Si vous disposez d'un abonnement actif chez OpenAI ou chez un autre fournisseur LLM, vous pouvez utiliser [CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI) pour exécuter un serveur proxy local qui achemine les requêtes via votre abonnement existant, sans nécessiter de clé API séparée. ### Opérations DAW de base @@ -277,7 +281,7 @@ Remarque : en raison des limitations CORS chez certains fournisseurs, Google Gem 1. Obtenez une clé API OpenAI auprès de [**OpenAI**](https://platform.openai.com/account/api-keys). Vous devrez peut-être créer un compte et ajouter un moyen de paiement pour générer une clé API. 2. Dans **Settings ⚙️ → General → LLM Provider**, sélectionnez **OpenAI** comme fournisseur. 3. Saisissez votre clé API dans **OpenAI → Key**. -4. Sélectionnez votre modèle préféré dans la liste **OpenAI → Model**. Pour un bon équilibre entre performances et coût, nous recommandons `gpt-5.4-mini`. +4. Sélectionnez votre modèle préféré dans la liste **OpenAI → Model**. Pour un bon équilibre entre performances et coût, nous recommandons `gpt-5.4`. 5. Vous pouvez également choisir d'activer Flex Mode dans **OpenAI → Flex Mode**. Flex Mode offre un tarif réduit, mais peut entraîner des temps de réponse plus lents ou des erreurs côté serveur. ### Utiliser OpenRouter @@ -291,9 +295,14 @@ OpenRouter est une plateforme qui fournit un accès unifié à un large éventai **Remarque :** chaque fournisseur de modèle peut avoir des politiques différentes en matière de conservation des données et de confidentialité. Veuillez les consulter avant utilisation. 5. Saisissez le nom du modèle choisi dans **OpenAI Compatible Server → Model**. Les séries recommandées comprennent : - `Anthropic: Claude Sonnet 4.6` (`anthropic/claude-sonnet-4.6`: [Link](https://openrouter.ai/anthropic/claude-sonnet-4.6)) — meilleur équilibre qualité/coût pour la série Claude - - `Qwen: Qwen3.5-35B-A3B` (`qwen/qwen3.5-35b-a3b`: [Link](https://openrouter.ai/qwen/qwen3.5-35b-a3b)) — modèle open source recommandé - - `Qwen: Qwen3-Next-80B-A3B` (MODÈLE GRATUIT : `qwen/qwen3-next-80b-a3b-instruct:free`: [Link](https://openrouter.ai/qwen/qwen3-next-80b-a3b-instruct:free)) — modèle gratuit recommandé - - `OpenAI: GPT-OSS 120B` (MODÈLE GRATUIT : `openai/gpt-oss-120b:free`: [Link](https://openrouter.ai/openai/gpt-oss-120b:free)) — modèle gratuit recommandé + - Modèles gratuits : + - `OpenAI: GPT-OSS 120B` (MODÈLE GRATUIT : `openai/gpt-oss-120b:free`: [Link](https://openrouter.ai/openai/gpt-oss-120b:free)) + - `Google: Gemma 4 26B A4B IT` (MODÈLE GRATUIT : `google/gemma-4-26b-a4b-it:free`: [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it:free)) + - `Google: Gemma 4 31B IT` (MODÈLE GRATUIT : `google/gemma-4-31b-it:free`: [Link](https://openrouter.ai/google/gemma-4-31b-it:free)) + - Pour l'auto-déploiement (nécessite ~24 Go de VRAM ou 24-32 Go de mémoire unifiée), nous recommandons : + - `Qwen: Qwen3.6 35B A3B` (`qwen/qwen3.6-35b-a3b`: [Link](https://openrouter.ai/qwen/qwen3.6-35b-a3b)) + - `Google: Gemma 4 26B A4B IT` (`google/gemma-4-26b-a4b-it`: [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it)) + - `Google: Gemma 4 31B IT` (`google/gemma-4-31b-it`: [Link](https://openrouter.ai/google/gemma-4-31b-it)) - Remarque : les fournisseurs de modèles gratuits peuvent collecter vos données ; consultez la page du modèle avant utilisation - Remarque : la disponibilité des modèles gratuits change fréquemment ; pour les options gratuites les plus récentes, consultez la [page des modèles OpenRouter](https://openrouter.ai/models) et utilisez le filtre **Prompt Pricing** 6. Saisissez l'URL de base `https://openrouter.ai/api/v1` dans **OpenAI Compatible Server → Base URL**. diff --git a/README-zh_cn.md b/README-zh_cn.md index a77f733..c4405c9 100644 --- a/README-zh_cn.md +++ b/README-zh_cn.md @@ -19,7 +19,7 @@ K.G.Studio 是一款轻量、现代化的 DAW,完全运行于浏览器中,并以 **K.G.Studio 音乐创作助手** 为核心。它提供基于 Tone.js sampler 的真实乐器回放、钢琴卷帘编辑器、支持完整撤销/重做的音轨与区域管理、基于 OPFS(Origin Private File System)的项目持久化、可配置的设置面板,以及可执行工具的内置 AI 助手。 -**K.G.Studio 音乐创作助手** 是一个面向和声、编曲与音符编辑的 AI 助手,但并不负责整首作品的全自动作曲。 +**K.G.Studio 音乐创作助手** 是一款具备项目感知能力的 AI 协同创作助手。它并不直接生成音频文件(如 WAV 格式),而是直接在结构化的音轨和音符层级进行操作——帮助您编写旋律、构建和弦进行以及编辑 MIDI 音符,同时将完整的控制权留给您,方便您后续轻松调整、微调和雕琢每一个音乐细节。
K.G.One Logo @@ -29,6 +29,15 @@ K.G.Studio 是一款轻量、现代化的 DAW,完全运行于浏览器中, ## 最新更新 +- **2026.06.05**: 大幅扩展 **K.G.Studio 音乐创作助手**,升级为完整的项目级 AI Agent: + - **音轨管理工具** — Agent 现在无需选择区域,即可列出、创建、更新和删除音轨,以及浏览所有可用乐器。 + - **全局轨工具** — 完整读取/写入/删除四条全局轨:**和弦进行**、**速度(BPM)**、**调号** 和 **Marker**。Agent 可在一次对话中重构整首编曲的和声与节奏骨架。 + - **工具确认机制** — 写操作在执行前会在聊天中显示确认步骤,让您在修改生效前有机会审查。 + - **Agent 待办清单** — 助手现在会在聊天中维护内联任务清单,以实时快照卡片的形式呈现,让多步骤计划一目了然。 + - **对话历史** — 聊天会话会按项目持久化保存,并可跨页面刷新恢复。 + - **自动上下文压缩** — 当对话接近上下文限制时,旧消息会自动摘要压缩,保证长会话持续运转。 + - **高效 Agent 模式** — 为较小/本地语言模型提供精简提示词与工具集,使用 Local Browser LLM 时自动启用。 + - **2026.05.30**: 新增 **国际化(i18n)支持** — K.G.Studio 现已提供四种语言版本:**English**、**简体中文**、**繁體中文** 和 **Français**。可在 **设置 ⚙️ → 通用 → 语言** 中配置首选语言,选择 `Auto` 时将自动检测浏览器语言。 - **2026.05.27**: @@ -49,13 +58,8 @@ K.G.Studio 是一款轻量、现代化的 DAW,完全运行于浏览器中,
- **2026.05.02**: 新增 **音频轨频谱可视化**。音频区域现在会在时间网格中显示实时频谱叠加层。另新增 **Piano Roll hybrid mode**:当您在 MIDI 区域中打开钢琴卷帘时,可将相邻音频区域的频谱作为参考层显示,从而参照音频形状编辑 MIDI 音符。新增 **钢琴卷帘缩放**,并保留当前视口位置,使画面始终锚定在当前播放头附近。还新增 **区域微调位置**,可用小步长推动区域以实现精确摆放;同时新增跨组件的播放头滚动同步,使主网格与钢琴卷帘在播放期间保持联动。 -- **2026.04.29**: 在 K.G.One Music Generator 面板中新增 **Remix** 和 **Repaint**(由 ACE-Step 1.5 驱动)。**Remix** 可让您用新风格重制现有音频区域,您只需选中音频区域、描述目标风格,并可选填写新歌词,ACE-Step 就会按提示重新演绎歌曲的配器与氛围。**Repaint** 可对歌曲中特定片段进行局部再生成,您可以先在时间线上设置 loop 范围作为重绘窗口,再描述该片段应有的声音效果,其余部分将保持不变。两项功能都支持与其他 K.G.One 标签页相同的导入流程:可在内置播放器中预览、拖拽到音轨上,或点击 **Import Aligned to Source**,自动将结果放置到原始区域下方的新音轨中。 -- **2026.04.24**: 新增 [**K.G.One Music Studio**](https://github.com/KGAudioLab/K.G.One) 集成。当 K.G.Studio 连接到本地 K.G.One 服务器后,**K.G.One Music Generator** 面板(工具栏中的魔杖按钮 ✦)即可启用,其中包含三项 AI 驱动工具:**Full Song Generation**(由 ACE-Step 1.5 驱动,可根据文本提示生成整首歌曲)、**Clip Generation**(由 Foundation-1 驱动,可根据文本生成乐器片段与 MIDI loop),以及 **Stem Separation**(由 python-audio-separator 驱动,可将任意音频拆分为人声、伴奏等多个 stem)。生成的音频和 MIDI 均可即时预览,并可直接拖拽到您的音轨中。K.G.One 完全运行在您自己的机器上(Windows/Linux,需要 CUDA GPU);详细设置说明请参阅 [K.G.One 仓库](https://github.com/KGAudioLab/K.G.One)。 -- **2026.04.11**: 将项目存储从 IndexedDB 迁移到 OPFS(Origin Private File System),并采用基于文件夹的结构,以更好地处理媒体文件。新增音频轨支持,包括 WAV/MP3 导入、回放、循环,以及非破坏性区域裁剪。另新增基于离线渲染的 WAV/MP3 导出功能。 -- **2026.04.05**: 将 AI agent 从基于 XML 的工具调用迁移到原生 OpenAI SDK function calling,以提升可靠性与兼容性。另新增 GPT-5.4 系列等 LLM 模型选项。 -- **2026.01.23**: 实现无缝循环播放。您可以拖动小节编号设置 loop 范围,或通过工具栏中的 Loop 按钮切换循环模式。循环播放基于 `Tone.js` 的原生 loop 机制,可实现采样级精确、无缝衔接的循环。 -- **2025.12.21**: 实现 MIDI 键盘支持。您现在可以连接 MIDI 键盘并直接用其演奏声音。请注意,由于 Safari 和部分其他浏览器对 Web MIDI 接口支持不完整,此功能在这些浏览器中可能无法达到最佳效果。 -- **2025.12.15**: 新增智能和弦助手,支持功能和声指导(T/S/D)。将鼠标悬停在琴键上即可查看上下文相关的和弦建议,并可一键创建完整和弦。 + +查看完整版本历史,请参阅 [**发布说明**](./docs/RELEASE_NOTES.md)。 ## 项目状态 @@ -120,7 +124,7 @@ K.G.Studio 可以借助 WebGPU 加速,在浏览器中直接运行 **Gemma 4 E4 - 在 **OpenAI Compatible Server → 基础 URL** 中输入 `https://openrouter.ai/api/v1`。 **提示:** -- 您也可以使用官方 OpenAI API、其他 OpenAI 兼容服务,或自托管 LLM 服务器(如 Ollama、vLLM)。请注意,不同模型的质量差异较大,并非所有模型都同样适合音乐编辑任务。对于本地部署,我们推荐 `qwen3.5-35b-a3b`,它在生成质量与硬件需求之间取得了较好的平衡。 +- 您也可以使用官方 OpenAI API、其他 OpenAI 兼容服务,或自托管 LLM 服务器(如 Ollama、vLLM)。请注意,不同模型的质量差异较大,并非所有模型都同样适合音乐编辑任务。对于自托管/本地部署(需要约 24G 显存或 24-32GB 统一内存),我们推荐:`qwen/qwen3.6-35b-a3b`、`google/gemma-4-26b-a4b-it` 或 `google/gemma-4-31b-it`。 - 如果您已订阅 OpenAI 或其他 LLM 提供方,可以使用 [CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI) 运行一个本地代理服务器,通过现有订阅转发请求,而无需单独准备 API Key。 ### 基本 DAW 操作 @@ -279,7 +283,7 @@ K.G.Studio 会从 `./public/config.json` 加载默认配置(内部也提供回 1. 在 [**OpenAI**](https://platform.openai.com/account/api-keys) 获取 OpenAI API Key。您可能需要先注册账号并添加支付方式,才能生成 API Key。 2. 在 **设置 ⚙️ → 通用 → LLM 提供方** 中选择 **OpenAI** 作为提供方。 3. 在 **OpenAI → 密钥** 中输入您的 API Key。 -4. 在 **OpenAI → 模型** 下拉中选择您偏好的模型。若希望在性能与成本之间取得较好平衡,我们推荐 `gpt-5.4-mini`。 +4. 在 **OpenAI → 模型** 下拉中选择您偏好的模型。若希望在性能与成本之间取得较好平衡,我们推荐 `gpt-5.4`。 5. 您也可以选择是否在 **OpenAI → Flex 模式** 中启用 Flex Mode。Flex Mode 可以降低价格,但也可能带来更慢的响应时间或更多服务端错误。 ### 使用 OpenRouter @@ -293,9 +297,14 @@ OpenRouter 是一个统一接入平台,可让您访问来自多个提供方的 **注意:** 不同模型提供方的数据保留与隐私政策可能不同,请在使用前自行查看。 5. 在 **OpenAI Compatible Server → 模型** 中输入您所选的模型名。推荐系列包括: - `Anthropic: Claude Sonnet 4.6` (`anthropic/claude-sonnet-4.6`: [Link](https://openrouter.ai/anthropic/claude-sonnet-4.6)) — Claude 系列中质量与成本平衡较好的选择 - - `Qwen: Qwen3.5-35B-A3B` (`qwen/qwen3.5-35b-a3b`: [Link](https://openrouter.ai/qwen/qwen3.5-35b-a3b)) — 推荐的开源模型 - - `Qwen: Qwen3-Next-80B-A3B`(免费模型:`qwen/qwen3-next-80b-a3b-instruct:free`: [Link](https://openrouter.ai/qwen/qwen3-next-80b-a3b-instruct:free))— 推荐的免费模型 - - `OpenAI: GPT-OSS 120B`(免费模型:`openai/gpt-oss-120b:free`: [Link](https://openrouter.ai/openai/gpt-oss-120b:free))— 推荐的免费模型 + - 免费模型: + - `OpenAI: GPT-OSS 120B`(免费模型:`openai/gpt-oss-120b:free`: [Link](https://openrouter.ai/openai/gpt-oss-120b:free)) + - `Google: Gemma 4 26B A4B IT`(免费模型:`google/gemma-4-26b-a4b-it:free`: [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it:free)) + - `Google: Gemma 4 31B IT`(免费模型:`google/gemma-4-31b-it:free`: [Link](https://openrouter.ai/google/gemma-4-31b-it:free)) + - 对于自托管/本地部署(需要约 24G 显存或 24-32GB 统一内存),我们推荐: + - `Qwen: Qwen3.6 35B A3B` (`qwen/qwen3.6-35b-a3b`: [Link](https://openrouter.ai/qwen/qwen3.6-35b-a3b)) + - `Google: Gemma 4 26B A4B IT` (`google/gemma-4-26b-a4b-it`: [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it)) + - `Google: Gemma 4 31B IT` (`google/gemma-4-31b-it`: [Link](https://openrouter.ai/google/gemma-4-31b-it)) - 注意:免费模型提供方可能会收集您的数据,使用前请先查看模型页面说明 - 注意:免费模型的可用性变化频繁。如需查看最新免费选项,请访问 [OpenRouter Models Page](https://openrouter.ai/models),并使用 **Prompt Pricing** 过滤当前免费模型 6. 在 **OpenAI Compatible Server → 基础 URL** 中填写 `https://openrouter.ai/api/v1`。 diff --git a/README-zh_hk.md b/README-zh_hk.md index 781f66f..bb811df 100644 --- a/README-zh_hk.md +++ b/README-zh_hk.md @@ -19,7 +19,7 @@ K.G.Studio 是一款輕量、現代化的 DAW,完全執行於瀏覽器中,並以 **K.G.Studio 音樂創作助手** 為核心。它提供基於 Tone.js sampler 的真實樂器回放、Piano Roll 編輯器、支援完整復原/重做的音軌與區域管理、基於 OPFS(Origin Private File System)的專案持久化、可設定的設定面板,以及可執行工具的內建 AI 助手。 -**K.G.Studio 音樂創作助手** 是一個面向和聲、編曲與音符編輯的 AI 助手,但並不負責整首作品的全自動作曲。 +**K.G.Studio 音樂創作助手** 是一款具備專案感知能力的 AI 協同創作助手。它並不直接生成音訊檔案(如 WAV 格式),而是直接在結構化的音軌和音符層級進行操作——幫助您編寫旋律、構建和聲進行以及編輯 MIDI 音符,同時將完整的控制權留給您,方便您後續輕鬆調整、微調和雕琢每一個音樂細節。
K.G.One Logo @@ -29,6 +29,15 @@ K.G.Studio 是一款輕量、現代化的 DAW,完全執行於瀏覽器中, ## 最新更新 +- **2026.06.05**: 大幅擴展 **K.G.Studio 音樂創作助手**,升級為完整的專案級 AI Agent: + - **音軌管理工具** — Agent 現在無需選取區域,即可列出、建立、更新和刪除音軌,以及瀏覽所有可用樂器。 + - **全域軌工具** — 完整讀取/寫入/刪除四條全域軌:**和弦進行**、**速度(BPM)**、**調號** 和 **Marker**。Agent 可在一次對話中重構整首編曲的和聲與節奏骨架。 + - **工具確認機制** — 寫操作在執行前會在聊天中顯示確認步驟,讓您在修改生效前有機會審查。 + - **Agent 待辦清單** — 助手現在會在聊天中維護內聯任務清單,以即時快照卡片的形式呈現,讓多步驟計劃一目了然。 + - **對話歷史** — 聊天會話會按專案持久化保存,並可跨頁面重新整理恢復。 + - **自動上下文壓縮** — 當對話接近上下文限制時,舊訊息會自動摘要壓縮,保證長對話持續運作。 + - **高效 Agent 模式** — 為較小/本地語言模型提供精簡提示詞與工具集,使用 Local Browser LLM 時自動啟用。 + - **2026.05.30**: 新增 **國際化(i18n)支援** — K.G.Studio 現已提供四種語言版本:**English**、**简体中文**、**繁體中文** 和 **Français**。可在 **設定 ⚙️ → 通用 → 語言** 中設定偏好語言,選擇 `Auto` 時將自動偵測瀏覽器語言。 - **2026.05.27**: @@ -49,13 +58,8 @@ K.G.Studio 是一款輕量、現代化的 DAW,完全執行於瀏覽器中,
- **2026.05.02**: 新增 **音訊軌頻譜可視化**。音訊區域現在會在時間網格中顯示即時頻譜疊加層。另新增 **Piano Roll hybrid mode**:當您在 MIDI 區域中開啟鋼琴卷簾時,可將相鄰音訊區域的頻譜作為參考層顯示,從而參照音訊形狀編輯 MIDI 音符。新增 **鋼琴卷簾縮放**,並保留目前視口位置,使畫面始終錨定在目前播放頭附近。還新增 **區域微調位置**,可用小步長推動區域以實現精確擺放;同時新增跨元件的播放頭捲動同步,使主網格與鋼琴卷簾在播放期間保持連動。 -- **2026.04.29**: 在 K.G.One Music Generator 面板中新增 **Remix** 和 **Repaint**(由 ACE-Step 1.5 驅動)。**Remix** 可讓您用新風格重製現有音訊區域,您只需選取音訊區域、描述目標風格,並可選填寫新歌詞,ACE-Step 就會按提示重新演繹歌曲的配器與氛圍。**Repaint** 可對歌曲中特定片段進行局部再生成,您可以先在時間線上設定 loop 範圍作為重繪視窗,再描述該片段應有的聲音效果,其餘部分將保持不變。兩項功能都支援與其他 K.G.One 分頁相同的匯入流程:可在內建播放器中預覽、拖曳到音軌上,或點擊 **Import Aligned to Source**,自動將結果放置到原始區域下方的新音軌中。 -- **2026.04.24**: 新增 [**K.G.One Music Studio**](https://github.com/KGAudioLab/K.G.One) 整合。當 K.G.Studio 連線到本地 K.G.One 伺服器後,**K.G.One Music Generator** 面板(工具列中的魔杖按鈕 ✦)即可啟用,其中包含三項 AI 驅動工具:**Full Song Generation**(由 ACE-Step 1.5 驅動,可根據文字提示生成整首歌曲)、**Clip Generation**(由 Foundation-1 驅動,可根據文字生成樂器片段與 MIDI loop),以及 **Stem Separation**(由 python-audio-separator 驅動,可將任意音訊拆分為人聲、伴奏等多個 stem)。生成的音訊和 MIDI 均可即時預覽,並可直接拖曳到您的音軌中。K.G.One 完全執行在您自己的機器上(Windows/Linux,需要 CUDA GPU);詳細設定說明請參閱 [K.G.One 倉庫](https://github.com/KGAudioLab/K.G.One)。 -- **2026.04.11**: 將專案儲存從 IndexedDB 遷移到 OPFS(Origin Private File System),並採用基於資料夾的結構,以更好地處理媒體檔案。新增音訊軌支援,包括 WAV/MP3 匯入、回放、循環,以及非破壞性區域裁剪。另新增基於離線渲染的 WAV/MP3 匯出功能。 -- **2026.04.05**: 將 AI agent 從基於 XML 的工具呼叫遷移到原生 OpenAI SDK function calling,以提升可靠性與相容性。另新增 GPT-5.4 系列等 LLM 模型選項。 -- **2026.01.23**: 實現無縫循環播放。您可以拖動小節編號設定 loop 範圍,或透過工具列中的 Loop 按鈕切換循環模式。循環播放基於 `Tone.js` 的原生 loop 機制,可實現取樣級精確、無縫銜接的循環。 -- **2025.12.21**: 實現 MIDI 鍵盤支援。您現在可以連接 MIDI 鍵盤並直接用其演奏聲音。請注意,由於 Safari 和部分其他瀏覽器對 Web MIDI 介面支援不完整,此功能在這些瀏覽器中可能無法達到最佳效果。 -- **2025.12.15**: 新增智慧和弦助手,支援功能和聲指導(T/S/D)。將滑鼠懸停在琴鍵上即可查看與上下文相關的和弦建議,並可一鍵建立完整和弦。 + +查看完整版本歷史,請參閱 [**發佈說明**](./docs/RELEASE_NOTES.md)。 ## 專案狀態 @@ -120,7 +124,7 @@ K.G.Studio 可以藉助 WebGPU 加速,在瀏覽器中直接執行 **Gemma 4 E4 - 在 **OpenAI Compatible Server → 基礎 URL** 中輸入 `https://openrouter.ai/api/v1`。 **提示:** -- 您也可以使用官方 OpenAI API、其他 OpenAI 相容服務,或自行託管的 LLM 伺服器(如 Ollama、vLLM)。請注意,不同模型的品質差異較大,並非所有模型都同樣適合音樂編輯任務。對於本地部署,我們推薦 `qwen3.5-35b-a3b`,它在生成品質與硬體需求之間取得了較好的平衡。 +- 您也可以使用官方 OpenAI API、其他 OpenAI 相容服務,或自行託管的 LLM 伺服器(如 Ollama、vLLM)。請注意,不同模型的品質差異較大,並非所有模型都同樣適合音樂編輯任務。對於自託管/本地部署(需要約 24G 顯存或 24-32GB 統一記憶體),我們推薦:`qwen/qwen3.6-35b-a3b`、`google/gemma-4-26b-a4b-it` 或 `google/gemma-4-31b-it`。 - 如果您已訂閱 OpenAI 或其他 LLM 提供方,可以使用 [CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI) 執行一個本地代理伺服器,透過現有訂閱轉發請求,而無需另外準備 API Key。 ### 基本 DAW 操作 @@ -277,7 +281,7 @@ K.G.Studio 會從 `./public/config.json` 載入預設設定(內部也提供回 1. 在 [**OpenAI**](https://platform.openai.com/account/api-keys) 取得 OpenAI API Key。您可能需要先註冊帳號並新增付款方式,才能生成 API Key。 2. 在 **設定 ⚙️ → 通用 → LLM 提供方** 中選擇 **OpenAI** 作為提供方。 3. 在 **OpenAI → 密鑰** 中輸入您的 API Key。 -4. 在 **OpenAI → 模型** 下拉中選擇您偏好的模型。若希望在效能與成本之間取得較好平衡,我們推薦 `gpt-5.4-mini`。 +4. 在 **OpenAI → 模型** 下拉中選擇您偏好的模型。若希望在效能與成本之間取得較好平衡,我們推薦 `gpt-5.4`。 5. 您也可以選擇是否在 **OpenAI → Flex 模式** 中啟用 Flex Mode。Flex Mode 可以降低價格,但也可能帶來更慢的回應時間或更多伺服器端錯誤。 ### 使用 OpenRouter @@ -291,9 +295,14 @@ OpenRouter 是一個統一接入平台,可讓您存取來自多個提供方的 **注意:** 不同模型提供方的資料保留與隱私政策可能不同,請在使用前自行查看。 5. 在 **OpenAI Compatible Server → 模型** 中輸入您所選的模型名。推薦系列包括: - `Anthropic: Claude Sonnet 4.6` (`anthropic/claude-sonnet-4.6`: [Link](https://openrouter.ai/anthropic/claude-sonnet-4.6)) — Claude 系列中品質與成本平衡較好的選擇 - - `Qwen: Qwen3.5-35B-A3B` (`qwen/qwen3.5-35b-a3b`: [Link](https://openrouter.ai/qwen/qwen3.5-35b-a3b)) — 推薦的開源模型 - - `Qwen: Qwen3-Next-80B-A3B`(免費模型:`qwen/qwen3-next-80b-a3b-instruct:free`: [Link](https://openrouter.ai/qwen/qwen3-next-80b-a3b-instruct:free))— 推薦的免費模型 - - `OpenAI: GPT-OSS 120B`(免費模型:`openai/gpt-oss-120b:free`: [Link](https://openrouter.ai/openai/gpt-oss-120b:free))— 推薦的免費模型 + - 免費模型: + - `OpenAI: GPT-OSS 120B`(免費模型:`openai/gpt-oss-120b:free`: [Link](https://openrouter.ai/openai/gpt-oss-120b:free)) + - `Google: Gemma 4 26B A4B IT`(免費模型:`google/gemma-4-26b-a4b-it:free`: [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it:free)) + - `Google: Gemma 4 31B IT`(免費模型:`google/gemma-4-31b-it:free`: [Link](https://openrouter.ai/google/gemma-4-31b-it:free)) + - 對於自託管/本地部署(需要約 24G 顯存或 24-32GB 統一記憶體),我們推薦: + - `Qwen: Qwen3.6 35B A3B` (`qwen/qwen3.6-35b-a3b`: [Link](https://openrouter.ai/qwen/qwen3.6-35b-a3b)) + - `Google: Gemma 4 26B A4B IT` (`google/gemma-4-26b-a4b-it`: [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it)) + - `Google: Gemma 4 31B IT` (`google/gemma-4-31b-it`: [Link](https://openrouter.ai/google/gemma-4-31b-it)) - 注意:免費模型提供方可能會蒐集您的資料,使用前請先查看模型頁面說明 - 注意:免費模型的可用性變化頻繁。如需查看最新免費選項,請造訪 [OpenRouter Models Page](https://openrouter.ai/models),並使用 **Prompt Pricing** 過濾目前免費模型 6. 在 **OpenAI Compatible Server → 基礎 URL** 中填入 `https://openrouter.ai/api/v1`。 diff --git a/README.md b/README.md index 683763f..b06cd94 100644 --- a/README.md +++ b/README.md @@ -19,7 +19,7 @@ English | [Français](./README-fr.md) | [简体中文](./README-zh_cn.md) | [繁 K.G.Studio is a lightweight, modern DAW that runs entirely in the browser with **K.G.Studio Musician Assistant** at its core. It features realistic instrument playback via Tone.js samplers, a piano‑roll editor, track and region management with full undo/redo, project persistence to OPFS (Origin Private File System), a configurable settings panel, and an integrated AI assistant with tool execution. -**K.G.Studio Musician Assistant** is an AI assistance agent for harmony, arrangement, and note editing — but not full auto‑composition. +**K.G.Studio Musician Assistant** is a project-aware AI co-creator designed to elevate your creative workflow. Rather than generating raw audio files, it operates directly at the structured track and note level—helping you draft melodies, build harmonic progressions, and edit MIDI notes, while leaving you in complete control to adjust, fine-tune, and perfect every single detail.
K.G.One Logo @@ -29,6 +29,15 @@ K.G.Studio is a lightweight, modern DAW that runs entirely in the browser with * ## Latest Updates +- **2026.06.05**: Significantly expanded the **K.G.Studio Musician Assistant** into a full project-level agent: + - **Track management tools** — the agent can now list, create, update, and delete tracks, and browse all available instruments, without requiring a region to be selected first. + - **Global track tools** — full read/write/remove access to all four global tracks: **Chord Progression**, **Tempo (BPM)**, **Key Signature**, and **Markers**. The agent can restructure an entire arrangement's harmonic and rhythmic skeleton in a single conversation. + - **Tool confirmation** — write operations surface a confirmation step in the chat before executing, giving you a chance to review before anything changes. + - **Agent todo list** — the assistant now maintains an inline task checklist rendered as live snapshot cards directly in the chat, making multi-step plans transparent and trackable. + - **Conversation history** — chat sessions are persisted per project and can be resumed across page reloads. + - **Automatic context compaction** — long conversations are summarised automatically when the context limit approaches, keeping sessions running without manual intervention. + - **Efficient agent mode** — a streamlined prompt and reduced tool set for smaller / local language models, activated automatically when using the Local Browser LLM. + - **2026.05.30**: Added **internationalization (i18n) support** — K.G.Studio now ships in four languages: **English**, **Simplified Chinese (简体中文)**, **Traditional Chinese (繁體中文)**, and **French (Français)**. The active language can be configured under **Settings ⚙️ → General → Language**, with an `Auto` option that automatically detects your browser's locale. - **2026.05.27**: @@ -49,13 +58,8 @@ K.G.Studio is a lightweight, modern DAW that runs entirely in the browser with *
- **2026.05.02**: Added **audio track spectrogram visualization** — audio regions now display a real-time spectrogram overlay in the track grid. Added **Piano Roll hybrid mode**: open the piano roll on a MIDI region while an adjacent audio region's spectrogram is shown as a reference layer, letting you edit MIDI notes against the visual shape of the audio. Added **piano roll zoom in/out** with viewport-position preservation so the view stays anchored to the current playhead. Added **fine-tune region position**: nudge regions by small increments for precise placement. Also added cross-component playhead scroll synchronization so the main grid and piano roll stay in sync during playback. -- **2026.04.29**: Added **Remix** and **Repaint** to the K.G.One Music Generator panel (powered by ACE-Step 1.5). **Remix** lets you cover an existing audio region in a new style — select an audio region, describe the target style and optionally provide new lyrics, and ACE-Step will re-perform the song with the prompted instrumentation and feel. **Repaint** lets you surgically re-generate a specific section of a song — set a loop range on the timeline to define the repaint window, then describe what you want that section to sound like; the rest of the song stays untouched. Both tools support the same import workflow as the other K.G.One tabs: preview the result in the built-in player, drag it onto a track, or click **Import Aligned to Source** to automatically place it below the original region in a new track. -- **2026.04.24**: Added [**K.G.One Music Studio**](https://github.com/KGAudioLab/K.G.One) integration! When K.G.Studio connects to a local K.G.One server, the **K.G.One Music Generator** panel (magic wand button ✦ in the toolbar) becomes available with three AI-powered tools: **Full Song Generation** (powered by ACE-Step 1.5 — generate full-length songs from text prompts), **Clip Generation** (powered by Foundation-1 — generate instrument clips and MIDI loops from text), and **Stem Separation** (powered by python-audio-separator — split any audio into vocals, instrumentals, and more). Generated audio and MIDI can be previewed instantly and dragged directly onto your tracks. K.G.One runs entirely on your own machine (Windows/Linux, CUDA GPU required); see the [K.G.One repository](https://github.com/KGAudioLab/K.G.One) for setup instructions. -- **2026.04.11**: Migrated project storage from IndexedDB to OPFS (Origin Private File System) with a folder-based structure for better media file handling. Added audio track support with WAV/MP3 import, playback, looping, and non-destructive region trimming. Added bounce-to-WAV/MP3 export via offline rendering. -- **2026.04.05**: Migrated the AI agent from XML-based tool calling to native OpenAI SDK function calling for improved reliability and compatibility. Added new LLM model options including GPT-5.4 series. -- **2026.01.23**: Implemented seamless loop playback! Drag on the bar numbers to set loop range, or toggle loop mode with the Loop button in the toolbar. Loop playback uses `Tone.js`'s native looping for sample-accurate, gap-free looping. -- **2025.12.21**: Implemented MIDI keyboard support! You can now connect a MIDI keyboard and use it to play sounds. Please note that this feature may not work optimally in Safari and some other browsers that lack complete Web MIDI interface support. -- **2025.12.15**: Added Intelligent Chord Assistant with functional harmony guidance (T/S/D). Hover over piano keys to see context-aware chord suggestions and create full chords with one click! + +For the full release history, see [**Release Notes**](./docs/RELEASE_NOTES.md). ## Project Status @@ -120,7 +124,7 @@ K.G.Studio can run **Gemma 4 E4B** entirely inside your browser using WebGPU acc - In **OpenAI Compatible Server → Base URL**, enter `https://openrouter.ai/api/v1`. **Tips:** -- You can also use the official OpenAI API, other OpenAI-compatible services, or a self-hosted LLM server (e.g., Ollama, vLLM). Note that model quality varies — not all models perform equally well for music editing tasks. For local hosting, we recommend `qwen3.5-35b-a3b` as a good balance between generation quality and hardware requirements. +- You can also use the official OpenAI API, other OpenAI-compatible services, or a self-hosted LLM server (e.g., Ollama, vLLM). Note that model quality varies — not all models perform equally well for music editing tasks. For self-deployment (requiring ~24G VRAM or 24-32GB Unified Memory), we recommend: `qwen/qwen3.6-35b-a3b`, `google/gemma-4-26b-a4b-it`, or `google/gemma-4-31b-it`. - If you have an active subscription with OpenAI or another LLM provider, you can use [CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI) to run a local proxy server that routes requests through your existing subscription, without needing a separate API key. ### Basic DAW operations @@ -277,7 +281,7 @@ Note: due to CORS limitations with some providers, Google Gemini and Anthropic C 1. Obtain an OpenAI API Key from [**OpenAI**](https://platform.openai.com/account/api-keys). You may need to create an account and add a payment method to generate an API Key. 2. In **Settings ⚙️ → General → LLM Provider**, select **OpenAI** as your provider. 3. Enter your API Key in **OpenAI → Key**. -4. Select your preferred model from the **OpenAI → Model** dropdown. For a good balance between performance and cost, we recommend `gpt-5.4-mini`. +4. Select your preferred model from the **OpenAI → Model** dropdown. For a good balance between performance and cost, we recommend `gpt-5.4`. 5. Optionally, choose whether to enable Flex Mode in **OpenAI → Flex Mode**. Flex Mode offers discounted pricing, but may result in slower response times or server-side errors. ### Using OpenRouter @@ -291,9 +295,14 @@ OpenRouter is a platform that provides unified access to a wide range of languag **Note:** Each model provider may have different data retention and privacy policies. Please review these policies before use. 5. Enter your chosen model name in **OpenAI Compatible Server → Model**. Recommended model series include: - `Anthropic: Claude Sonnet 4.6` (`anthropic/claude-sonnet-4.6`: [Link](https://openrouter.ai/anthropic/claude-sonnet-4.6)) — best balance of quality and cost for the Claude series - - `Qwen: Qwen3.5-35B-A3B` (`qwen/qwen3.5-35b-a3b`: [Link](https://openrouter.ai/qwen/qwen3.5-35b-a3b)) — recommended open source model - - `Qwen: Qwen3-Next-80B-A3B` (FREE MODEL: `qwen/qwen3-next-80b-a3b-instruct:free`: [Link](https://openrouter.ai/qwen/qwen3-next-80b-a3b-instruct:free)) — recommended free model - - `OpenAI: GPT-OSS 120B` (FREE MODEL: `openai/gpt-oss-120b:free`: [Link](https://openrouter.ai/openai/gpt-oss-120b:free)) — recommended free model + - Free Models: + - `OpenAI: GPT-OSS 120B` (FREE MODEL: `openai/gpt-oss-120b:free`: [Link](https://openrouter.ai/openai/gpt-oss-120b:free)) + - `Google: Gemma 4 26B A4B IT` (FREE MODEL: `google/gemma-4-26b-a4b-it:free`: [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it:free)) + - `Google: Gemma 4 31B IT` (FREE MODEL: `google/gemma-4-31b-it:free`: [Link](https://openrouter.ai/google/gemma-4-31b-it:free)) + - For self-deployment (requiring ~24G VRAM or 24-32GB Unified Memory), we recommend: + - `Qwen: Qwen3.6 35B A3B` (`qwen/qwen3.6-35b-a3b`: [Link](https://openrouter.ai/qwen/qwen3.6-35b-a3b)) + - `Google: Gemma 4 26B A4B IT` (`google/gemma-4-26b-a4b-it`: [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it)) + - `Google: Gemma 4 31B IT` (`google/gemma-4-31b-it`: [Link](https://openrouter.ai/google/gemma-4-31b-it)) - Note: free model providers may collect your data; check the model page for details before use - Note: free model availability changes frequently — for the latest free options, visit the [OpenRouter Models Page](https://openrouter.ai/models) and use the **Prompt Pricing** filter to find currently free models 6. Input the base URL `https://openrouter.ai/api/v1` **OpenAI Compatible Server → Base URL**. diff --git a/docs/RELEASE_NOTES.md b/docs/RELEASE_NOTES.md new file mode 100644 index 0000000..f53341f --- /dev/null +++ b/docs/RELEASE_NOTES.md @@ -0,0 +1,49 @@ +# K.G.Studio — Release Notes + +This file contains the complete release history of K.G.Studio. For a summary of recent highlights, see the [main README](../README.md). + +--- + +- **2026.06.05**: Significantly expanded the **K.G.Studio Musician Assistant** into a full project-level agent: + - **Track management tools** — the agent can now list, create, update, and delete tracks, and browse all available instruments, without requiring a region to be selected first. + - **Global track tools** — full read/write/remove access to all four global tracks: **Chord Progression**, **Tempo (BPM)**, **Key Signature**, and **Markers**. The agent can restructure an entire arrangement's harmonic and rhythmic skeleton in a single conversation. + - **Tool confirmation** — write operations surface a confirmation step in the chat before executing, giving you a chance to review before anything changes. + - **Agent todo list** — the assistant now maintains an inline task checklist rendered as live snapshot cards directly in the chat, making multi-step plans transparent and trackable. + - **Conversation history** — chat sessions are persisted per project and can be resumed across page reloads. + - **Automatic context compaction** — long conversations are summarised automatically when the context limit approaches, keeping sessions running without manual intervention. + - **Efficient agent mode** — a streamlined prompt and reduced tool set for smaller / local language models, activated automatically when using the Local Browser LLM. + +- **2026.05.30**: Added **internationalization (i18n) support** — K.G.Studio now ships in four languages: **English**, **Simplified Chinese (简体中文)**, **Traditional Chinese (繁體中文)**, and **French (Français)**. The active language can be configured under **Settings ⚙️ → General → Language**, with an `Auto` option that automatically detects your browser's locale. + +- **2026.05.27**: + - Added **Global Track System** — introduce four global tracks: **Marker**, **Tempo**, **Key Signature**, and **Chord** (chord-symbol span regions). + - Added **Audio Chord Detection** feature — open piano roll window for an audio or a MIDI region and click "..." -> **Detect Chords** to automatically analyse the recording and populate the Chord Track using a zero-dependency FFT pipeline with configurable sensitivity, stability, and seventh-chord detection. + - Added **Tempo Detection with Auto-Align Beats** feature — open piano roll window for an audio region and click "..." -> **Detect Tempo** in the toolbar to analyse the audio for BPM and optionally realign the project's Tempo Track regions to match. + - Added **Demucs 4S** as a second local browser-embedded stem-separation model — the existing two-stem UVR-MDX-NET model is now joined by the four-stem `htdemucs_4s` model (~172 MB, vocals / drums / bass / others), both running entirely in-browser via ONNX Runtime WebGPU. + +- **2026.05.15**: Added **browser-embedded AI models** — two AI models now run entirely in the browser with no external service, no API key, and no K.G.One server required. The **K.G.Studio Musician Assistant** gains a new **Local LLM (Browser)** provider powered by **Gemma 4 E4B** via LiteRT-LM with WebGPU acceleration; the model is downloaded once and cached in OPFS for instant subsequent launches, with configurable context length (32 k / 64 k / 128 k tokens) and live inference performance statistics. **Stem separation** now also runs locally through a browser-embedded **UVR-MDX-NET-Inst_HQ_3** ONNX model with WebGPU acceleration — open the **Music Generator** panel (✦ button), download the model once, and separate vocals from instruments entirely on-device. Both features require a WebGPU-capable browser (Chrome 113+ or Edge 113+) and a secure context (HTTPS or localhost). Recommended hardware: a GPU with at least 8 GB VRAM or a system with at least 16 GB unified RAM. + +- **2026.05.10**: Added **staff notation (sheet music) view** — the piano roll now offers a full standard notation mode. Switch between Piano Roll and Sheet Music views using the toggle in the piano roll toolbar. In sheet music mode, notes are engraved via VexFlow with automatic clef selection (treble or bass) based on the active instrument, key signature rendering, automatic beam grouping, ties across bar lines, and configurable quantization for note-value resolution. Enable **Track Scope** to render all MIDI regions on the track as a continuous score rather than a single isolated region. + +- **2026.05.09**: Added **audio recording** — record directly from your microphone into an audio track. A live waveform preview grows in real time as you record, and the region is committed to the timeline as a standard audio region when you stop. Added **audio I/O device selection** in Settings so you can choose your preferred microphone input and audio output device. + +- **2026.05.08**: Added **MIDI automation** — draw and edit pitch bend and MIDI CC curves (CC1 Modulation, CC2 Breath, CC7 Volume, CC11 Expression, CC64 Sustain) in an editable automation lane below the piano grid. Added **track-level automation**: each track now has a dedicated automation panel where you can view and edit the same curves directly on the timeline. Real-time MIDI controller input (pitch wheel, CC pedals) is recorded and played back with per-lane interpolation. Added the **Event List Panel** — a tabbed sidebar (Notes / Pitch Bend / Controller) for inspecting and inline-editing all events in the active MIDI region. Added **region multi-select** with lasso and bulk move/resize, and **merge MIDI regions**. +
+ K.G.Studio Logo +
+ +- **2026.05.02**: Added **audio track spectrogram visualization** — audio regions now display a real-time spectrogram overlay in the track grid. Added **Piano Roll hybrid mode**: open the piano roll on a MIDI region while an adjacent audio region's spectrogram is shown as a reference layer, letting you edit MIDI notes against the visual shape of the audio. Added **piano roll zoom in/out** with viewport-position preservation so the view stays anchored to the current playhead. Added **fine-tune region position**: nudge regions by small increments for precise placement. Also added cross-component playhead scroll synchronization so the main grid and piano roll stay in sync during playback. + +- **2026.04.29**: Added **Remix** and **Repaint** to the K.G.One Music Generator panel (powered by ACE-Step 1.5). **Remix** lets you cover an existing audio region in a new style — select an audio region, describe the target style and optionally provide new lyrics, and ACE-Step will re-perform the song with the prompted instrumentation and feel. **Repaint** lets you surgically re-generate a specific section of a song — set a loop range on the timeline to define the repaint window, then describe what you want that section to sound like; the rest of the song stays untouched. Both tools support the same import workflow as the other K.G.One tabs: preview the result in the built-in player, drag it onto a track, or click **Import Aligned to Source** to automatically place it below the original region in a new track. + +- **2026.04.24**: Added [**K.G.One Music Studio**](https://github.com/KGAudioLab/K.G.One) integration! When K.G.Studio connects to a local K.G.One server, the **K.G.One Music Generator** panel (magic wand button ✦ in the toolbar) becomes available with three AI-powered tools: **Full Song Generation** (powered by ACE-Step 1.5 — generate full-length songs from text prompts), **Clip Generation** (powered by Foundation-1 — generate instrument clips and MIDI loops from text), and **Stem Separation** (powered by python-audio-separator — split any audio into vocals, instrumentals, and more). Generated audio and MIDI can be previewed instantly and dragged directly onto your tracks. K.G.One runs entirely on your own machine (Windows/Linux, CUDA GPU required); see the [K.G.One repository](https://github.com/KGAudioLab/K.G.One) for setup instructions. + +- **2026.04.11**: Migrated project storage from IndexedDB to OPFS (Origin Private File System) with a folder-based structure for better media file handling. Added audio track support with WAV/MP3 import, playback, looping, and non-destructive region trimming. Added bounce-to-WAV/MP3 export via offline rendering. + +- **2026.04.05**: Migrated the AI agent from XML-based tool calling to native OpenAI SDK function calling for improved reliability and compatibility. Added new LLM model options including GPT-5.4 series. + +- **2026.01.23**: Implemented seamless loop playback! Drag on the bar numbers to set loop range, or toggle loop mode with the Loop button in the toolbar. Loop playback uses `Tone.js`'s native looping for sample-accurate, gap-free looping. + +- **2025.12.21**: Implemented MIDI keyboard support! You can now connect a MIDI keyboard and use it to play sounds. Please note that this feature may not work optimally in Safari and some other browsers that lack complete Web MIDI interface support. + +- **2025.12.15**: Added Intelligent Chord Assistant with functional harmony guidance (T/S/D). Hover over piano keys to see context-aware chord suggestions and create full chords with one click! diff --git a/public/chat/help-fr_fr.md b/public/chat/help-fr_fr.md index 18452a8..222059c 100644 --- a/public/chat/help-fr_fr.md +++ b/public/chat/help-fr_fr.md @@ -45,7 +45,16 @@ OpenRouter donne accès à de nombreux modèles via une API unique, y compris ce 2. Dans **Réglages ⚙️ → Général → Fournisseur LLM**, choisissez **Serveur compatible OpenAI**. 3. Saisissez votre clé dans **Serveur compatible OpenAI → Clé**. 4. Consultez les modèles disponibles sur la [**page des modèles OpenRouter**](https://openrouter.ai/models). -5. Saisissez le nom du modèle dans **Serveur compatible OpenAI → Modèle**. +5. Saisissez le nom du modèle dans **Serveur compatible OpenAI → Modèle**. Les séries recommandées comprennent : + - `Anthropic: Claude Sonnet 4.6` (`anthropic/claude-sonnet-4.6` : [Link](https://openrouter.ai/anthropic/claude-sonnet-4.6)) — meilleur équilibre qualité/coût pour la série Claude + - Modèles gratuits : + - `OpenAI: GPT-OSS 120B` (MODÈLE GRATUIT : `openai/gpt-oss-120b:free` : [Link](https://openrouter.ai/openai/gpt-oss-120b:free)) + - `Google: Gemma 4 26B A4B IT` (MODÈLE GRATUIT : `google/gemma-4-26b-a4b-it:free` : [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it:free)) + - `Google: Gemma 4 31B IT` (MODÈLE GRATUIT : `google/gemma-4-31b-it:free` : [Link](https://openrouter.ai/google/gemma-4-31b-it:free)) + - Pour l'auto-déploiement (nécessite ~24 Go de VRAM ou 24-32 Go de mémoire unifiée), nous recommandons : + - `Qwen: Qwen3.6 35B A3B` (`qwen/qwen3.6-35b-a3b` : [Link](https://openrouter.ai/qwen/qwen3.6-35b-a3b)) + - `Google: Gemma 4 26B A4B IT` (`google/gemma-4-26b-a4b-it` : [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it)) + - `Google: Gemma 4 31B IT` (`google/gemma-4-31b-it` : [Link](https://openrouter.ai/google/gemma-4-31b-it)) 6. Saisissez `https://openrouter.ai/api/v1` dans **Serveur compatible OpenAI → URL de base**. ### Opérations DAW de base diff --git a/public/chat/help-zh_cn.md b/public/chat/help-zh_cn.md index cebab8e..2852454 100644 --- a/public/chat/help-zh_cn.md +++ b/public/chat/help-zh_cn.md @@ -48,9 +48,14 @@ OpenRouter 提供统一接口,可访问多个语言模型提供方的模型, **注意:** 各模型提供方的数据保留与隐私策略可能不同,使用前请自行查看。 5. 在 **OpenAI 兼容服务 → 模型** 中填入模型名。推荐系列包括: - `Anthropic: Claude Sonnet 4.6`(`anthropic/claude-sonnet-4.6`)—— Claude 系列里质量和成本平衡较好 - - `Qwen: Qwen3.5-35B-A3B`(`qwen/qwen3.5-35b-a3b`)—— 推荐开源模型 - - `Qwen: Qwen3-Next-80B-A3B`(免费:`qwen/qwen3-next-80b-a3b-instruct:free`)—— 推荐免费模型 - - `OpenAI: GPT-OSS 120B`(免费:`openai/gpt-oss-120b:free`)—— 推荐免费模型 + - 免费模型: + - `OpenAI: GPT-OSS 120B`(免费:`openai/gpt-oss-120b:free`) + - `Google: Gemma 4 26B A4B IT`(免费:`google/gemma-4-26b-a4b-it:free`) + - `Google: Gemma 4 31B IT`(免费:`google/gemma-4-31b-it:free`) + - 对于自托管/本地部署(需要约 24G 显存或 24-32GB 统一内存),我们推荐: + - `Qwen: Qwen3.6 35B A3B`(`qwen/qwen3.6-35b-a3b`) + - `Google: Gemma 4 26B A4B IT`(`google/gemma-4-26b-a4b-it`) + - `Google: Gemma 4 31B IT`(`google/gemma-4-31b-it`) - 注意:免费模型会经常变化,请以 OpenRouter 模型页中的 **Prompt Pricing** 过滤结果为准 - 注意:免费模型提供方可能会收集您的数据,使用前请先查看模型页面说明 6. 在 **OpenAI 兼容服务 → 基础 URL** 中填写 `https://openrouter.ai/api/v1`。 diff --git a/public/chat/help-zh_hk.md b/public/chat/help-zh_hk.md index 1cdbdf9..378287a 100644 --- a/public/chat/help-zh_hk.md +++ b/public/chat/help-zh_hk.md @@ -48,9 +48,14 @@ OpenRouter 提供統一接口,可訪問多個語言模型提供方的模型, **注意:** 各模型提供方的資料保留與隱私策略可能不同,使用前請自行查看。 5. 在 **OpenAI 兼容服務 → 模型** 中填入模型名。推薦系列包括: - `Anthropic: Claude Sonnet 4.6`(`anthropic/claude-sonnet-4.6`)—— Claude 系列裡質量和成本平衡較好 - - `Qwen: Qwen3.5-35B-A3B`(`qwen/qwen3.5-35b-a3b`)—— 推薦開源模型 - - `Qwen: Qwen3-Next-80B-A3B`(免費:`qwen/qwen3-next-80b-a3b-instruct:free`)—— 推薦免費模型 - - `OpenAI: GPT-OSS 120B`(免費:`openai/gpt-oss-120b:free`)—— 推薦免費模型 + - 免費模型: + - `OpenAI: GPT-OSS 120B`(免費:`openai/gpt-oss-120b:free`) + - `Google: Gemma 4 26B A4B IT`(免費:`google/gemma-4-26b-a4b-it:free`) + - `Google: Gemma 4 31B IT`(免費:`google/gemma-4-31b-it:free`) + - 對於自託管/本地部署(需要約 24G 顯存或 24-32GB 統一記憶體),我們推薦: + - `Qwen: Qwen3.6 35B A3B`(`qwen/qwen3.6-35b-a3b`) + - `Google: Gemma 4 26B A4B IT`(`google/gemma-4-26b-a4b-it`) + - `Google: Gemma 4 31B IT`(`google/gemma-4-31b-it`) - 注意:免費模型會經常變化,請以 OpenRouter 模型頁中的 **Prompt Pricing** 過濾結果為準 - 注意:免費模型提供方可能會收集您的資料,使用前請先查看模型頁面說明 6. 在 **OpenAI 兼容服務 → 基礎 URL** 中填寫 `https://openrouter.ai/api/v1`。 diff --git a/public/chat/help.md b/public/chat/help.md index 28943e9..62cd64f 100644 --- a/public/chat/help.md +++ b/public/chat/help.md @@ -48,9 +48,14 @@ OpenRouter is a platform that provides unified access to a wide range of languag **Note:** Each model provider may have different data retention and privacy policies. Please review these policies before use. 5. Enter your chosen model name in **OpenAI Compatible Server → Model**. Recommended model series include: - `Anthropic: Claude Sonnet 4.6` (`anthropic/claude-sonnet-4.6`: [Link](https://openrouter.ai/anthropic/claude-sonnet-4.6)) — best balance of quality and cost for the Claude series - - `Qwen: Qwen3.5-35B-A3B` (`qwen/qwen3.5-35b-a3b`: [Link](https://openrouter.ai/qwen/qwen3.5-35b-a3b)) — recommended open source model - - `Qwen: Qwen3-Next-80B-A3B` (FREE MODEL: `qwen/qwen3-next-80b-a3b-instruct:free`: [Link](https://openrouter.ai/qwen/qwen3-next-80b-a3b-instruct:free)) — recommended free model - - `OpenAI: GPT-OSS 120B` (FREE MODEL: `openai/gpt-oss-120b:free`: [Link](https://openrouter.ai/openai/gpt-oss-120b:free)) — recommended free model + - Free Models: + - `OpenAI: GPT-OSS 120B` (FREE MODEL: `openai/gpt-oss-120b:free`: [Link](https://openrouter.ai/openai/gpt-oss-120b:free)) + - `Google: Gemma 4 26B A4B IT` (FREE MODEL: `google/gemma-4-26b-a4b-it:free`: [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it:free)) + - `Google: Gemma 4 31B IT` (FREE MODEL: `google/gemma-4-31b-it:free`: [Link](https://openrouter.ai/google/gemma-4-31b-it:free)) + - For self-deployment (requiring ~24G VRAM or 24-32GB Unified Memory), we recommend: + - `Qwen: Qwen3.6 35B A3B` (`qwen/qwen3.6-35b-a3b`: [Link](https://openrouter.ai/qwen/qwen3.6-35b-a3b)) + - `Google: Gemma 4 26B A4B IT` (`google/gemma-4-26b-a4b-it`: [Link](https://openrouter.ai/google/gemma-4-26b-a4b-it)) + - `Google: Gemma 4 31B IT` (`google/gemma-4-31b-it`: [Link](https://openrouter.ai/google/gemma-4-31b-it)) - Note: free model availability changes frequently — for the latest free options, visit the [OpenRouter Models Page](https://openrouter.ai/models) and use the **Prompt Pricing** filter - Note: free model providers may collect your data; check the model page for details before use 6. Input the base URL `https://openrouter.ai/api/v1` in **OpenAI Compatible Server → Base URL**.