音乐局部重绘

编辑和组合现有歌曲的片段

使用 music_v2 或 music_v2_5 模型进行音乐局部重绘,可修改歌曲的特定部分,同时保持其余内容不变。存储生成的歌曲后,可在作曲计划中引用其中的片段,以保持原样、重新生成,或基于原始音频生成新音频。

工作原理

music_v2 或 music_v2_5 作曲计划是由 片段 组成的有序列表。每个片段属于以下两种类型之一:

  • 生成片段 — 根据 text 和风格生成新音频。可用于重新生成某个部分或添加新内容。
  • 音频引用片段 — 原样插入已存储歌曲的一段音频。可用于完全保留现有歌曲的某个部分。

快速入门

1

存储歌曲以进行局部重绘

有两种方式可存储用于局部重绘的歌曲:使用 store_for_inpainting 生成新歌曲,或上传现有音频文件。

import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(api_key=os.environ.get("ELEVENLABS_API_KEY"))
# Generate a song and store it for later inpainting
response = elevenlabs.music.compose_detailed(
prompt="An upbeat pop song with verse and chorus",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
# Save the audio
with open("original.mp3", "wb") as f:
f.write(response.audio)
2

保留和重新生成片段

创建一个混合音频引用片段(保留)与生成片段(重新生成)的计划,然后通过 model_id="music_v2_5" 将其传递给 compose:

# Keep the first 30 seconds, regenerate the rest with a new style
composition_plan = {
"chunks": [
# Keep the original first 30 seconds unchanged
{
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 30000}
},
# Regenerate the chorus with a new style
{
"text": "[Chorus]\nWe're rising up tonight\nNothing can stop us now",
"duration_ms": 30000,
"positive_styles": ["bigger drums", "layered vocals", "anthemic"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(
composition_plan=composition_plan,
model_id="music_v2_5",
)
with open("edited.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)

条件控制

生成片段可通过 conditioning_ref 基于已存储音频的一段内容进行条件控制。模型会重新生成该片段,同时尽量贴近引用音频的音乐特征。使用 condition_strength(low、medium、high 或 xhigh)控制片段遵循引用的紧密程度。

{
"text": "[Chorus]\nThis is my moment\nI won't let it go",
"duration_ms": 15000,
"positive_styles": ["powerful vocals", "full band", "anthemic"],
"negative_styles": [],
"context_adherence": "high",
"conditioning_ref": {
"song_id": "vVtPM1Sas70E2LIhQFch",
"range": { "start_ms": 30000, "end_ms": 45000 }
},
"condition_strength": "high"
}

第一个片段会影响所有后续片段的生成。如需让整首歌基于某个引用,请从第一个片段开始应用 conditioning_ref。

条件引用最长可达 30 秒(30,000ms)。

上下文遵循度

每个生成片段都有 context_adherence 级别,用于控制其遵循相邻片段的紧密程度:

  • high(默认)— 与周围片段保持一致。适用于保留音频和重新生成音频之间的平滑过渡。
  • medium — 平衡一致性与创作自由度。
  • low — 让片段偏离上下文,发挥更多创意。

示例

编辑单个部分

生成电影预告片,然后仅使用不同歌词重新生成片尾。

1

生成原始版本

composition_plan = {
"chunks": [
{
"text": "[Intro]\nIn a world beyond code\nWhere sound becomes life",
"duration_ms": 15000,
"positive_styles": ["cinematic", "epic", "orchestral", "low strings", "suspenseful"],
"negative_styles": ["acoustic", "pop", "minimalistic"],
"context_adherence": "high"
},
{
"text": "[Build]\nTechnology awakens the future\nShaping every word into power",
"duration_ms": 20000,
"positive_styles": ["rising brass", "full orchestra", "epic"],
"negative_styles": ["acoustic", "pop"],
"context_adherence": "high"
},
{
"text": "[Bridge]\n(ah ah ah ah)",
"duration_ms": 15000,
"positive_styles": ["ethereal choir", "crescendo"],
"negative_styles": [],
"context_adherence": "high"
},
{
"text": "[Outro]\nThe voice of tomorrow, unleashed\nElevenLabs",
"duration_ms": 10000,
"positive_styles": ["deep narration", "epic finale"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
response = elevenlabs.music.compose_detailed(
composition_plan=composition_plan,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

仅编辑片尾

使用一个音频引用片段保留前三个部分(第 0–50 秒),然后重新生成片尾:

edited_plan = {
"chunks": [
# Keep the intro, build, and bridge unchanged
{
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 50000}
},
# Regenerate the outro with new lyrics
{
"text": "[Outro]\nThe future has arrived\nElevenLabs",
"duration_ms": 10000,
"positive_styles": ["deep narration", "epic finale"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=edited_plan, model_id="music_v2_5")

延长歌曲

为现有歌曲添加新的前奏和片尾。

1

生成原始版本

response = elevenlabs.music.compose_detailed(
prompt="Berlin night club techno",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

延长并添加新的前奏和片尾

在两个新的生成片段之间插入原始歌曲的保留片段:

extend_plan = {
"chunks": [
# New intro
{
"text": "[Intro]",
"duration_ms": 30000,
"positive_styles": ["techno", "building tension", "filtered synths"],
"negative_styles": [],
"context_adherence": "high"
},
# Keep the core of the original (seconds 10-50)
{
"song_id": song_id,
"range": {"start_ms": 10000, "end_ms": 50000}
},
# New outro
{
"text": "[Outro]",
"duration_ms": 30000,
"positive_styles": ["techno", "fading out", "sparse"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=extend_plan, model_id="music_v2_5")

创建无缝循环

生成一段音乐乐句,并使用“衔接”片段连接两次重复的同一音频片段,从而创建循环。

1

生成短片段

composition_plan = {
"chunks": [
{
"text": "[Solo Acoustic]",
"duration_ms": 10000,
"positive_styles": ["acoustic guitar", "fingerpicking", "warm tone", "soft dynamics"],
"negative_styles": ["electric", "drums", "electronic"],
"context_adherence": "high"
}
]
}
response = elevenlabs.music.compose_detailed(
composition_plan=composition_plan,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

使用衔接片段创建循环

loop_plan = {
"chunks": [
# Loop start - keep a slice of the original
{
"song_id": song_id,
"range": {"start_ms": 3000, "end_ms": 8000}
},
# Glue - generate a smooth transition between the two slices
{
"text": "[Glue]",
"duration_ms": 3000,
"positive_styles": ["acoustic guitar", "fingerpicking", "smooth transition"],
"negative_styles": [],
"context_adherence": "high"
},
# Loop end - keep the same slice again
{
"song_id": song_id,
"range": {"start_ms": 3000, "end_ms": 8000}
}
]
}
audio = elevenlabs.music.compose(composition_plan=loop_plan, model_id="music_v2_5")

生成相似歌曲

基于现有歌曲的一小段音频生成全新的歌曲,以延续其音乐特征。由于第一个片段会影响之后的所有片段,即使只有该片段引用已存储音频,在第一个片段上应用 conditioning_ref 也会塑造整次生成。

1

生成原始版本

response = elevenlabs.music.compose_detailed(
prompt="An upbeat pop song with bright synths and driving drums",
music_length_ms=60000,
model_id="music_v2_5",
store_for_inpainting=True
)
song_id = response.song_id
2

基于原始版本条件控制生成新歌

similar_plan = {
"chunks": [
# The first chunk conditions the whole song on the reference
{
"text": "[Verse]\nSalt on my skin from a borrowed sea\nCounting the heartbeats it takes to break free",
"duration_ms": 30000,
"positive_styles": ["pop", "energetic", "bright synths", "driving drums"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high",
"conditioning_ref": {
"song_id": song_id,
"range": {"start_ms": 0, "end_ms": 10000}
},
"condition_strength": "high"
},
{
"text": "[Chorus]\nWe're rising up tonight\nNothing can stop us now",
"duration_ms": 30000,
"positive_styles": ["bigger drums", "layered vocals", "anthemic"],
"negative_styles": ["sparse", "minimal"],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(composition_plan=similar_plan, model_id="music_v2_5")

分块参考

一个 music_v2 计划最多包含 30 个分块。每个分块可以是生成分块或音频参考分块。

生成分块

字段类型描述
textstring方括号中的段落名称([Verse 1])、歌词行,以及花括号中的内联指示({scratching})。
duration_msnumber时长,单位为毫秒(3,000 - 120,000)。
positive_stylesarray要包含的风格和指示(最多 50 个)。
negative_stylesarray要避免的风格和指示(最多 50 个)。默认为空。
context_adherencestringlow、medium 或 high(默认)。分块对其周围分块的遵循程度。
conditioning_refobject | null可选的存储音频 { song_id, range } 片段,用作条件参考。默认为 null。
condition_strengthstring | nulllow、medium(默认)、high 或 xhigh。分块对条件参考的遵循程度。

第一个分块的风格最重要,它决定整体基调和流派。在方向确定前, 建议早期分块至少使用 6 - 7 种风格。

音频参考分块

字段类型描述
song_idstring要从中获取音频的已存储歌曲 ID。
rangeobject要原样插入的已存储歌曲 { start_ms, end_ms } 片段。

限制

限制条件值
每个计划的最大分块数30
最短分块时长3 秒(3,000ms)
最长分块时长2 分钟(120,000ms)
最大条件参考时长30 秒(30,000ms)
最小时间范围50ms

后续步骤