作曲计划

通过结构化 JSON 精确控制音乐生成

作曲计划可精细控制音乐生成。music_v2 或 music_v2_5 计划是由多个片段组成的有序列表,每个片段定义歌曲的一个部分,并拥有各自的风格、歌词和时长。快速原型设计可使用文本提示词;需要特定片段结构、精确歌词时序或复杂编曲时,请使用作曲计划。

作曲计划和文本提示词不能同时使用。二选一即可。

{
"chunks": [
{
"text": "[Verse 1]\nWoke up today with a feeling inside\nSomething is changing I cannot hide\nThe sun on my face and the wind at my back\nI'm finally ready to get on track",
"duration_ms": 16000,
"positive_styles": [
"upbeat pop",
"female vocalist with clear tone",
"acoustic guitar and light synths",
"gentle and conversational vocals",
"light drums in background",
"polished production",
"120 BPM",
"C major"
],
"negative_styles": ["dark", "aggressive", "slow tempo", "a cappella"],
"context_adherence": "high"
},
{
"text": "[Verse 2]\nUsed to be scared of the world outside\nBuilding up walls where I used to hide\nBut now I see clearly what I need to do\nTake that first step into something new",
"duration_ms": 16000,
"positive_styles": [
"confident vocals",
"fuller guitar strumming",
"steady drum beat",
"bass joins in"
],
"negative_styles": ["a cappella", "sparse", "quiet"],
"context_adherence": "high"
},
{
"text": "[Pre-Chorus]\nNo more waiting for tomorrow\nThis is my time now",
"duration_ms": 8000,
"positive_styles": [
"building intensity",
"rising synth melody",
"driving drums",
"full band playing"
],
"negative_styles": ["a cappella", "dropping out"],
"context_adherence": "high"
},
{
"text": "[Chorus]\nI'm breaking through\nNothing's gonna stop me now\nI'm breaking through\nFinally found out how",
"duration_ms": 16000,
"positive_styles": [
"powerful and anthemic vocals",
"full band at maximum energy",
"punchy drums and bass",
"layered synths and guitar"
],
"negative_styles": ["a cappella", "minimal", "stripped back"],
"context_adherence": "high"
},
{
"text": "[Outro]",
"duration_ms": 8000,
"positive_styles": [
"instrumental fade out",
"guitar melody repeating",
"drums softening",
"gentle ending"
],
"negative_styles": ["vocals", "abrupt ending", "building"],
"context_adherence": "high"
}
]
}

基于片段的作曲计划需要 music_v2 或 music_v2_5。作曲时传入 model_id="music_v2_5"。

快速入门

按照音乐快速入门设置 API 密钥并安装 SDK,然后使用作曲计划获得更精细的控制。

1

使用作曲计划生成音乐

import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(api_key=os.environ.get("ELEVENLABS_API_KEY"))
composition_plan = {
"chunks": [
{
"text": "[Verse]\nWalking down an empty street\nWondering who I'll meet",
"duration_ms": 15000,
"positive_styles": ["pop", "upbeat", "female vocals", "soft vocals", "acoustic guitar"],
"negative_styles": ["dark", "slow"],
"context_adherence": "high"
},
{
"text": "[Chorus]\nThis is my moment\nI won't let it go",
"duration_ms": 15000,
"positive_styles": ["powerful vocals", "full band"],
"negative_styles": [],
"context_adherence": "high"
}
]
}
audio = elevenlabs.music.compose(
composition_plan=composition_plan,
model_id="music_v2_5",
# with_timestamps=True, # Optional: return word-level timestamps
)
with open("output.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)
2

根据提示词生成计划

根据文本描述生成作曲计划,再在生成前修改:

plan = elevenlabs.music.composition_plan.create(
prompt="An upbeat pop song about summer adventures",
music_length_ms=60000,
model_id="music_v2_5"
)
# Modify the generated plan
plan["chunks"][0]["text"] = "[Verse 1]\nCustom lyrics here"
audio = elevenlabs.music.compose(composition_plan=plan, model_id="music_v2_5")

结构参考

片段

作曲计划是最多包含 30 个片段的有序列表。每个片段会根据其 text 和风格生成歌曲的一个部分。第一个片段最为重要:它的风格决定整首歌的整体氛围和流派。

字段类型说明
textstring方括号内的段落名称([Verse 1])、歌词行,以及花括号内的行内指示({scratching})。
duration_msnumber时长,单位为毫秒(3,000 - 120,000)。
positive_stylesarray要包含的风格和指示(最多 50 项)。
negative_stylesarray要避免的风格和指示(最多 50 项)。默认为空。
context_adherencestringlow、medium 或 high(默认)。表示片段遵循相邻片段的程度。

一首歌最多可有 30 个片段。总时长必须在 3 秒至 10 分钟之间,每个片段则须在 3 至 120 秒之间。

第一个片段的风格最重要,因为它决定整体氛围和流派。在方向确定前,前几个片段最好至少使用 6 - 7 种风格。可在列表中附加“great production quality”等通用风格作为默认项。

片段还可以引用已存储歌曲中的音频,以保持现有部分不变,或基于这些音频生成新音频。编辑和组合现有歌曲,请参阅音乐局部重绘。

编写歌词

text 字段结合了段落名称、歌词和行内指示:

  • 方括号内的 段落名称:[Verse 1]、[Chorus]、[Bridge]
  • 以纯文本形式编写的 歌词,每行以换行符分隔(\n)
  • 括号内的 语音化声音:(hmmm hmmm)、(ooh)、(yeah)
  • 花括号内的 行内指示:{guitar solo}、{scratching}、{instrumental break}

简短的行内提示请使用花括号。对于应用于整个片段的更广泛特征——流派、配器或整体演唱风格——则使用 positive_styles。

{
"text": "[Verse]\n(soft female vocals) I've been waiting\n(instrumental break)\nfor you"
}

在修正后的示例中,整体演唱风格移至 positive_styles,而简短的行内提示仍保留在 text 中,并使用花括号而非圆括号。

风格建议

请使用具体的风格描述:

{
"positive_styles": [
"warm acoustic guitar with light fingerpicking",
"soft female vocals with intimate delivery",
"gentle percussion with brushed snare",
"80 BPM"
]
}

积极使用负向风格来避免不需要的声音。风格必须使用英文(歌词可以使用任何语言)。

如果在风格中包含受版权保护的内容,API 会返回 bad_composition_plan 错误,并提供建议的替代方案。参阅处理受版权保护的素材。

示例

电影感纯音乐

{
"chunks": [
{
"text": "[Tension Build]",
"duration_ms": 15000,
"positive_styles": [
"cinematic",
"orchestral",
"epic",
"low strings tremolo",
"building intensity",
"80 BPM",
"D minor"
],
"negative_styles": ["vocals", "lyrics", "pop", "electronic", "bright"],
"context_adherence": "high"
},
{
"text": "[Climax]",
"duration_ms": 15000,
"positive_styles": ["full orchestra", "brass fanfare", "triumphant"],
"negative_styles": ["quiet", "vocals"],
"context_adherence": "high"
},
{
"text": "[Resolution]",
"duration_ms": 10000,
"positive_styles": ["gentle strings", "piano melody", "fading out"],
"negative_styles": ["intense", "vocals"],
"context_adherence": "high"
}
]
}

带旁白配音的广告

{
"chunks": [
{
"text": "[Intro]",
"duration_ms": 5000,
"positive_styles": [
"upbeat",
"modern pop",
"energetic",
"120 BPM",
"instrumental",
"catchy hook"
],
"negative_styles": ["sad", "slow", "dark", "vocals"],
"context_adherence": "high"
},
{
"text": "[Voiceover]\nIntroducing the future of productivity\nWork smarter, not harder",
"duration_ms": 10000,
"positive_styles": ["spoken voiceover", "confident male voice", "background music"],
"negative_styles": ["singing"],
"context_adherence": "high"
},
{
"text": "[Outro]",
"duration_ms": 5000,
"positive_styles": ["musical sting", "memorable"],
"negative_styles": ["vocals"],
"context_adherence": "high"
}
]
}

后续步骤