H3 提示词写作技巧

发布于 2026-09-10

ref-en.txt

MiniMax H3 官方提供,适用于 Full-Reference 全参考模式任务的规范 ref-en.txt,包含中文对照翻译,调整了部分格式,便于阅读和理解,Agent 实际使用时以原文为准。

ref_en.txt

Full-Reference Mode Rewrite Output Format Guide

全参考模式改写输出格式指南

This guide explains how rewrite outputs are organized and written in full-reference mode.

本指南说明在全参考模式(Full-Reference Mode)下,改写输出应如何组织和编写。

Write all six rewrite sections in English.

所有六个改写部分均使用英文编写。

Preserve the original language only for dialogue and lyrics inside <d> and for text visibly present in the scene.

只有 <d> 标签中的对白和歌词,以及场景中实际可见的文字,保留其原始语言。

Description detail: Make detailed_description as detailed and explicit as possible.

描述详细程度: detailed_description 应尽可能详细、明确。

For each shot, clearly establish the current composition, subject appearance and position, environment and lighting, actions and state changes, camera movement, current sound, and the points where referenced content actually appears or takes effect.

对于每个镜头,都应清楚说明当前构图、主体的外观与位置、环境与光照、动作与状态变化、镜头运动、当前声音,以及参考内容实际出现或产生作用的具体位置。

Avoid reducing the description to a plot summary or a list of reference relationships.

避免将描述简化为剧情概述,或仅仅罗列参考素材之间的对应关系。

The basic formats for shots, camera movement, speakers, dialogue, and ordinary sound are shared with the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA). This guide focuses on the reference labels, analysis sections, and format differences specific to full-reference mode.

镜头、镜头运动、说话者、对白以及普通声音的基本格式,与《Video Prompt Writing Guide(T2VA / I2VA / FL2VA/L2VA)》保持一致。本指南重点说明全参考模式特有的参考标签、分析部分以及格式差异。

1. Overall Structure

整体结构

A complete rewrite output consists of six sections in the following order:

一个完整的改写输出由以下六个部分组成,并严格按照以下顺序排列:
Section 部分Purpose 用途
subject_definitionsDefines referenced content and its reference labels 定义参考内容及其对应的参考标签
summarySummarizes the task type, target video, and main reference relationships 概括任务类型、目标视频以及主要的参考关系
retention_analysisDescribes how referenced content is preserved, transferred, or reused描述参考内容如何被保留、迁移或复用
detailed_descriptionDescribes visuals, actions, shots, sound, and dialogue in playback order按视频播放顺序描述画面、动作、镜头、声音和对白
overall_soundscapeSummarizes ambience and physical sounds概括环境声和物理动作产生的声音
non_diegetic_musicDescribes background music audible only to the audience描述只有观众能够听到的非画面内背景音乐

2. Reference Labels and Definitions (subject_definitions)

参考标签与定义

Full-reference rewrites use four types of labels to identify the source and role of referenced content:

全参考模式的改写使用以下四种标签,用于标识参考内容的来源及其作用:
Label 标签Meaning 含义
<Subject N>Visible content abstracted from reference assets that can be reused or modified in the target video 从参考素材中抽象出的可见内容,可以在目标视频中复用或修改
<Picture N>A reference image used as a concrete target frame or shot-planning anchor作为具体目标帧或镜头规划锚点使用的参考图像
<Video N>A reference video that provides an editing source, continuation starting point, or whole-video temporal structure提供剪辑素材、续写起点或完整视频时间结构的参考视频
<Audio N>An audio signal that is copied or referenced 被复制或引用的音频信号

Once a reference label is assigned to a piece of content, it keeps the same meaning across subject_definitions, summary, retention_analysis, detailed_description, and the audio sections.

一旦为某项参考内容分配了参考标签,该标签在subject_definitions、summary、retention_analysis、detailed_description以及音频相关部分中,都必须始终保持相同的含义。

subject_definitions defines each piece of referenced content that must be tracked separately later, such as a person, an environment, a source video's structure, or an audio track.

subject_definitions 用于定义后续需要单独追踪的每一项参考内容,例如人物、环境、源视频的结构或音轨。

Give each item its own line and explain what its label denotes, its reference role, and the main features to follow; name the corresponding source asset when its provenance needs to be made explicit.

每一项参考内容应单独占一行,并说明:

  • 该标签具体表示什么;
  • 它在参考关系中承担什么作用;
  • 后续需要遵循的主要特征;
  • 当有必要明确其来源时,应注明对应的源素材。

If <Picture N> or <Video N> only identifies the source of another referenced item and will not be analyzed or used separately later, cite it inside that item's definition without adding a separate line.

如果 <Picture N> 或 <Video N> 仅用于标识另一项参考内容的来源,并且后续不会被单独分析或使用,则直接在对应参考项的定义中引用即可,无需再为其单独增加一行。

retention_analysis records where each referenced item appears and whether it is fully preserved, partially preserved, transferred, or reused.

retention_analysis 用于记录每项参考内容在目标视频中的出现位置,以及它是被完整保留、部分保留、迁移还是复用。

2.1 <Subject N>

<Subject N> is used for reusable visible content, including:

<Subject N> 用于表示可复用的可见内容,包括:
  • People, animals, or objects

    人物、动物或物体
  • Scenes, backgrounds, or environments

    场景、背景或环境
  • Clothing, props, interfaces, or visual effects

    服装、道具、界面或视觉效果
  • Styles, actions, expressions, or poses

    风格、动作、表情或姿态

It represents a content unit that will actually be used in the target video, rather than the source file itself.

它表示的是实际会在目标视频中使用的内容单元,而不是参考源文件本身。

One subject may be defined by multiple reference assets, and one reference asset may provide multiple subjects.

一个 <Subject N> 可以由多个参考素材共同定义,而一个参考素材也可以提供多个 <Subject N>。
PROMPT 片段

<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.

<Subject 1> 是 <Picture 1> 中的年轻女子,她留着深色长发,身穿蓝色开衫,并佩戴一条细银项链。

When the same subject comes from multiple assets, combine the sources and state what each asset provides:

当同一个主体的信息来自多个参考素材时,应将这些来源组合起来,并明确说明每个素材分别提供了什么:
PROMPT 片段

<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.

<Subject 1> 是该女子,她的外观来自 <Picture 1>,而她的行走动作来自 <Video 1>。

2.2 <Picture N>

Use a standalone <Picture N> when the reference image itself serves as a shot's first frame, keyframe, last frame, edited keyframe, or composition anchor:

当参考图像本身作为某个镜头的首帧、关键帧、尾帧、编辑后的关键帧或构图锚点时,应单独定义一个 <Picture N>:
PROMPT 片段

<Picture 2> is the first frame of [Shot 1], showing a woman seated beside a café window.

<Picture 2> 是 [Shot 1] 的首帧,画面中一名女子坐在咖啡馆窗边。

If an image is used only to define a character, scene, costume, or style, do not create a standalone picture entry.

如果一张图像仅用于定义角色、场景、服装或风格,则不需要单独创建 <Picture N> 条目。

Instead, cite the image source inside the corresponding <Subject N> definition.

应直接在对应的 <Subject N> 定义中引用该图像作为来源。

When an image acts as a storyboard or shot-planning reference, state which shots it maps to and what planning information it provides:

当一张图像作为故事板(storyboard)或镜头规划参考时,应明确说明它对应哪些镜头,以及它提供了哪些镜头规划信息:
PROMPT 片段

<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.

<Picture 3> 是 [Shot 1] 和 [Shot 2] 的故事板参考,用于确定这些镜头的视角、主体位置以及镜头顺序。

2.3 <Video N>

<Video N> is reserved for whole-video relationships, such as:

<Video N> 专门用于表示完整视频层面的参考关系,例如:
  • Editing an original video

    对原始视频进行编辑
  • Continuing from the end of an original video

    从原始视频的结尾继续生成
  • Referencing the original video's camera movement, cuts, rhythm, or temporal structure

    参考原始视频的镜头运动、镜头切换、节奏或时间结构
PROMPT 片段

<Video 1> is the source video for the target video edit.

<Video 1> 是目标视频编辑所使用的源视频。

If a person, object, scene, action, or effect from a reference video is reused as visible content, it still belongs under <Subject N>.

如果复用参考视频中的人物、物体、场景、动作或视觉效果,并将其作为目标视频中实际可见的内容,那么这些内容仍应定义为 <Subject N>。

<Video N> identifies the asset or structural source and does not replace subject labels.

<Video N> 用于标识参考视频素材本身或其结构来源,不能替代 <Subject N> 主体标签。

2.4 <Audio N>

<Audio N> represents a standalone audio asset or an enabled synchronized audio track from a reference video. Common uses include:

<Audio N> 表示独立的音频素材,或参考视频中启用的同步音轨。常见用途包括:
  • Copying all or part of an audio signal

    复制全部或部分音频信号
  • Referencing a background-music style

    参考背景音乐的风格
  • Referencing a speaker's voice timbre and delivery

    参考说话者的声音音色和表达方式
  • Using dialogue, lyrics, or sound effects from the original audio

    使用原始音频中的对白、歌词或音效
  • Referencing beat, rhythm, or audio continuity

    参考节拍、节奏或音频连续性

When an <Audio N> explicitly corresponds to a target speaker, reuse that speaker's global ID in the definition: write <Subject N> (Sx) when the speaker maps to a defined subject, or use a stable voice description followed by (Sx) otherwise.

当某个 <Audio N> 明确对应目标视频中的某位说话者时,应在定义中复用该说话者的全局 ID:如果说话者对应已经定义的主体,则写作 <Subject N> (Sx);否则,应先使用稳定的声音特征描述,再在后面标注 (Sx)。

The ID comes from the target video's global speaker order and is not independently assigned or renumbered in the audio definition.

这里的 ID 来自目标视频中说话者的全局出现顺序,不能在音频定义中单独分配或重新编号。

See Section 5.4 for the speaker-numbering rules:

说话者编号规则参见第 5.4 节:
PROMPT 片段

<Audio 1> is the voice-timbre reference for <Subject 1> (S1).

<Audio 1> 是 <Subject 1> (S1) 的声音音色参考。

When one audio asset serves multiple roles, describe those roles in one natural sentence rather than creating additional subsections.

当同一份音频素材同时承担多种作用时,应使用一句自然、完整的句子描述这些作用,而不是为不同作用额外创建新的子部分。

2.5 Visual and Audio Tracks from the Same Reference Video

来自同一参考视频的视觉轨道与音频轨道

<Video N> and <Audio N> are numbered independently.

<Video N> 和 <Audio N> 分别独立编号。

Each index indicates only the label's order within its own category and does not encode a pairing between the two categories.

每个编号只表示该标签在自身类别中的顺序,并不表示两个类别之间存在配对关系。

The same reference video may therefore correspond to <Video 1> and <Audio 2>;

因此,同一个参考视频完全可以同时对应 <Video 1> 和 <Audio 2>。

different indices do not prevent them from coming from the same source asset.

编号不同,并不意味着它们不能来自同一个源素材。

An ordinary reference video does not create <Audio N> merely because the file contains sound.

普通的参考视频不会仅仅因为文件本身包含声音,就自动创建对应的 <Audio N>。

An <Audio N> definition primarily states the audio's role and does not have to name the <Video N> it comes from.

<Audio N> 的定义主要用于说明音频所承担的作用,并不要求必须注明它来自哪个 <Video N>。

State the shared source only when needed to remove provenance ambiguity, for example:

只有在需要消除素材来源歧义时,才需要明确说明二者来自同一个源素材。例如:
PROMPT 片段

<Video 1> is the source video for the target video edit.

<Video 1> 是目标视频编辑所使用的源视频。

<Audio 2> is the synchronized audio track of <Video 1> and is reused in the target video.

<Audio 2> 是 <Video 1> 的同步音轨,并在目标视频中被复用。

3. summary

This section uses one short English paragraph to summarize the target video and its reference relationships.

这一部分使用一个简短的英文段落,概括目标视频及其与参考素材之间的关系。

It begins with a square-bracketed task-type prefix:

内容以一个放在方括号中的任务类型前缀开始:
PROMPT 片段

[reference generation] ...

[引用生成]...

[video editing + reference generation + audio reuse] ...

[视频编辑 + 引用生成 + 音频复用]...

Choose task types according to the actual role each reference asset plays in the target video:

应根据每项参考素材在目标视频中实际承担的作用选择任务类型:
Task type 任务类型When to use it 使用场景
keyframe completionAn image serves as the target video's first frame, keyframe, last frame, edited keyframe, or another concrete frame anchor 图像作为目标视频的首帧、关键帧、尾帧、编辑后的关键帧或其他具体帧锚点
reference generationAn image, video, or audio asset provides generation guidance for a character, scene, style, action, camera movement, storyboard, and so on, without serving as a concrete frame or as the source video being edited or continued 图像、视频或音频素材为角色、场景、风格、动作、镜头运动、故事板等提供生成参考,但它既不作为具体帧,也不是被直接编辑或续写的源视频
video editingAn existing source video is directly modified; editing an image or generating between still keyframes does not belong to this type 直接修改已有的源视频;编辑图像或在静态关键帧之间生成视频不属于此类型
video continuationNew content continues, extends, resumes, or transitions from an existing source video 从已有源视频继续、延长、恢复生成或过渡到新的内容
audio reuseThe same audio signal is reused in full or in part 完整或部分复用相同的音频信号
audio referenceThe audio signal is not copied directly; only its music style, timbre, dialogue or lyric content, sound-effect texture, beat, or continuity is referenced 不直接复制音频信号,只参考其音乐风格、音色、对白或歌词内容、音效质感、节拍或连续性

When a task satisfies multiple relationships, combine the task types with + and do not repeat a type.

当一个任务同时满足多种关系时,使用 + 组合对应的任务类型,并且不要重复同一种类型。

For example, continuing from a source video while using an image as the last frame is written as [video continuation + keyframe completion].

例如,从源视频继续生成,同时使用一张图片作为目标视频的最后一帧,可以写为:[video continuation + keyframe completion]。

Editing a source video while retaining its original audio may be written as [video editing + audio reuse].

对源视频进行编辑,同时保留其原始音频,可以写为:[video editing + audio reuse]。

The mere presence of video or audio does not automatically create a corresponding task type.

仅仅存在视频或音频,并不会自动产生对应的任务类型。

If a reference video provides only camera movement, cuts, or rhythm, it normally belongs to reference generation.

如果参考视频仅提供镜头运动、剪辑或节奏方面的参考,通常应归入 reference generation。

Use video editing or video continuation only when that video is directly edited or continued.

只有当该视频被直接编辑或续写时,才使用 video editing 或 video continuation。

When editing a source video, use audio reuse as well if its original audio remains audible.

编辑源视频时,如果其原始音频在目标视频中仍然可听,则还应使用 audio reuse。

When continuing a source video without directly copying the audio signal, use audio reference if the new audio only continues the original track's audible characteristics.

续写源视频时,如果没有直接复制原始音频信号,而只是让新生成的音频延续原音轨的可听特征,则应使用 audio reference。

The summary uses the previously defined <Subject N>, <Picture N>, <Video N>, and <Audio N> labels to describe the main subjects, shot flow, and roles of the reference assets.

summary 使用前面已经定义好的 <Subject N>、<Picture N>、<Video N> 和 <Audio N> 标签,描述主要主体、镜头流程以及各项参考素材所承担的作用。

Do not introduce new reference labels in this section.

不要在这一部分引入新的参考标签。

For video-editing tasks, begin the summary after the task-type prefix with:

对于视频编辑(video-editing)任务,在任务类型前缀之后,应以下面的句子作为 summary 的开头:
PROMPT 片段

The target video is an edited version of <Video 1>.

目标视频是 <Video 1> 的编辑版本。

4. retention_analysis

保留分析

This section describes how each piece of referenced content is preserved, transferred, copied, or referenced in the target video.

本节描述每项参考内容在目标视频中如何被保留、迁移、复制或引用。

Use one line for each reference label and preserve the meaning established in subject_definitions.

每个参考标签使用一行,并保持与 subject_definitions 主体定义 中确立的含义一致。

4.1 Visible Content

可见内容

<Subject N>, <Picture N>, and <Video N> use the following relationship markers.

<Subject N>(主体 N)、<Picture N>(图片 N)和 <Video N>(视频 N)使用以下关系标记。

These markers are fixed English values in the output format:

这些标记是输出格式中固定使用的英文值:
Relationship marker 关系标记Meaning 含义
fully_preserved 完全保留The defined role of the referenced content is fully preserved参考内容所定义的作用被完整保留
partially_preserved部分保留)The referenced content is still used, but some defined characteristics are changed or only partially retained 参考内容仍被使用,但其部分已定义特征发生变化,或仅保留其中一部分
attribute_transfer属性迁移Referenced characteristics are transferred to a different identifiable target subject 参考内容的特征被迁移到另一个可明确识别的目标主体上
weak_reference弱参考Only broad similarity in style, category, composition, or atmosphere is retained 仅保留风格、类别、构图或氛围等方面的大致相似性

Subject entry:

主体条目:
PROMPT 片段

<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...

<Subject 1>(在 [Shot 1]、[Shot 3] 中出现):fully_preserved - ...

Picture entry:

图片条目
PROMPT 片段

<Picture 2> ([Shot 1] first frame): fully_preserved - ...

<Picture 2>(在 [Shot 1] 第一个镜头中出现):fully_preserved - ...

Video-structure entry:

视频结构条目
PROMPT 片段

<Video 1> (cut and pacing structure): weak_reference - ...

<Video 1>(剪辑和节奏结构):weak_reference - ...

4.2 Audio

音频

<Audio N> uses the following relationship markers:

<Audio N> 使用以下关系标记:
Relationship marker 关系标记Meaning 含义
fully_copyThe complete source audio serves as the target video's complete final audio track 完整的源音频直接作为目标视频最终的完整音轨
partially_copyOnly part of the timeline or selected audio layers are copied, or other sounds are added, removed, or replaced after copying 仅复制部分时间段或选定的音频层,或者复制后又添加、删除或替换了其他声音
referenceThe signal is not copied directly; only timbre, rhythm, music style, dialogue content, or sound texture is referenced 不直接复制原始音频信号,只参考其音色、节奏、音乐风格、对白内容或声音质感
weak_referenceOnly broad similarity in category or atmosphere is retained 仅保留类别或氛围层面的宽泛相似性
PROMPT 片段

<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.

<Audio 1> 以 1:1 的方式完整复用,作为目标视频最终的完整音轨。

PROMPT 片段

<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.

目标说话者参考 <Audio 2> 的声音音色和从容平稳的表达方式,但不复制原始音频信号

Choose each relationship marker only within the reference role already defined for that label in subject_definitions.

每个关系标记的选择,都必须限定在 subject_definitions 中已经为该标签定义的参考作用范围内。

Do not treat newly added actions, backgrounds, or plot events in the target video as losses of reference fidelity.

不要因为目标视频中新增了动作、背景或剧情事件,就将这些新增内容视为参考保真度的损失。

5. detailed_description

详细描述

This is the main body of a full-reference rewrite.

这是全参考模式改写的主体部分。

It describes visuals, actions, sound, and dialogue shot by shot in target-video playback order and inserts reference labels where they apply.

它按照目标视频的实际播放顺序,逐镜头描述画面、动作、声音和对白,并在参考内容实际产生作用的位置插入相应的参考标签。

5.1 Basic Format

基本格式

The basic format follows the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA):

基本格式遵循《Video Prompt Writing Guide(T2VA / I2VA / FL2VA / L2VA)》中的规则:
  • Write the body in English. Preserve the original language of dialogue, lyrics, and visible text.

    主体内容使用英文编写。对白、歌词以及画面中实际可见的文字保留其原始语言。
  • [Shot 1] marks the opening shot and has no timestamp. Later shots use [Shot N] At MM:SS.mmm, ... to mark cut times.

    [Shot 1] 表示开场镜头,不添加时间标记。后续镜头使用 [Shot N] At MM:SS.mmm, ... 标记切镜时间。
  • Write camera movement as natural English within the current shot, including movement type, amplitude, and speed when they need to be expressed.

    镜头运动应以自然的英文描述融入当前镜头中;需要明确表达时,应包含运动类型、幅度和速度。
  • Give vocal sources stable (S1), (S2), and subsequent IDs. Write dialogue and lyrics as <d>[Language] ...</d>.

    为所有人声来源分配稳定的 (S1)、(S2) 以及后续编号。对白和歌词使用 <d>[Language] ...</d> 格式。
  • Use <scenetrans>, <cutoff>, and the corresponding continuity descriptions for dialogue crossing a cut, speech truncated by the video ending, and continuous audio across shots.

    对于跨镜头延续的对白、因视频结束而被截断的语音,以及跨镜头连续播放的音频,应使用 <scenetrans>、<cutoff> 以及相应的连续性描述。

For complete rules and examples covering camera vocabulary, group speech, voice-over, dialogue across cuts, and visible text, see the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA).

关于镜头运动词汇、多人共同发声、画外音、跨镜头对白以及画面可见文字的完整规则和示例,请参阅《Video Prompt Writing Guide(T2VA / I2VA / FL2VA / L2VA)》。

5.2 Full-Reference Mode Differences

全参考模式的差异
Dimension 维度T2VAFull-reference mode 全参考模式
Main field 主字段integrated_multimodal_descriptiondetailed_description
Style opening 风格描述位置Written after [Shot 1]写在 [Shot 1] 之后Established in one or two English sentences before [Shot 1]在 [Shot 1] 之前,先用一到两句英文确定整体风格
Reference information 参考信息Does not use full-reference labels 不使用全参考模式的参考标签Inserts <Subject N>, <Picture N>, <Video N>, and <Audio N> at their first appearance and where their roles apply 在参考内容首次出现以及其参考作用实际生效的位置,插入 <Subject N>、<Picture N>、<Video N> 和 <Audio N>
Audio relationships 音频关系Describes the target video's own sound 描述目标视频自身的声音Cites <Audio N> in the corresponding shot or audio phase and states whether the signal is copied or referenced 在对应镜头或音频阶段引用 <Audio N>,并说明音频信号是被复制还是仅作为参考

Opening example:

开头示例:
PROMPT 片段

The target video is in a cinematic, literary music-video style with soft lighting and a slightly desaturated color palette.

目标视频的风格为电影文学风格的音乐视频,采用柔和的灯光和略微偏灰的色彩调色板。

[Shot 1] The scene opens in a crowded urban street...

[Shot 1] 场景在拥挤的都市街道上展开...

[Shot 2] At 00:09.000, the shot cuts to an extreme close-up...

[Shot 2] 在 00:09.000 时,镜头切到极端特写...

For generation tasks, detailed_description is normally 350-500 English words.

对于生成类任务,detailed_description 通常为 350–500 个英文单词。

Dialogue-dense content prioritizes fitting the complete spoken timeline rather than mechanically reaching a word count.

如果内容包含大量对白,应优先确保完整的对白时间线能够合理容纳在视频时长内,而不是机械地达到规定字数。

Video-editing descriptions scale with the complexity of the source video and do not have to follow the generation-task range.

对于视频编辑任务,描述长度应根据源视频本身的复杂程度进行调整,不要求遵循生成类任务的 350–500 词范围。

A single shot does not automatically justify a shorter description; distribute detail across multiple shots according to their information load.

只有一个镜头,并不意味着描述就应该自动缩短。应根据各个镜头承载的信息量合理分配描述细节。

5.3 Using Reference Labels in Shots

在镜头中使用参考标签

At the first clear appearance of an important <Subject N>, describe its referenced characteristics, position in the frame, and current action within what is actually visible in the shot.

当某个重要的 <Subject N> 首次清晰出现在画面中时,应结合该镜头实际可见的内容,描述其参考特征、画面中的位置以及当前动作。

Continue using the same label in later shots without redefining what the label represents.

在后续镜头中继续使用相同的标签即可,不需要重新定义该标签所代表的内容。

Use natural phrasing for concrete frame anchors:

对于具体的帧锚点,应使用自然的表达方式,例如:
PROMPT 片段

the shot begins from <Picture 1>

镜头从 <Picture 1> 所对应的画面开始;

the shot's keyframe corresponds to <Picture 2>

该镜头的关键帧对应 <Picture 2>;

the shot ends on <Picture 3>

该镜头最终落在 <Picture 3> 所对应的画面上。

When editing or continuing an original video, cite <Video N> naturally where its source state, structure, or continuation relationship applies.

当对原始视频进行编辑或续写时,应在其源视频状态、视频结构或续写关系实际生效的位置,自然地引用 <Video N>。

Cite <Audio N> in the shot or semantic phase where the audio relationship is active.

对于 <Audio N>,则应在对应的音频关系实际生效的镜头或语义阶段中进行引用。

5.4 Speakers, Audio Sources, and Dialogue

说话者、音频来源与对白

The basic speaker-ID and <d> formats follow T2VA.

基本的说话者 ID 和 <d> 格式遵循 T2VA 的规则。

When a referenced subject physically speaks, retain both the visual reference label and the speaker ID:

当一个被引用的 <Subject N> 实际发声时,应同时保留其视觉参考标签和说话者 ID:
PROMPT 片段

<Subject 2> (S1) turns toward the woman and says, <d>[English] Last summer, I went to my grandfather's house. He talked about you.</d>

<Subject 2> (S1) 转向女性并说道,<d>[英文] 去年夏天,我去爷爷家拜访。他谈到了你。</d>

<Subject N> identifies the referenced subject, while (Sx) identifies the actual speaker.

<Subject N> 用于标识被引用的主体;(Sx) 用于标识实际的说话者。

When the subject speaks, write <Subject N> (Sx).

当该主体说话时,写作 <Subject N> (Sx)。

If the same subject speaks off-screen, keep the same form and mark it as off-screen.

如果同一个主体在画外说话,仍然使用相同的形式,并将其标记为 off-screen。

When the speaker does not correspond to a defined subject, use a stable voice description followed by (Sx).

如果说话者并不对应已经定义的 <Subject N>,则使用稳定的声音特征描述,并在其后标注 (Sx)。

When verbal content is only a cue within a directly reused BGM or complete soundtrack, and no person, character, narrator, or other independent vocal source physically produces it, use <Audio N> as the audible source and do not invent an additional (Sx).

当语言内容仅仅是直接复用的背景音乐(BGM)或完整音轨中的一个声音提示,并且没有人物、角色、旁白者或其他独立的人声来源实际发出这些声音时,应使用 <Audio N> 作为可听声音的来源,不要额外虚构一个 (Sx)。

If a concrete person, character, narrator, or other independent vocal source produces the voice, assign and reuse (Sx) for that source:

如果声音确实由某个具体人物、角色、旁白者或其他独立的人声来源发出,则应为该来源分配 (Sx),并在后续持续复用:
PROMPT 片段

When <Audio 1> reaches the phrase <d>[English] I'm lonely lonely lonely lonely lonely I'm lonely</d>, <Subject 1> performs the corresponding hand gesture without becoming a separate speaker source.

当 <Audio 1> 播放到 <d> 中的这句歌词时,<Subject 1> 做出与之对应的手势,但 <Subject 1> 不会因此被视为一个独立的说话者或人声来源。

When dialogue, narration, or lyrics from reference audio are directly reused, or when the input prompt explicitly requests their reperformance, preserve the exact source words and original language inside <d>.

当参考音频中的对白、旁白或歌词被直接复用,或者输入提示词明确要求重新演绎这些内容时,应在 <d> 中逐字保留源素材的原始内容和原始语言。对于无法听清的片段,应写作 [unclear],不要自行猜测或改写。

Write [unclear] for unintelligible spans instead of guessing or paraphrasing them. Standardize punctuation to the basic written marks needed to express the sentence, such as ,, ., ?, and !; remove repeated tildes, emoji, bullets, and repeated or decorative punctuation.

标点符号应规范为表达句子所需的基本书面标点,例如 ,、.、? 和 !;应删除重复的波浪号、Emoji、项目符号,以及重复或装饰性的标点。

End complete statements, questions, and exclamations with ., ?, or ! respectively before </d>.

完整陈述句以 . 结尾;疑问句以 ? 结尾;感叹句以 ! 结尾。在 </d> 之前。

When only timbre, rhythm, emotion, or delivery is referenced, do not carry the original dialogue from the reference audio into the target video.

当参考音频仅用于参考音色、节奏、情绪或表达方式时,不要将参考音频中的原始对白带入目标视频。

Assign (Sx) once according to the order of actual vocal events in the target video.

(Sx) 应按照目标视频中实际人声事件的出现顺序一次性分配。

Reuse the corresponding ID at every actual vocal event in detailed_description;

之后,在 detailed_description 中每次该人声来源实际发声时,都应复用对应的 ID。

an <Audio N> definition bound to a target speaker in subject_definitions also reuses the same (Sx) but never assigns a new one independently.

如果 subject_definitions 中的某个 <Audio N> 已经绑定到目标视频中的某位说话者,那么该定义也应复用相同的 (Sx),但不能在音频定义中独立分配新的说话者 ID。

Do not write (Sx) in retention_analysis.

不要在 retention_analysis 中写 (Sx)。

Verbal cues that exist only within a directly reused BGM or complete soundtrack use <Audio N>; voices physically produced by a concrete person, character, narrator, or other independent vocal source use (Sx).

区分原则可以概括为:仅存在于直接复用的 BGM 或完整音轨中的语言内容 → 使用 <Audio N>;由具体人物、角色、旁白者或其他独立人声来源实际发出的声音 → 使用 (Sx)。

6. overall_soundscape and non_diegetic_music

整体声景 和 非画面内音乐

The definitions of these two sound categories follow the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA).

这两类声音的定义遵循《Video Prompt Writing Guide(T2VA / I2VA / FL2VA / L2VA)》中的规则。

overall_soundscape summarizes ambience and physical sounds across the full video. Dialogue, singing, and sound events synchronized to a particular shot remain in detailed_description:

overall_soundscape(整体声景)用于概括整个视频中的环境声和物理动作声。对白、歌唱以及与特定镜头同步发生的声音事件,仍然写在 detailed_description(详细描述)中:
PROMPT 片段

overall_soundscape: Quiet indoor room tone and a low ventilation hum continue throughout the video.

整个视频中持续存在安静的室内环境底噪和低沉的通风设备嗡鸣声。

non_diegetic_music describes background music that the characters cannot hear and that is audible only to the audience.

non_diegetic_music(非画面内音乐)用于描述角色无法听到、只有观众能够听到的背景音乐。

When music is present, state its instrumentation, tempo, and dynamic development:

当存在此类音乐时,应说明其使用的乐器、速度以及动态变化:
PROMPT 片段

non_diegetic_music: A restrained solo-piano score at a slow tempo, with sustained low cello underneath and no swell.

缓慢速度的克制钢琴独奏配乐,下方持续铺有低音大提琴,没有明显的音量渐强。

When reference audio is used, state its copy or reference relationship only in the section that matches the audible layer:

当使用参考音频时,只应在与实际可听声音层相对应的部分中说明其复制或参考关系:

ambience and sound effects belong in overall_soundscape, while audience-only score belongs in non_diegetic_music.

环境声和音效属于 overall_soundscape(整体声景);只有观众能够听到的配乐属于 non_diegetic_music(非画面内音乐)。

If the same audio provides both kinds of content, describe the corresponding relationship in each section:

如果同一份音频同时提供这两类声音内容,则应分别在对应部分中描述其参考关系:
PROMPT 片段

overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.

overall_soundscape(整体声景):从 <Audio 1> 复制的环境声层贯穿整个目标视频。

non_diegetic_music: <Audio 2> is directly reused as the complete audience-only score.

non_diegetic_music(非画面内音乐):<Audio 2> 被直接完整复用为只有观众能够听到的配乐。

Write complete dialogue and lyrics only inside <d> in detailed_description;

完整的对白和歌词只能写在 detailed_description(详细描述)的 <d> 中;

do not repeat them in these two sections.

不要在 overall_soundscape(整体声景)和 non_diegetic_music(非画面内音乐)中重复这些内容。

7. Complete Example 完整示例

Show the complete example 查看完整示例
PROMPT · I2VA
subject_definitions(主体定义):

[<Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.

<Subject 1> 是 <Picture 1> 中的咖啡馆环境,包括裸露的砖墙、带有图案抱枕的橙色簇绒沙发、霓虹灯标牌以及木质咖啡桌。

<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.

<Subject 2> 是 <Picture 2>、<Picture 3> 和 <Picture 4> 中毛茸茸的白色萨摩耶犬,具有浓密的白色毛发、尖耳朵、深色鼻子和卷曲的尾巴。

<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.

<Subject 3> 是 <Video 1> 中的年轻金发女子,留着金色长发,身穿卷起袖子的浅粉色纽扣衬衫。

<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.

<Subject 4> 是 <Video 2> 中的年轻男子,留着棕色短卷发,身穿带抽绳的深灰色连帽衫。

<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.

<Audio 1>:reference(参考) - 其人声音色用于指导 <Subject 3> 的对白表达,但不复制原始音频信号。

detailed_description(详细描述):

The target video uses a realistic multi-camera sitcom style with warm indoor lighting.

目标视频采用写实的多机位情景喜剧风格,并使用温暖的室内灯光。

[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.

[Shot 1] 一个中景镜头首先展现 <Subject 1>:咖啡馆内有裸露的砖墙、带图案抱枕的橙色簇绒沙发、霓虹灯标牌以及木质咖啡桌。<Subject 3> (S1) 是一名留着金色长发、身穿卷袖浅粉色纽扣衬衫的年轻女子,她坐在沙发上,手中拿着一块巧克力豆曲奇。从画面左侧,<Subject 4> 进入画面。这名年轻男子留着棕色短卷发,身穿带抽绳的深灰色连帽衫,手中牵着 <Subject 2> 的狗绳。<Subject 2> 是一只毛发浓密的白色萨摩耶犬,有尖耳朵、深色鼻子和卷曲的尾巴。狗猛地扑向曲奇,将狗绳拉得绷紧。<Subject 3> (S1) 猛地把手缩回,并使用参考自 <Audio 1> 的清晰、年轻的声音音色,带着轻微的不满喊道:<d>[English] Hey! Watch your dog!</d> 她闭上嘴,把曲奇护住,同时 <Subject 4> 将狗拉了回来。

[Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.

[Shot 2] At 00:03.000,镜头切换到 <Subject 4> (S2) 的特写。这名年轻男子就是 Shot 1 中身穿深灰色连帽衫的人,此时他坐在沙发上 <Subject 3> 的旁边,并将 <Subject 2> 稳稳抱在怀中。<Subject 4> (S2) 使用随意的年轻男性声音,以略带玩笑的语气和自然轻松的谈话节奏说道:<d>[English] He just likes cookies more than me.</d> 说完后,他闭上嘴,露出带有歉意的微笑,并抚摸着狗浓密的白色毛发。

[Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.

[Shot 3] At 00:05.000,镜头切换到 <Subject 3> (S1) 的特写,她就是 Shot 1 中身穿浅粉色衬衫的金发女子。当她看向萨摩耶犬时,原本不满的神情逐渐缓和。<Subject 3> (S1) 继续使用参考自 <Audio 1> 的同一种清晰、年轻的声音音色,以带有笑意的节奏回应道:<d>[English] Well, he has good taste at least.</d> 她微笑着举起曲奇,做出一个类似举杯的小动作。台词结束后,经典的罐头观众笑声立即响起,并一直持续到最后一帧。

overall_soundscape(整体声景):

Soft indoor coffee-shop room tone continues throughout the scene.

整个场景中持续存在轻柔的咖啡馆室内环境底噪。

non_diegetic_music(非画面内音乐):

N/A

N/A

中文为对照译文,复制内容为英文原文。