H3 提示词写作技巧

发布于 2026-09-10

base-en.txt

MiniMax H3 官方提供,适用于 Base Model 类型任务的规范 base-en.txt, 包含中文对照翻译,调整了部分格式,便于阅读和理解,Agent 实际使用时以原文为准。

base-en.txt

Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)

视频提示词写作指南(T2VA / I2VA / FL2VA / L2VA)

1. Task Overview

任务概述
  • T2VA: Builds a complete audiovisual timeline from text.

    根据文本构建完整的视听时间线。
  • I2VA: T2VA body + first-frame instruction + a visual path that develops forward from the first frame.

    T2VA 主体结构 + 首帧指令 + 从首帧开始向后发展的视觉演变路径。
  • FL2VA: T2VA body + first-and-last-frame instruction + a continuous path from the first frame to the last frame.

    T2VA 主体结构 + 首尾帧指令 + 从首帧连续过渡到尾帧的视觉演变路径。
  • L2VA: T2VA body + last-frame instruction + a path that converges from a plausible preceding state to the last frame. T2VA 主体结构 + 尾帧指令 + 从合理的前置状态逐步过渡并最终收敛到尾帧的视觉演变路径。

2. Final Prompt Structure

最终提示词结构

2.1 Part One Is the Instruction

第一部分是指令

T2VA has no image-alignment instruction and begins directly with the three core fields.

T2VA 不包含图像对齐指令,直接从三个核心字段开始。

I2VA always uses (始终使用):

PROMPT 片段

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

对于目标视频,在 0.00 秒处完整参考 `<Picture 1>`(来自 [Shot 1])。

FL2VA always uses (始终使用):

PROMPT 片段

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.

参考图片与目标视频的对齐方式——Picture 1(来自 Shot 1)与目标视频的 0.00 秒处对齐;Picture 2(来自 Shot N)与目标视频的 S.SS 秒处对齐。

L2VA always uses (始终使用):

PROMPT 片段

How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.

参考图片与目标视频的对齐方式——<Picture 1>(来自 [Shot N])与目标视频的 S.SS 秒处对齐。

Here, N is the index of the actual final shot, and S.SS is the effective video duration formatted to exactly two decimal places.

其中,N 表示实际最后一个镜头的序号,S.SS表示视频的实际有效时长,并严格格式化为两位小数。

The instruction must be the first line of the final prompt, followed by one blank line before the core fields.

该指令必须作为最终提示词的第一行,之后空一行,再开始编写核心字段。

2.2 Part Two Contains the Three Core Fields

第二部分包含三个核心字段
PROMPT 片段

integrated_multimodal_description: [Shot 1] ...

overall_soundscape: ...

non_diegetic_music: ...

  • integrated_multimodal_description: Describes visuals, actions, shots, speakers, dialogue, singing, and diegetic audio along the timeline.

    沿时间线描述画面、动作、镜头、说话者、对话、歌唱以及画面内(diegetic)声音。
  • overall_soundscape: Summarizes ambient sound, physical action sounds, and non-verbal human sounds across the entire video.

    概括整个视频中的环境声、物理动作产生的声音,以及非语言类人声。
  • non_diegetic_music: Describes background music that the characters cannot hear and only the audience can hear. 描述非画面内(non-diegetic)的背景音乐,即角色无法听到、只有观众能够听到的音乐。

3. How to Incorporate Keyframes into the Multimodal Description

如何将关键帧融入多模态描述

3.1 I2VA: Begin from the Image and Develop Forward

从图像开始并向后发展

<Picture 1> is the actual first frame of the video at 0.00 seconds and belongs to [Shot 1].

<Picture 1> 是视频在 0.00 秒处的实际首帧,属于 [Shot 1]。

The description should first establish the style, subjects, composition, and scene anchors in the image, then describe the next action.

描述时应首先明确图像中的风格、主体、构图和场景锚点,然后描述接下来发生的动作。

Character identity, clothing, colors, key objects, and spatial relationships should remain consistent.

角色身份、服装、色彩、关键物体以及空间关系应保持一致。

Recommended structure: first-frame anchor → action onset → continuous development → result or reaction.

推荐结构:首帧锚定 → 动作开始 → 连续发展 → 结果或反应。

3.2 FL2VA: Describe the Path Between the First and Last Frames

描述首帧与尾帧之间的演变路径

Picture 1 is the opening, and Picture 2 is the ending.

Picture 1 是起始画面,Picture 2 是结束画面。

Focus on how the subject moves, how poses change, how objects are manipulated, how the composition evolves, and how the scene or lighting transitions.

重点描述主体如何移动、姿态如何变化、物体如何被操作、构图如何演变,以及场景或光线如何过渡。

FL2VA generally favors a single shot so the model can interpolate continuously from the first frame to the last frame.

FL2VA 通常优先采用单镜头,以便模型能够从首帧连续地插值过渡到尾帧。

Use multiple shots only when they are explicitly specified.

只有在明确指定多个镜头时才使用多镜头。

The last frame must be reached by the final [Shot N] at the end of the video.

视频结束时,必须在最后一个 [Shot N] 到达所提供的尾帧状态。

Recommended structure: first-frame state → observable intermediate changes → progressively narrowing differences → last-frame state.

推荐结构:首帧状态 → 可观察的中间变化 → 与目标尾帧的差异逐步缩小 → 尾帧状态。

3.3 L2VA: Infer the Opening and Land on the Image at the End

推断起始状态,并在视频结尾落到参考图像

<Picture 1> is the final frame of the video and belongs to the last [Shot N]; it does not inherently belong to Shot 1.

<Picture 1> 是视频的最终帧,属于最后一个 [Shot N],并不天然属于 Shot 1。

Infer a plausible earlier state from the user's intent and the last frame, then describe how the characters, objects, camera, and scene gradually approach the reference image.

应根据用户意图和尾帧内容推断一个合理的前置状态,然后描述角色、物体、镜头和场景如何逐步向参考图像中的最终状态靠拢。

Recommended structure: plausible preceding state → explicit action and transition path → gradual convergence in the final shot → last-frame landing.

推荐结构:合理的前置状态 → 明确的动作与过渡路径 → 在最后一个镜头中逐步收敛 → 落到尾帧状态。

4. How to Write the Three Shared Core Sections

如何编写三个通用核心部分

4.1 Develop the Multimodal Description Along the Timeline

沿时间线展开多模态描述

integrated_multimodal_description is the main body of the rewritten prompt.

integrated_multimodal_description 是改写后提示词的主体部分。

Every detail should correspond to something visible or audible: visual style, initial composition, subject appearance and position, scene and key props, actions and reactions, shot changes, spoken language, and synchronized diegetic sound.

每一个细节都应对应实际可见或可听的内容,包括:视觉风格、初始构图、主体的外观与位置、场景与关键道具、动作与反应、镜头切换、对白语言,以及与画面同步的画面内声音(diegetic sound)。

At the beginning of [Shot 1], state the overall style and initial composition. Common styles include Cinematic, live-action, 2D-animated, 3D CG, claymation, watercolor, and vintage film.

在 [Shot 1] 开始时,应说明整体视觉风格和初始构图。常见风格包括 Cinematic、live-action、2D-animated、3D CG、claymation、watercolor 和 vintage film。

For keyframe tasks, derive the style from the reference image; for T2VA, select it from the user's text.

对于关键帧任务,应根据参考图像确定风格;对于 T2VA,则根据用户的文本描述选择合适的风格。
PROMPT 片段

[Shot 1] Live-action, cinematic, a medium-wide shot frames...

[Shot 1] 实况拍摄,电影感,中景镜头

4.2 Shots and Cuts

镜头与切换

Do not add a timestamp to the first shot. Use sequential shot numbers for later shots, and begin each one with a strictly increasing cut time that falls within the video duration:

第一个镜头不要添加时间戳。后续镜头使用连续递增的镜头编号,并且每个镜头都应以严格递增、且位于视频总时长范围内的切换时间开始:

PROMPT 片段

[Shot 2] At 00:03.500, the camera cuts to...

[Shot 2] 在 00:03.500 时,镜头切换到...

For ordinary cuts, use the camera cuts to, the shot cuts to, the shot transitions to, the shot changes to, or the shot switches to.

普通镜头切换可使用 the camera cuts to、the shot cuts to、the shot transitions to、the shot changes to 或 the shot switches to。

When explicitly requested by the user, cross-dissolve, fade, or wipe may also be used.

当用户明确要求时,也可以使用交叉溶解(cross-dissolve)、淡入淡出(fade)或擦除转场(wipe)。

A cut should introduce new information about the subject, space, state, viewpoint, or time.

一次镜头切换应带来关于主体、空间、状态、视角或时间的新信息。

If only the distance or a slight angle needs to change, prefer camera motion.

如果只是需要改变拍摄距离或轻微调整角度,应优先使用镜头运动,而不是切换镜头。

4.3 Camera Motion: Motion Type + Amplitude + Speed

镜头运动:运动类型 + 幅度 + 速度

A complete camera-motion expression has three dimensions:

一个完整的镜头运动表达包含三个维度:
  • the motion type defines how the camera moves;

    运动类型(motion type) 定义镜头如何运动
  • amplitude defines the range of compositional change;

    幅度(amplitude)定义构图变化的范围
  • and speed defines the pacing of that change.

    速度(speed) 定义这一变化的节奏

Add amplitude and speed only when they are meaningful; medium amplitude and normal speed are usually omitted.

只有在幅度和速度具有明确意义时才需要写出;中等幅度和正常速度通常可以省略。
Dimension 维度Available Expression 可用表达Description 说明
Motion type 运动类型Zoom In / Zoom OutThe focal length changes while the camera body remains stationary 机位保持不动,通过改变焦距实现放大 / 缩小
Motion typePush In / Pull OutThe camera moves forward / backward 摄像机向前 / 向后移动
Motion typePan Left / Pan RightThe camera remains in place while the lens pivots horizontally 摄像机位置不变,镜头水平向左 / 向右转动
Motion typeTruck Left / Truck RightThe camera translates horizontally 摄像机整体水平向左 / 向右移动
Motion typeTilt Up / Tilt DownThe camera remains in place while the lens pivots vertically >摄像机位置不变,镜头垂直向上 / 向下转动
Motion typePedestal Up / Pedestal DownThe entire camera moves upward / downward 摄像机整体向上 / 向下移动
Motion typeArc ShotThe camera moves in an arc around the subject 摄像机围绕主体做弧线运动
Motion typeTracking ShotThe camera follows a moving subject 摄像机跟随运动中的主体移动
Motion typeStatic ShotThe camera position and lens remain still 摄像机位置和镜头均保持静止
Motion typeShake Slightly / Shake StronglySlight / strong camera shake 镜头轻微 / 强烈抖动
Motion typePOVThe subject's point of view 采用主体的第一人称视角
Motion typeRoll Clockwise / Roll CounterclockwiseThe camera rolls clockwise / counterclockwise around the lens axis 摄像机绕镜头光轴顺时针 / 逆时针旋转
Amplitude 幅度with small amplitudeSmall-range change 小幅度变化
Amplitudewith large amplitudeLarge-range change 大幅度变化
Speed 速度at slow speedSlow movement 缓慢运动
Speedat fast speedFast movement 快速运动

Camera motion should be written as a natural English action within the shot, rather than stacked as separate labels at the end of a sentence:

镜头运动应作为镜头描述中的自然英文动作来编写,而不是在句子末尾将多个标签单独堆叠:
PROMPT 片段

The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.

摄像机以小幅度、慢速向前推进(Push In),逐渐靠近她手中折叠的信件。

PROMPT 片段

The camera pans right with large amplitude at fast speed, revealing the open doorway.

摄像机以大幅度、快速向右摇摄(Pan Right),露出敞开的门口。

PROMPT 片段

The camera holds a static shot as the runner exits the frame.

摄像机保持固定镜头(Static Shot),直到跑步者离开画面。

4.4 Speakers, Dialogue, and Singing

说话者、对白与歌唱

Subjects who speak, sing, or produce an off-screen human voice use stable IDs such as (S1) and (S2).

对于说话、唱歌或发出画外人声的主体,使用稳定的 ID,例如 (S1) 和(S2)。

When multiple already-numbered speakers speak or sing together, use a compound ID such as (S1,S2).

当多个已经编号的说话者一起说话或唱歌时,使用组合 ID,例如 (S1,S2)。

A speaker keeps the same ID across shots; characters who never vocalize receive no speaker ID.

同一个说话者在不同镜头中始终保持相同的ID;从未发声的角色不需要分配说话者 ID。

When a speaker first appears, provide enough information from the visual and audio context to establish a stable identity, such as character type, age, gender, whether the person is on-screen, pitch, timbre, speaking rate, or accent.

当说话者首次出现时,应根据视觉和音频上下文提供足够的信息,以建立稳定的身份,例如角色类型、年龄、性别、是否出现在画面中、音高、音色、语速或口音。

Place the speaker's identifying phrase, ID, action, and delivery outside <d>.

说话者的身份描述、ID、动作和表达方式应放在 <d> 之外。

Inside <d>, include only the language tag and the actual user-provided spoken content.

<d> 内只包含语言标签和用户实际提供的对白内容。

Preserve every original word and punctuation mark verbatim; do not translate or rewrite them.

所有原始文字和标点符号都必须逐字保留,不得翻译或改写。
PROMPT 片段

The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>

一名声音轻柔、带有气声的年轻女子 (S1) 说道:<d>[English] 我在下一站下车。</d>

PROMPT 片段

The two children (S1,S2) shout together, <d>[English] Wait for us!</d>

两个孩子 (S1,S2) 一起喊道:<d>[English] 等等我们!</d>

For voiceover, use the exact phrase says in an off-screen voiceover.

对于画外音(voiceover),使用固定短语 says in an off-screen voiceover。

Immediately after every voiceover <d> block, state that the corresponding on-screen character's lips remain closed:

每个画外音 <d> 块之后,应立即说明对应的画面内角色嘴唇始终保持闭合:
PROMPT 片段

The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.

男子 (S1) 以画外音说道:<d>[English] 我仍然记得那条路。</d>,同时画面中他的嘴唇始终完全闭合。

When the same line of dialogue or lyrics crosses a cut, use <scenetrans> at the connecting points in both parts and explicitly state that the audio continues across the cut.

当同一句对白或歌词跨越镜头切换时,应在前后两部分的衔接位置使用 <scenetrans>,并明确说明音频在切镜过程中持续播放。

Use <cutoff> when speech is truncated by the end of the video.

如果说话内容因视频结束而被截断,则使用 <cutoff>。

Continuity may be expressed with:

对白或歌词跨镜头的连续性可以使用以下表达:
  • continues seamlessly across the cut

    在切镜过程中无缝延续
  • continues uninterrupted into the next shot

    不间断地延续到下一个镜头
  • carries over from the previous shot

    从上一个镜头延续过来
  • remains audible across the transition.

    在镜头过渡期间始终保持可听

4.5 On-Screen Text

画面内文字

Place any banner, sign, label, subtitle, or neon text that is actually visible on screen in English double quotation marks.

对于实际出现在画面中的横幅、标牌、标签、字幕或霓虹灯文字,应使用英文双引号括起来。

Preserve the original text and punctuation verbatim, without translation.

原始文字和标点符号必须逐字保留,不得翻译。
PROMPT 片段

A red neon sign reading "营业中" glows above the doorway.

门口上方有一块写着 "营业中" 的红色霓虹灯招牌正在发光。

4.6 overall_soundscape

Use 1–4 English sentences in one continuous paragraph to summarize the ambient sound, physical action sounds, and non-verbal human sounds across the full video, such as wind, rain, traffic, footsteps, fabric movement, impacts, breathing, laughter, or panting.

使用 1–4个英文句子,以一个连续段落概括整个视频中的环境声、物理动作产生的声音以及非语言类人声,例如风声、雨声、交通声、脚步声、衣物摩擦声、撞击声、呼吸声、笑声或喘息声。

Dialogue, singing, and diegetic music already belong in the multimodal description and should not be repeated here.

对白、歌唱以及画面内音乐(diegetic music)已经属于多模态描述的一部分,不应在这里重复。

Use N/A only when the user explicitly requests complete silence throughout the video.

只有当用户明确要求整个视频完全静音时,才使用 N/A。
PROMPT 片段

overall_soundscape: Steady rain taps against the café windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.

持续的雨滴敲打着咖啡馆的窗户,同时伴随着低沉的室内环境声。入口处的门铃响了一次,随后传来湿漉漉的脚步声和椅子轻轻摩擦地面的声音。

4.7 non_diegetic_music

Use 1–3 English sentences to describe background music that the characters cannot hear and only the audience can hear.

使用 1–3 个英文句子描述角色无法听到、只有观众能够听到的背景音乐。

Focus on instrumentation, speed, rhythm, and dynamic changes; do not use abstract mood words or explain the emotional function of the score.

重点描述乐器、速度、节奏以及动态变化;不要使用抽象的情绪词,也不要解释配乐所承担的情感作用。

Singing, instruments, radio, television, or phone music audible to the characters are diegetic events and should appear in the multimodal description.

如果歌声、乐器声、收音机、电视或手机播放的音乐能够被角色听到,那么它们属于画面内事件(diegetic events),应写入多模态描述中。

Use N/A when there is no non-diegetic music.

当不存在非画面内音乐(non-diegetic music)时,使用 N/A。
PROMPT 片段

non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.

稀疏的钢琴音符以缓慢的速度响起,随后加入持续的低音弦乐,音量逐渐增强,最后慢慢淡出。

5. Cases

Case 1: T2VA

With no reference image, construct the complete timeline directly from the text.

在没有参考图像的情况下,直接根据文本构建完整的时间线。

You may add scene, character, action, and sound details that remain consistent with the user's intent.

可以补充场景、角色、动作和声音等细节,但这些内容应与用户的原始意图保持一致。
PROMPT · T2VA
integrated_multimodal_description:

[Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle-aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: <d>[English] First batch of the morning.</d>

[Shot 1] 真人实拍、电影感,中广角镜头拍摄日出前,一名面包师正在打开街边一家小面包店的百叶窗。摄像机以小幅度、慢速向前推进(Push In)。一名声音平静、略带沙哑的中年面包师 (S1) 将一条刚出炉的面包放到木质柜台上,并说道:<d>[English] 今天早上的第一炉。</d>

[Shot 2] At 00:05.000, the camera cuts to a close-up of steam rising from the sliced bread while the baker's final words carry over from the previous shot.

[Shot 2] 在 00:05.000,镜头切换到切开的面包特写,热气正从面包中缓缓升起,同时面包师最后说出的几个词从上一个镜头延续到这里。

overall_soundscape:

Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced.

木质百叶窗摩擦着打开,街道十分安静,同时面包店内传来托盘轻微碰撞的声音。门铃响了一次,随后传来轻盈的脚步声,以及切开面包时清脆的声音。

non_diegetic_music:

A soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes and a gentle fade at the end.

中等速度的轻柔原声吉他节奏,随后加入稀疏的低音提琴音符,并在结尾处轻柔地淡出。

中文为对照译文,复制内容为英文原文。

Case 2: I2VA

Write the first-frame instruction first, then use the subject, composition, and scene in Picture 1 as the starting point of Shot 1 before describing how the scene continues to develop.

首先编写首帧指令,然后以 Picture 1 中的主体、构图和场景作为 [Shot 1] 的起始状态,再描述场景如何继续向后发展。

PROMPT · I2VA

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

对于目标视频,在目标视频的 0.00 秒处,完整参考 <Picture 1>(来自 [Shot 1])。

integrated_multimodal_description:

[Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: <d>[English] I get off at the next station.</d> She folds the letter along its existing crease.

[Shot 1] 真人实拍、电影感,<Picture 1> 中的年轻女子仍坐在布满雨水的列车窗边,保持她的外貌、服装、座位位置以及车厢布局不变。摄像机以小幅度、慢速向右平移(Truck Right)。她将视线从折叠的信件上抬起,望向窗外掠过的城市灯光。她的倒影在玻璃上移动,这名声音轻柔、带有气声的年轻女子 (S1) 说道:<d>[English] 我在下一站下车。</d> 随后,她沿着信件原有的折痕将其折起。

overall_soundscape:

The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.

在低沉的通风系统嗡鸣声下,列车车轮发出稳定而富有节奏的金属声。雨滴敲打着车窗,同时她手中的纸张发出轻微的沙沙声。

non_diegetic_music:

Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.

缓慢节奏的持续大提琴音符,配合间隔较大的钢琴音,音量逐渐降低。

中文为对照译文,复制内容为英文原文。

Case 3: FL2VA

The two images anchor the opening and ending respectively. The body should not repeat two static image descriptions; instead, it should supply the motion path that connects them. The following example is an eight-second single shot.

两张图像分别作为视频的起始帧和结束帧。主体描述不应只是重复描述两张静态图像,而应补充连接二者的连续运动与变化路径。下面的示例是一个 8 秒的单镜头视频。

PROMPT · FL2VA

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.

参考图片与目标视频的对齐方式——Picture 1(来自 Shot 1)与目标视频的 0.00 秒处对齐;Picture 2(来自 Shot 1)与目标视频的 8.00 秒处对齐。

integrated_multimodal_description:

[Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot.

[Shot 1] 真人实拍、电影感。一名被雨水淋湿的骑行者以 Picture 1 所确定的位置和构图作为起始状态,站在一辆银色自行车旁,手中拿着一把收起的黑色雨伞。摄像机以小幅度、慢速向后拉远(Pull Out)。她松开自行车把手,将雨伞举到肩膀上方,并向上推动伞的滑套,直到伞面完全撑开。雨水从逐渐展开的伞面上滑落,她迈步走到伞下,转动伞柄使其达到最终角度,并在镜头结束时逐渐调整到 Picture 2 所确定的姿态、间距和构图。

overall_soundscape:

Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes.

雨水持续落在路面上,随后传来雨伞滑套移动时的金属咔嗒声,以及伞面撑开时轻微而清脆的弹响。雨水从自行车车架上滴落,远处不时传来车辆驶过的声音。

non_diegetic_music:

N/A

中文为对照译文,复制内容为英文原文。

Case 4: L2VA

The image anchors only the final moment. First establish a compatible earlier state, then let the actions, object states, and composition gradually land on Picture 1 in the final shot. The following example is a six-second single shot.

参考图像只用于锚定视频的最终时刻。首先建立一个与尾帧相符的合理前置状态,然后让动作、物体状态和构图逐步变化,并在最后一个镜头中收敛到 Picture 1 所确定的最终状态。下面的示例是一个 6 秒的单镜头视频。

PROMPT · L2VA

How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 6.00-second mark of the target video.

参考图片与目标视频的对齐方式——<Picture 1>(来自 [Shot 1])与目标视频的 6.00 秒处对齐。

integrated_multimodal_description:

[Shot 1] Live-action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in <Picture 1> approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by <Picture 1>.

[Shot 1] 真人实拍、电影感。特写镜头开始时,一只完好的玻璃杯放在深色木桌的边缘附近,同时 <Picture 1> 中可见的同一只手和衣袖从画面右侧靠近。摄像机以小幅度、慢速向前推进(Push In),与此同时,指尖碰到玻璃杯的杯沿。玻璃杯倾斜、掉落,并猛烈撞击地面;裂纹迅速蔓延,碎片向四周滑动。接近视频结尾时,仍在移动的碎片逐渐失去动量并停止,最终准确落到 <Picture 1> 所确定的玻璃破碎形态、手部位置、摄像机角度、光照以及最终构图。

overall_soundscape:

Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor.

指尖轻碰玻璃杯,随后玻璃杯摩擦桌面、掉落,并伴随着一声清脆而猛烈的碎裂声。细小的玻璃碎片向四周散开,并逐渐停止在地面上滑动。

non_diegetic_music:

A low electronic pulse at a slow tempo, ending immediately after the glass breaks.

缓慢节奏的低沉电子脉冲音,在玻璃杯破碎后立即结束。

中文为对照译文,复制内容为英文原文。