音乐视频说唱口型同步
一个专业的音乐视频提示词,场景为白色无限摄影棚中的说唱歌手,具备精准的口型同步、镜像舞者回声效果以及高级时尚灯光。
- Category
- Charts & Infographics
- Model
- seedance-2.0
- Creator
- Kiber Alla
- Source language
- en
- Source ID
- 7343
- Published
- Jul 16, 2026
Full prompt
SHE RAPS THE VOCAL ON CAMERA — PRECISE LIPSYNC IS THE TOP PRIORITY. Mouth articulates every syllable of (Audio1) exactly on time; face visible and sharp through every vocal line, no cutaways mid-word. LIPSYNC MAP: 0.0–0.5 instrumental. 0.5–2.8 "I'm standing on the edge / Say it with your chest / Or keep it on the deck" + "Hey!" 3.2–6.7 "I walk in, whole room gets tense / I don't need luck, I'm the consequence / If you really want to test my intent / Come correct, come correct or get bent" + "Woo!" 7.5–13.7 same hook verbatim second time, escalated. 14.5–15.0 instrumental hold. Music-video route: performance with mirrored echoes. Director thesis: white infinity studio — she raps at the lens while five dancers in white repeat her last pose one beat late, a human delay effect behind her voice. Visual world: white cyclorama infinity studio, seamless floor and walls, one hard fashion key light with clean shadows, subtle floor reflections. 5 female background dancers in all-white utilitarian streetwear, hair slicked, deliberately similar but never identical to her; her indigo denim makes her instantly readable. Palette: white on white, skin tones, indigo. No neon, no particles. Shot flow: 0–2.8s symmetrical wide-to-medium push-in: she raps the opening lines at the apex of a tight wedge, the five echoes frozen in her exact stance behind her. 2.8–3.2s on "Hey!" all five snap chins up in unison. 3.2–6.7s medium: she raps while throwing an angular vogue accent at the end of each line, and the echoes replay that exact accent one beat later, rippling backward through the wedge; camera slowly orbits 45 degrees keeping her mouth front and center; finger to lens on "come correct". 7.5–10s cut on the kick to a chest-up close frame: second hook with doubled intensity, the echoes now a soft-focus rhythmic blur behind her articulation. 10–13.7s the echoes carousel slowly around her while she stands still at center rapping the final lines, camera counter-rotating, her face never leaving focus. 13.7–15s on "Woo!" she freezes arm high; the five echoes freeze in five different mid-move poses around her — she is the only resolved image; micro push-in, hold. Loopable. Performance rules: dominant, stoic, immaculate diction; echoes expressionless and precise, never mouth the words — only she raps. Continuity: same six women, same wardrobe, same white studio. Audio intent: (Audio1) only, her mouth locked to it; faint studio room tone. Quality bar: expensive fashion-campaign rap video, no AI gloss, no glow.
Translations
音乐视频说唱口型同步
enSHE RAPS THE VOCAL ON CAMERA — PRECISE LIPSYNC IS THE TOP PRIORITY. Mouth articulates every syllable of (Audio1) exactly on time; face visible and sharp through every vocal line, no cutaways mid-word. LIPSYNC MAP: 0.0–0.5 instrumental. 0.5–2.8 "I'm standing on the edge / Say it with your chest / Or keep it on the deck" + "Hey!" 3.2–6.7 "I walk in, whole room gets tense / I don't need luck, I'm the consequence / If you really want to test my intent / Come correct, come correct or get bent" + "Woo!" 7.5–13.7 same hook verbatim second time, escalated. 14.5–15.0 instrumental hold. Music-video route: performance with mirrored echoes. Director thesis: white infinity studio — she raps at the lens while five dancers in white repeat her last pose one beat late, a human delay effect behind her voice. Visual world: white cyclorama infinity studio, seamless floor and walls, one hard fashion key light with clean shadows, subtle floor reflections. 5 female background dancers in all-white utilitarian streetwear, hair slicked, deliberately similar but never identical to her; her indigo denim makes her instantly readable. Palette: white on white, skin tones, indigo. No neon, no particles. Shot flow: 0–2.8s symmetrical wide-to-medium push-in: she raps the opening lines at the apex of a tight wedge, the five echoes frozen in her exact stance behind her. 2.8–3.2s on "Hey!" all five snap chins up in unison. 3.2–6.7s medium: she raps while throwing an angular vogue accent at the end of each line, and the echoes replay that exact accent one beat later, rippling backward through the wedge; camera slowly orbits 45 degrees keeping her mouth front and center; finger to lens on "come correct". 7.5–10s cut on the kick to a chest-up close frame: second hook with doubled intensity, the echoes now a soft-focus rhythmic blur behind her articulation. 10–13.7s the echoes carousel slowly around her while she stands still at center rapping the final lines, camera counter-rotating, her face never leaving focus. 13.7–15s on "Woo!" she freezes arm high; the five echoes freeze in five different mid-move poses around her — she is the only resolved image; micro push-in, hold. Loopable. Performance rules: dominant, stoic, immaculate diction; echoes expressionless and precise, never mouth the words — only she raps. Continuity: same six women, same wardrobe, same white studio. Audio intent: (Audio1) only, her mouth locked to it; faint studio room tone. Quality bar: expensive fashion-campaign rap video, no AI gloss, no glow.
音乐视频说唱口型同步
zh-CN她对着镜头说唱 —— 精准的口型同步是首要任务。嘴部动作需与 (Audio1) 的每一个音节完全吻合;面部在每一句歌词中都清晰可见,严禁在词句中间进行切镜。 口型同步时间轴: 0.0–0.5 乐器演奏。 0.5–2.8 “I'm standing on the edge / Say it with your chest / Or keep it on the deck” + “Hey!” 3.2–6.7 “I walk in, whole room gets tense / I don't need luck, I'm the consequence / If you really want to test my intent / Come correct, come correct or get bent” + “Woo!” 7.5–13.7 重复第二遍副歌,情绪升级。 14.5–15.0 乐器停顿。 音乐视频路线:带有镜像回声效果的表演。导演构思:白色无限摄影棚 —— 她对着镜头说唱,身后五名身穿白衣的舞者比她慢一拍重复她的动作,形成一种人声背后的“人体延迟”效果。 视觉世界:白色无缝背景无限摄影棚,地面与墙面浑然一体,采用单一硬质时尚主光,阴影干净利落,地面有细微反射。5 名女性伴舞身穿全白实用主义街头服饰,发型利落,动作与她刻意相似但绝不完全相同;她身穿的靛蓝色牛仔装使其在画面中极具辨识度。色调:白底白调、肤色、靛蓝色。无霓虹灯,无粒子特效。 镜头流程: 0–2.8 秒:对称的广角到中景推镜头:她站在紧凑楔形队列的顶点说唱开场白,五名回声舞者在她身后保持着与她完全一致的姿势。 2.8–3.2 秒:在“Hey!”处,五名舞者同时抬头。 3.2–6.7 秒:中景:她一边说唱,一边在每句歌词末尾加入棱角分明的 Vogue 风格强调动作,回声舞者在慢一拍后重复该动作,像涟漪一样向后传递;摄像机缓慢旋转 45 度,始终将她的嘴部保持在画面正中央;在“come correct”处手指镜头。 7.5–10 秒:在鼓点处切换至胸部以上特写:第二段副歌强度加倍,回声舞者此时在她的发音身后形成柔焦的节奏模糊感。 10–13.7 秒:当她站在中心说唱最后几句歌词时,回声舞者围绕她缓慢旋转,摄像机反向旋转,她的面部始终保持对焦。 13.7–15 秒:在“Woo!”处她手臂高举定格;五名回声舞者在她周围定格在五个不同的动作中 —— 她是画面中唯一清晰的焦点;微距推镜头,保持。可循环。 表演准则:气场强大、冷峻、发音完美;回声舞者表情冷漠且精准,严禁对口型 —— 只有她进行说唱。连续性:相同的六名女性,相同的服装,相同的白色摄影棚。音频意图:仅限 (Audio1),口型需与之锁定;背景保留微弱的摄影棚环境音。质量标准:昂贵的时尚广告说唱视频质感,拒绝 AI 感,拒绝发光特效。














