Generate Multimodal Video Script

Generate a multimodal video script for short-form content, leveraging insights from text, images, audio, and existing video inputs.

תחום: תוכן

מתי להשתמש

Use when the user needs a detailed script for a short AI-generated video, incorporating various media elements.

מתי זה מופעל

when the user asks to generate a script for an AI multimodal video

תגיות: video-generation, multimodal-content, ai-scripting, short-form-video

SKILL.md

--- name: Generate Multimodal Video Script description: Generate a multimodal video script for short-form content, leveraging insights from text, images, audio, and existing video inputs. when_to_trigger: when the user asks to generate a script for an AI multimodal video --- ## Overview Generate a compelling, short-form video script (up to 10 seconds) designed for multimodal AI video generation platforms like Gemini Omni Flash. Integrate provided text, image, audio, and video inputs to create a cohesive narrative and visual flow. Focus on engagement and platform-specific best practices. ## When to use Use when a user provides various content elements (text, images, audio, or existing video snippets) and requests a script for a short, AI-generated video. This is ideal for quickly prototyping video content or repurposing existing assets into new, dynamic formats. ## How it works 1. Receive and analyze user-provided text, image descriptions, audio cues, and existing video input details. 2. Identify key themes, emotions, and narrative arcs present across all input modalities. 3. Brainstorm potential visual and auditory elements that can be generated or enhanced by AI, based on the identified themes. 4. Structure the video script into concise segments, each with a clear visual direction, audio suggestion (narration, music, sound effects), and on-screen text/graphics. 5. Ensure the total script duration is appropriate for short-form content (aim for 5-10 seconds). 6. Incorporate dynamic transitions and AI-powered enhancements where relevant, explicitly noting these in the script. 7. Format the output as a detailed script including a timeline, visual descriptions, audio cues, and potential on-screen text. ## Example usage **User:** "Generate a 10-second video script for a new coffee product. I have a slogan 'Awaken Your Day', an image of a steaming mug, and a short audio clip of gentle jazz music. I'd like a calming and uplifting feel." **Output:** ```markdown ## Multimodal Video Script: Awaken Your Day Coffee **Video Duration:** 8 seconds **Theme:** Calming, Uplifting, Fresh Start --- **Scene 1 (0-2 seconds):** * **Visuals:** Transition from dark, abstract swirls to a close-up of steam rising from a porcelain mug (from user image input, enhanced with soft, warm lighting). AI-generated subtle slow-motion effect. * **Audio:** Gentle jazz music fades in (from user audio input). * **On-screen Text:** "Awaken Your Day" (slow fade-in, elegant font). **Scene 2 (2-5 seconds):** * **Visuals:** AI-generated animation of coffee beans swirling and transforming into a rich, dark liquid filling the mug. Sunlight gradually illuminates the scene from the side. * **Audio:** Jazz music continues. Soft, subtle pouring sound effect. * **On-screen Text:** "Rich Aroma. Smooth Taste." (quick fade-in/out). **Scene 3 (5-8 seconds):** * **Visuals:** Overhead shot of a hand gently reaching for the mug. Soft focus on the mug, background blurred. AI-generated shimmering light effect. * **Audio:** Jazz music gently fades out. Sound of a contented sigh (subtle, optional). * **On-screen Text:** "Start your morning right. [Website/Brand Name]" (bold, clear call to action, remains on screen for 2 seconds). --- ``` ## References * [Gemini Omni Flash: multimodal video generation for content workflows](https://letsdatascience.com/news) * [How I created the AI Cornflakes video in Chef's Games? (Full explanation)](https://www.youtube.com/watch?v=nP4DAPZml7o)