The Multimodal AI Prompt Playbook: Text, Images, and Video in One Workflow

Most AI users treat text, images, and video as separate worlds. You prompt ChatGPT for copy, then switch to Midjourney for visuals, then to some video tool for clips—and the moment you try to combine them, the whole workflow collapses. That's where multimodal AI changes everything. These systems handle text, images, and video in a single context, so you can generate an image, ask about its contents, edit it, and produce a video script—all in one conversation.

But here's the catch: most prompts are written for single-modality models. They fall apart when the input includes both a JPEG and a text block. To get real results, you need prompts that explicitly reference the different modalities and their relationships. This guide gives you ready-to-use prompts for practical tasks—content creation, analysis, education, and development—with real examples and sources to back them up.

1. From Idea to Storyboard: Generating a Visual Concept from a Text Prompt

When to use: You have a rough idea and need a visual concept—like a storyboard for a product video, a series of illustrations for a blog post, or a character design for an animation.

Why it works: The prompt forces the AI to translate abstract text into a concrete visual description, and then into a storyboard format. This is a classic multimodal pattern: text → structured visual specification → image generation.

The prompt:

Convert the following product description into a 6-scene storyboard for a 30-second video. For each scene, provide: (1) a detailed visual description including camera angle, lighting, and composition; (2) a short text overlay or voiceover; (3) a Midjourney-style prompt for generating the scene image. Use a consistent character and color palette across all scenes.

Product: [your product description]

Example: If your product is a smart water bottle, the AI might generate scene 1 as 'Close-up of the bottle on a desk, morning light, phone notification about hydration' and a Midjourney prompt like 'smart bottle, morning light, photorealistic, shallow depth of field'.

2. Reverse Prompting: Extracting a Prompt from an Existing Image

When to use: You have an image (e.g., from a stock site or a competitor) and want to recreate a similar style or understand its composition for your own project. This is a key reverse-engineering technique.

Why it works: Multimodal models can analyze an image and describe it in terms of style, subject, and composition—essentially generating a prompt that could have created it.

The prompt:

Analyze the attached image and provide a detailed description in the following categories: (1) subject and setting; (2) art style and medium (e.g., photography, digital art, watercolor); (3) color palette with hex codes; (4) lighting and mood; (5) composition and focal points. Then write a single text-to-image prompt that would recreate this image, optimized for [Midjourney / DALL·E / Stable Diffusion].

Example: Upload a photo of a neon-lit city street. The AI might return: 'Cyberpunk, rainy night, neon signs, blue and magenta palette, wide-angle, cinematic lighting' and a prompt like 'a wet cyberpunk street at night, neon signs, photorealistic, cinematic, 8k'.

3. Visual Q&A: Asking Questions About an Image for Analysis

When to use: You have an image and need specific information—like identifying objects, reading text, assessing damage, or understanding a chart. This is the most common multimodal task.

Why it works: The prompt structures the analysis, ensuring the AI doesn't give a generic description but answers your exact question.

The prompt:

Look at the attached image and answer the following questions with specific evidence from the image:
1. What is the main subject and its approximate size?
2. Are there any text elements? If so, transcribe them exactly.
3. What is the estimated color distribution (dominant colors, approximate percentages)?
4. Is there anything unusual or unexpected in the image? Describe it.
5. If this image is a chart/graph, summarize the key data points and trends.

If you cannot determine something, say 'Cannot determine' rather than guessing.

Example: For a bar chart of quarterly sales, the AI might say 'The chart shows sales in Q1 2026 at $120K, Q2 at $150K, a 25% increase...'—with specific numbers extracted from the image.

4. Script to Video: Generating a Video Script from a Text Outline and a Reference Image

When to use: You need a video script that aligns with a specific visual style or brand image. The prompt combines a text outline with an image reference to produce a script with visual cues.

Why it works: The reference image anchors the visual style, and the outline provides structure. The AI generates a script that includes camera directions, on-screen text, and narration.

The prompt:

Here is a text outline for a video: [paste outline]. Here is a reference image of our brand's visual style: [attach image]. Write a 60-second video script that includes:
- Visual description for each shot, referencing the style of the attached image
- Narration or voiceover text (max 150 words)
- On-screen text overlays (max 10 words per screen)
- Transition notes (e.g., cut, fade, zoom)

Keep the tone [formal/casual/inspirational].

Example: If your outline is 'Intro → Problem → Solution → CTA' and the image is a pastel minimal style, the script will specify 'soft pastel backgrounds, minimalist props' in each shot.

5. Visual Critique: Getting Constructive Feedback on an Image

When to use: You've created an image (e.g., for a social media post) and want professional feedback on composition, color, and message clarity before publishing.

Why it works: The prompt applies design principles to the image analysis, giving you actionable suggestions rather than generic praise.

The prompt:

Act as an expert graphic designer. Analyze the attached image and provide:
1. Three strengths (specific, e.g., 'strong focal point on the product')
2. Three areas for improvement (specific, e.g., 'text contrast is low on the left side')
3. A suggested color correction (if any) and why
4. An alternative composition that might improve visual impact
5. A rating on a scale of 1-10 for: composition, color harmony, clarity of message, and overall appeal.

Be brutally honest and specific.

Example: For a product photo, the AI might say 'The background is too cluttered; try a shallower depth of field' and 'The product is not centered; consider rule of thirds.'

6. Image-to-Text for Accessibility: Generating Alt Text and Descriptions

When to use: You need alt text for website images, social media, or documentation. This is crucial for accessibility (WCAG) and SEO.

Why it works: The prompt ensures the AI generates concise, descriptive alt text that includes relevant keywords without keyword stuffing.

The prompt:

Write alt text for the attached image according to WCAG 2.1 guidelines. The alt text should be:
- Concise (under 125 characters)
- Descriptive of the content and function
- Include relevant keywords naturally, but do not stuff
- Avoid phrases like 'image of' or 'photo of' unless necessary

Also provide a longer description (2-3 sentences) for complex images that might be used as a caption.

Example: For a photo of a laptop on a desk, the alt text might be 'Laptop on a wooden desk in a home office with a coffee mug.'

7. Video Content Analysis: Summarizing a Video File

When to use: You have a video file (e.g., a webinar recording, a product demo) and need a summary, key points, or a timestamped list of topics. This is a powerful time-saver.

Why it works: Multimodal models with video understanding can process frames and audio, extracting the narrative structure.

The prompt:

Analyze the attached video file. Provide:
1. A 100-word executive summary
2. A list of key topics with approximate timestamps (mm:ss)
3. Any important quotes or statements, with timestamps
4. The overall tone and target audience
5. Suggestions for how this video could be repurposed into blog posts, social media clips, or infographics

Example: For a 30-minute webinar, the AI might return '0:00 Intro, 2:15 Problem statement, 5:30 Solution overview...' and a summary of the main takeaways.

8. Visual Content Repurposing: Turning an Image into a Social Media Post

When to use: You have a single image (e.g., an infographic, a photo) and want to create a series of social media posts (different platforms) from it.

Why it works: The prompt guides the AI to extract key information from the image and adapt it to different formats and tones.

The prompt:

You have the attached image. Create 3 social media posts based on it:
1. LinkedIn post (professional, 150-200 words) with a hook, key points, and a call to action.
2. Instagram post (shorter, 50-80 words) with relevant hashtags.
3. Twitter/X post (under 280 characters) that is catchy and shareable.

Each post should reference the image content and include a suggestion for the visual (e.g., crop, caption overlay).

Example: For an infographic about AI trends, the LinkedIn post might start with 'AI is changing the game—here's what you need to know' and list key stats from the image.

9. Cross-Modal Translation: Converting a Chart into a Narrative Report

When to use: You have a data visual (chart, graph) and need a written analysis for a report or presentation.

Why it works: The prompt forces the AI to interpret the visual data and present it in a structured textual format.

The prompt:

The attached image is a data visualization. Convert it into a professional report section with:
1. A one-sentence summary of the main trend
2. A bulleted list of 3-5 key findings, each with a specific data point from the chart
3. A short paragraph explaining the implications for [industry/domain]
4. A caveat about any limitations (e.g., missing data, unclear axis labels)

Use formal business language.

Example: From a line chart showing website traffic over time, the AI might say 'Traffic peaked in March at 45K visits, a 20% increase from January' and suggest implications for marketing strategy.

10. Teaching with Visuals: Creating an Explainer from a Screenshot

When to use: You have a screenshot of a software interface, a diagram, or a process flow, and you need to create a tutorial or explanation for non-technical users.

Why it works: The prompt combines the visual with didactic structure, producing a step-by-step guide.

The prompt:

You are an expert technical trainer. Analyze the attached screenshot and create a beginner-friendly tutorial that includes:
1. An introduction to what the user is looking at
2. Step-by-step instructions for the main task (numbered)
3. Annotated callouts referencing specific parts of the image (e.g., 'the red button in the top-right corner')
4. Common mistakes and how to avoid them
5. A summary and a practice exercise

Assume the user has no prior knowledge.

Example: For a screenshot of a CRM dashboard, the tutorial might say 'The left sidebar shows navigation; click "Deals" to see your pipeline. The green button at the top right creates a new deal...'

11. Multimodal Prompt for Code Generation: Describing an Image of a UI to Get Code

When to use: You have a picture of a user interface (from a sketch, a design mockup, or even a competitor's app) and want to generate HTML/CSS or React code that approximates it.

Why it works: Multimodal models can analyze the layout and visual elements and translate them into code structure.

The prompt:

Analyze the attached UI mockup image and generate the following:
1. The HTML structure (semantic elements) with class names
2. The CSS styling (flexbox/grid, colors, spacing) that matches the visual design
3. Any JavaScript functionality that is implied (e.g., buttons, forms) as a separate block
4. A list of accessibility improvements

Target framework: [React/Vue/vanilla]. Output code in separate code blocks.

Example: From a mockup of a simple landing page, the AI might generate a header, hero section, and footer with corresponding CSS, and suggest adding alt text to images.

12. Video Script from a Series of Images: Building a Slideshow Video

When to use: You have a set of images (e.g., product photos, event pictures) and want to create a video slideshow with narration and music suggestions.

Why it works: The prompt uses the images as story beats and generates a script that sequences them into a narrative.

The prompt:

I have attached [N] images that I want to turn into a 30-60 second slideshow video. Generate:
1. A sequence order for the images with a reason for each transition
2. A narration script (50-100 words) that tells a story across the images
3. Suggested background music (genre, mood, tempo)
4. Text overlays for each image (max 5 words)
5. Transition effects (e.g., fade, zoom) for each change

Target audience: [describe].

Example: For 5 images from a team offsite, the AI might order them as 'team photo → working session → fun moment → group dinner → logo', with narration about teamwork and success.

13. Visual Brainstorming: Generating Ideas from an Image

When to use: You have an image that inspires you but need creative ideas for a project (e.g., a blog post, a marketing campaign, a product feature).

Why it works: The prompt stimulates divergent thinking by connecting the visual context to your domain.

The prompt:

Look at the attached image and brainstorm 10 creative ideas for [your project/goal]. For each idea, provide:
- A one-sentence description
- How it connects to the image (e.g., color, subject, mood)
- A potential implementation step

Be wild and unconventional; we'll filter later.

Example: For a photo of a mountain landscape, ideas for a blog post might include 'The psychology of peak experiences' or 'Lessons from mountaineering for startup founders'.

14. Multimodal Data Extraction: Creating a Structured Table from a Document Image

When to use: You have a scanned document, a table image, or a receipt, and need structured data (like a CSV) for analysis or bookkeeping.

Why it works: The prompt leverages the model's ability to read text and understand table structure.

The prompt:

Extract the data from the attached image into a structured table. Use the following format:

| Column 1 | Column 2 | Column 3 |
|----------|----------|----------|
| data     | data     | data     |

Preserve all text exactly as it appears, including currency symbols and units. If a cell is empty, leave it blank. After the table, provide a brief note about any ambiguities (e.g., 'The date format is unclear').

Example: From a photo of an invoice, the AI extracts 'Invoice #1234, Date: 08/20/2026, Items: ...' into a Markdown table.

15. Brand Consistency: Generating a Visual Style Guide from an Existing Image

When to use: You want to establish or enforce a brand style across content. The AI analyzes an image (e.g., your logo or a brand asset) and produces a style guide.

Why it works: The prompt extracts design tokens (colors, fonts, imagery style) and formalizes them.

The prompt:

Analyze the attached image as a brand asset. Generate a visual style guide that includes:
1. Primary and secondary color palettes with hex codes
2. Recommended typography (font families, weights, sizes) that matches the brand's tone
3. Imagery style (e.g., photography, illustration, flat design) with examples
4. Logo usage rules (clear space, minimum size, do's and don'ts)
5. A sample text-to-image prompt that would generate an image consistent with this brand

Example: For a tech startup logo, the style guide might specify 'Primary color: #00A3E0 (cyan)', 'Font: Inter, bold for headings', and 'Imagery: minimal, futuristic, using 3D renders'.


Final Thoughts

The prompts above are starting points, not magic bullets. The key is to be explicit about the relationship between the modalities—tell the AI what to do with the image, how to use the text, and what output format you need. With practice, you'll develop your own variations for specific tools like OpenAI's GPT-4 with vision, Google's Gemini, or open-source models like LLaVA.

Remember to always verify the output, especially for technical or factual content. Multimodal AI is powerful, but it's still a tool—your expertise is the final filter. Start with one prompt from this list, adapt it to your workflow, and you'll see how seamlessly text, images, and video can work together in your projects.

If you're ready to take it further, explore how AI can automate entire content pipelines, from idea to published video. The future is multimodal, and you're already one step ahead.

← All posts

Comments