We’re excited to launch Muse Image and preview Muse Video, the first media generation models developed by Meta Superintelligence Labs.

Muse Image is our most advanced image generation model yet: it follows instructions faithfully, edits with precision, composes from multiple references, and draws on Instagram for social context. It also brings agentic tool use capabilities and integrates with Muse Spark. Muse Video, built on the same pretraining base, delivers exceptional visual fidelity with native audio support.

Muse Image is available today across the Meta AI app and on meta.ai, Instagram Stories in the US, and WhatsApp in limited countries, and is coming soon to Facebook. Muse Video is coming soon to creators and Meta AI.

Muse Image: Agentic Image Generation

Instead of directly mapping prompts to images, Muse Image operates as an agent: it invokes search and coding tools to improve accuracy, self-refines its own generations, and improves through scaling test-time compute. Muse Image also integrates with Muse Spark, allowing the two models to share tools and plan jointly for powerful agentic media generation.

Tool Use

Muse Image learns to write and execute code that produces accurate plots and QR codes during reinforcement learning. It also learns to search the web to ground generated images in factual and real-time information.

Self-Refinement

Muse Image reflects on and improves upon its own work within its chain of thought. This behavior emerged during RL training because self-refinement produced better images and therefore higher reward.

Test-Time Compute Scaling

Like language models, Muse Image improves the more it thinks at inference time. Increasing reasoning strength improves human-preference Elo scores and shows an approximately log-linear scaling relationship.

Image Editing and Multi-Reference Composition

Muse Image edits images with precision and can compose elements from many input reference images, supporting interleaving text and images inline in prompts.

Muse Image holds the No. 2 spot on Arena for text-to-image, single-image editing, and multi-image editing as measured by human preference Elo rankings at the time of writing.

Previewing Muse Video

Alongside the release of Muse Image, we’re sharing an early preview of Muse Video. It offers competitive performance in prompt adherence, visual fidelity, and temporal consistency with native audio support. We’re investing in areas with current performance gaps, such as audio-video synchronization and physically accurate fast motion.

On Arena, Muse Video ranks No. 3 in human-preference Elo for text-to-video at the time of writing.

Content Seal

Muse Image includes Content Seal, an invisible watermarking system. Images created by Muse Image in the Meta AI app and on meta.ai carry a hidden provenance signal that stays intact even when cropped, compressed, resized, or screenshotted. Meta plans to extend Content Seal to video soon and is previewing a detection tool.

Meta Ecosystem Integration

Muse Image connects deeply with the Meta ecosystem. Combined with social tools in Meta AI, users can create images with friends and reimagine their Instagram photos. Marketing assets for small businesses can use @-mention of public Instagram accounts for personalized presets.