Back to Nodes
MMAudio V2

MMAudio V2

Official

Generate audio that follows video, image, or text context.

Nodespell AI
AI / Audio / Mmaudio

Generate audio that follows video, image, or text context.

MMAudio V2 is useful when a silent clip or still needs matching sound rather than music or speech. It accepts prompt plus video or image input, with duration, steps, guidance strength, seed, and negative prompt controls.

Use it for ambience, action sounds, and synchronized sound design drafts. For isolated sound effects, a text-to-SFX node is usually more direct.

Model Examples (3)

Example Index01 / 03
Example 01

Floating monastery ambience

Video-to-audio environmental sound for a fantasy location clip.

Source Inputs02
Prompt

Wind across rope bridges, distant monastery bells, fabric flutter, soft wooden creaks, high-altitude air, no music, no voices.

Video
Example input
Parameters05
Prompt
Wind across rope bridges, distant monastery bells, fabric flutter, soft wooden creaks, high-altitude air, no music, no voices.
Duration
8
Num Steps
25
Cfg Strength
4.5
Negative Prompt
music, dialogue, singing
video-to-audioambience
Response
Inputs (3)

prompt

String

Text prompt for generated audio

Multi InputMin: 0Max: 100

video

String

Optional video file for video-to-audio generation

Min: 0Max: 100

image

String

Optional image file for image-to-audio generation (experimental)

Min: 0Max: 100
Parameters (6)

Seed

Number

Random seed. Use -1 or leave blank to randomize the seed

Prompt

String

Text prompt for generated audio

Default:

Duration

Number

Duration of output in seconds

Default: 8

num_steps

Number

Number of inference steps

Default: 25

cfg_strength

Number

Guidance strength (CFG)

Default: 4.5

negative_prompt

String

Negative prompt to avoid certain sounds

Default: music
Outputs (1)

response

Inferred

response

Nodespell Team

Creator profile

Type

Node

Status

Official

Package

Nodespell AI

Category

AI / Audio / Mmaudio

Input

VideoText

Output

Audio

Keywords

Video EditSound Effect GenerationAudio EnhancementMultimodal GenerationConditional GenerationLength Control
Use in Workflow