black-forest-labs/FLUX-2-dev
Brand-new Flux2 Dev introduces a faster, more modular architecture for next-generation image generation pipelines. It delivers improved performance, cleaner control APIs, and a significantly more flexible development workflow for custom inference setups.

Try the model with a sample request using the same API shape shown in the reference page.
Text to convert to speech
Please upload an image file
Lyrics to sing, with optional [verse]/[chorus] structure tags. Empty means the model writes its own lyrics (or, for instrumental requests, none).
Text-based edit instruction (e.g., 'make the sky blue', 'add a cat'). This parameter serves as the text prompt.. (Default: empty)
Please upload an image file
List of coordinates to define the area to mask.
Input image URLs for image editing task. Provide a list of image URLs.
You can add more items with the button on the right
Please upload an audio file
You can add more items with the button on the right
Required. Publicly accessible URL of the input video.
Required. A URL with mask video defining the area to erase.
URL of the input video.
Video orientation.
Generate a lower-quality draft at reduced cost. 720p draft is $0.005/s vs $0.02/s; 1080p draft is $0.01/s vs $0.04/s.
URL of initial video clip for video continuation (MP4/MOV, 2-10s, max 100MB). (Default: empty)
Exaggeration factor for the speech (Default: 0, 0 ≤ exaggeration ≤ 1)
CFG factor for the speech (Default: 0, 0 ≤ cfg ≤ 1)
Output format (only pcm supported). (Default: pcm)
Whether the output video includes audio. true (default): Includes audio; false: No audio track.
Optional URL of the image to use as the last frame of the video
Select the desired voice for the speech output.
Minimum number of tokens for the output. (Default: empty)
Whether to preserve audio from the input when doing image-to-video generation.
Sample rate for the output audio (Default: 24000)
List of reference media. Must contain at least one reference_image or reference_video. The total of reference_image + reference_video must be <= 5. At most one first_frame is allowed.
Temperature of the generation (Default: 0.9)
Maximum number of tokens for the generation (Default: 2000, 0 < max_tokens ≤ 4096)
Repetition penalty for the generation (Default: 1.1, 0 ≤ repetition_penalty)
Whether to return word-level timestamps
Maximum number of tokens for the output. (Default: empty)
Transcript of the given speaker audio. If not provided then the speaker audio will be used as is.. (Default: empty)
Maximum audio length in milliseconds (Default: 10000)
Please upload an audio file
Target tempo in beats per minute. Empty means the model chooses. (Default: empty, 30 ≤ bpm ≤ 300)
If set, the generated image will match the size of the input image at this index (0-based). (Default: empty, 0 ≤ match_image_size < 4)
temperature to use for sampling (Default: 0)
optional text to provide as a prompt for the first window.. (Default: empty)
chunk level, either 'segment' or 'word'
chunk length in seconds to split audio (Default: 30, 1 ≤ chunk_length_s ≤ 30)
Language code for multilingual model (e.g., 'en', 'fr', 'zh'). Only used with chatterbox-multilingual.. (Default: empty)
Temperature for the speech (Default: 0.8, 0 ≤ temperature ≤ 2)
Speed of the speech (Default: empty, 0.25 ≤ speed ≤ 4)
Optional natural language instruction describing the desired speaking style, emotion, role-play, dialect, etc. Sent in the `user` role; not synthesized.. (Default: empty)
Whether to stream the output.
Absolute maximum number of tokens for the output. (Default: empty)
Speaking rate of the speech (Default: 1, 0.5 ≤ speaking_rate ≤ 1.5)
Let the planning model arrange structure and phrasing before generation. Higher quality but slower; disable it for faster results.
Generate instrumental music with no vocals, regardless of lyrics.
whether to use a fixed camera angle
Select the desired voice for the speech output. You can select multiple to combine and mix voices.
Select the desired voice for the speech output. You can select multiple to combine and mix voices.
Select the desired format for the speech output. Supported formats include mp3, opus, flac, wav, and pcm.
task to perform
The safety checker is always enabled in Playground. It can only be disabled by setting false through the API.
Output audio format. mp3 is small and universally playable; flac and wav are lossless.
Colbert is a system of contextualized vectors where every token in the input is represented by its own BERT-derived embedding.
Language for auto-written lyrics. Ignored when you supply lyrics or for instrumental tracks.
Whether to perform upsampling on the prompt. If active, automatically modifies the prompt for more creative generation
Output format for the generated image
Please upload an image file
Please upload an image file
Please upload an image file
Please upload an image file
Aspect ratio of the generated video. Default 16:9. Ignored if a first_frame is provided (the first frame's ratio is used).
Sparse is a collection of high-dimensional vectors where each word in the input is assigned a lexical weight, with most values being zero.
Please upload an image file
Script for the person to say when no audio is uploaded.. (Default: empty)
The desired output container and codec, e.g., 'mp4/h264'.
The scale for text guidance during image generation. Higher values increase adherence to the prompt. (Default: empty, 0 ≤ text_guidance_scale ≤ 10)
The number of diffusion steps to use for image generation. If not provided, the default number of steps for the model will be used. (Default: empty, 0 ≤ steps_num)
undefined (Default: 5, 1 ≤ scale ≤ 5)
Integer scale factor for upscaling
Structured JSON prompt in string format.. (Default: empty)
Output resolution (e.g., 1024x1024, 1920x1080).. (Default: empty)
Generation speed preset: 'standard', 'fast', or 'hq'.
A string containing the structured edit instruction in JSON format. Use this instead of instruction for precise, programmatic control.. (Default: empty)
Tolerance level for input and output moderation. Between 0 and 6, 0 being most strict, 6 being least strict (Default: 2, 0 ≤ safety_tolerance ≤ 6)
The aspect ratio of the generated images.. (Default: 1:1)
Question about the provided image
Please upload an image file
Output language for generated speech. Defaults to 'English (US)'.
Please upload an image file
You can add more items with the button on the right
A configurable parameter. Defaults to true in the Playground.
Speaking style, tone, pacing or emotion instructions for the generated speech.. (Default: empty)
Length of the track in seconds (30-300). Leave empty to let the model choose. (Default: empty, 30 ≤ duration ≤ 300)
Whether to enhance the prompt with additional context
Please upload an image file
number of denoising steps (Default: 28, 1 ≤ num_inference_steps ≤ 100)
The frame number to apply key points to. (Default: 0)
Value between -10 and 10 (Default: 0, -10 ≤ fractality ≤ 10)
Optional prompt for the video.. (Default: empty)
Video output format.
If true, the request is rejected immediately with HTTP 429 when the model has no spare capacity, instead of waiting in the queue. Opt-in; the default (false) keeps standard queueing behavior.
Disable safety filter for prompts and input image. Defaults to true.
Alternative to scale_factor, value between 0.001 and 200 (Default: empty, 0.001 ≤ target_megapixels ≤ 200)
Please upload an image file
Value between -10 and 10 (Default: 0, -10 ≤ creativity ≤ 10)
If true, videos longer than 5 seconds are trimmed to the first 5 seconds.
When true, skip the prompt upsampler and pass the raw user prompt to the video model.
Number of samples to generate, must be between 1 and 4 (Default: empty, 1 ≤ sample_count ≤ 4)
You can add more items with the button on the right
Output video quality. 1080p supports up to 15 seconds.
Disable safety checker for generated images
Whether to allow adult faces are generated in the video
Video duration in seconds.
image width in px (Default: 1024, 128 ≤ width ≤ 1920)
image height in px (Default: 1024, 128 ≤ height ≤ 1920)
Whether the generated video includes audio synchronized with the visuals
Predefined string only - one of the enum values below. Hex values are not supported
Whether to enable prompt rewriting using LLM
You can add more items with the button on the right
You can add more items with the button on the right
Whether the model is allowed to generate multiple clips
URL of an audio file to condition the video on. When provided, the duration parameter is ignored.. (Default: empty)
Resolution of the generated video in width*height format. Default is 1920*1080
whether the generated video includes audio synchronized with the visuals
Whether to add AI Generated watermark. Default is false
The service tier used for processing the request. When set to 'priority', the request will be processed with higher priority (only applies to models that support it).
text prompt
classifier-free guidance, higher means follow prompt more closely (Default: 2.5, 0 ≤ guidance_scale ≤ 20)
Frames per second of the generated video.
Size of the generated image. (Default: 2K)
URL of audio file for synchronized audio and video. (Default: empty)
First-frame image: URL or base64-encoded image data. Used for first-frame and first+last-frame video generation modes. (Default: empty)
Whether to enable prompt rewriting for better quality. Default is true
Whether to enable thinking mode for better editing precision
Bounding boxes for interactive precise editing, one list per input image. Each bounding box is [x1, y1, x2, y2] in absolute pixels. Use empty list [] for images without bounding boxes. Each image supports up to 2 bounding boxes.
Number of steps for the image generation process (Default: 25, 1 ≤ steps ≤ 50)
language that the audio is in; uses detected language if None; use two letter language code (ISO 639-1) (e.g. en, de, ja)
Last-frame image: URL or base64-encoded image data. Used together with first_frame or first_clip. (Default: empty)
Acceleration level to use. The more acceleration, the faster the generation, but with lower quality. The recommended value is 'none'. Default value: "none"
Duration of the output video
Please upload an image file
A key to identify prompt cache for reuse across requests. If provided, the prompt will be cached and can be reused in subsequent requests with the same key.. (Default: empty)
Interval is a setting that increases the variance in possible outputs letting the model be a tad more dynamic in what outputs it may produce in terms of composition, color, detail, and prompt interpretation. Setting this value low will ensure strong prompt following with more consistent outputs, setting it higher will produce more dynamic or varied outputs (Default: 2, 1 ≤ interval ≤ 4)
Please upload an image file
whether to return the last frame image of the generated video
number of images to generate (Default: 1, 1 ≤ num_images ≤ 4)
URL of audio file used as a driving source for lip-sync and timing (WAV/MP3, 2-30s, max 15MB). Only valid with first_frame. (Default: empty)
Tweak the overall style and tone of the conversation by giving some 'master' instructions. (Default: Be a helpful assistant)
Enter a JSON array
Aspect ratio for the image.
Dense is a set of low-dimensional vectors where every token in the input is represented by a fully populated embedding derived from a neural model.
A list of input video URLs. Furthermore, the total length of the three videos must not exceed 30 seconds.
whether to normalize the computed embeddings
Whether to preserve the original audio track.
Sample from the set of tokens with highest probability such that sum of probabilies is higher than p. Lower values focus on the most probable tokens.Higher values sample more low-probability tokens (Default: 0.9, 0 < top_p ≤ 1)
Repetition Penalty for the speech (Default: 1.2, 0 ≤ repetition_penalty ≤ 5)
random seed, empty means random (Default: empty, 0 ≤ seed < 18446744073709552000)
Top K for the speech (Default: 1000, 0 ≤ top_k ≤ 1000)
Style preset: default, portrait, or anime
The number of dimensions in the embedding. If not provided, the model's default will be used.If provided bigger than model's default, the embedding will be padded with zeros. (Default: empty, 32 ≤ dimensions ≤ 8192)
maximum length of the newly generated generated text.If explicitly set to None it will be the model's max context length minus input length or 65536, whichever is smaller (Default: empty, 1 ≤ max_new_tokens ≤ 1000000)
A unique identifier representing your end-user, which can help monitor and detect abuse. Avoid sending us any identifying information. We recommend hashing user identifiers.. (Default: empty)
Whether to generate AI audio synchronized with the video.
Min P for the speech (Default: 0, 0 ≤ min_p ≤ 1)
A list of input audio URLs. Furthermore, the total length of the three audios must not exceed 30 seconds.
Upload the image you want to use as the last frame of the video.
Whether to return the last frame of the video. When draft=true, this parameter cannot be set to true.
Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics. (Default: 0, -2 ≤ presence_penalty ≤ 2)
Positive values penalize new tokens based on how many times they appear in the text so far, increasing the model's likelihood to talk about new topics. (Default: 0, -2 ≤ frequency_penalty ≤ 2)
How to format the response
Enable online search?
temperature to use for sampling. 0 means the output is deterministic. Values greater than 1 encourage more diversity (Default: 0.7, 0 ≤ temperature ≤ 1)
Generate multi-camera video.This field must be set to false when both start and end frames are provided.
The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt.
First and last frame image URLs. Required when elements are referenced in the prompt (using @element_name syntax). When multi_shots is false: if length is 2, index 0 is the first frame and index 1 is the last frame; if length is 1, the array item serves as the first frame. When multi_shots is true: only the first frame is supported.
Upload the image you want to use as the first frame of the video.
Create multiple shots of the same video (up to 5 shots). The total duration of all shots cannot exceed 15 seconds
Optional video input. Only 1 video is allowed and it uses 2 image slots.
Whether to use the model's prompt optimizer
Styles, genres, moods, or instruments to steer the song away from — a compositional nudge, not an audio filter. Ignored when the model writes its own lyrics.. (Default: empty)
Smart Storyboarding
Adjust settings and run the demo to see the model output generated here in real-time.