September 28, 2026
Dictation and prompt improvement
The microphone in the generation panel turns speech into prompt text, and the assistant turns a short idea into a prompt for the selected model, taking into account the mode, settings and attached materials.
In the bottom row of the generation panel, next to the launch button, there are two text tools: the microphone and the prompt assistant. On a phone they are in the bottom row of the generation panel. Both buttons are available in any mode where the model has a text field.
The microphone records up to a minute of speech. Recording stops with the Stop and transcribe button or on its own after 60 seconds, then the text is transcribed and appended to the end of the prompt after a space, and what you wrote earlier stays. You do not need to choose a language. The same button is in the Agent chat and in Studio, in the Avatars and Products sections.
The assistant is the button with the Write a prompt for the selected model tooltip. It takes the text in the field as the task and writes a full prompt for the selected model from it. Every explicit fact from your text is kept: the character, the action, lines, text in the shot, song lyrics. The assistant does not add a plot, product or event of its own.
The assistant adds what is missing for execution based on the type of result. For an image: composition, environment, light, palette and materials. For video: action over time, framing, camera behavior and sound. For music: style, tempo, instruments, structure and dynamics.
The assistant sees the selected model and its settings: format, duration, resolution, sound, number of outputs. The values of these settings are not carried into the text; the panel controls manage them. Attached references and slot materials also go into the context. If the model requires references like @Image1, a result without them is not inserted.
The model sets the prompt language. For most models the assistant writes in English, and for Nano Banana, Qwen Image 3 and Seedance it writes in the interface language.
The finished prompt goes into the field, and a Restore previous prompt button appears next to it: it returns the text and settings to their state before the assistant. If the result does not pass the check, because it came out empty, longer than the model's limit or lost a required reference, the field stays as it was, and a hint names the reason.
The assistant also works from the project media library. On a generation card that the filter rejected or that came back empty, there is a Repair prompt with AI action: it rewrites the original prompt and opens the panel in that generation's mode. In the finished image viewer, the Extract prompt button builds a standalone prompt from the content and style of the image.
The Prompt improvement field in project settings sets standing conditions for every assistant run in this project, for example no cartoon look, English text, photorealism.
Separately from the assistant, the panel checks the text against the settings. If the prompt says 15 seconds but a different duration is selected, mentions @image2 that is not attached, or describes video while image mode is selected, that place is underlined, and a hint names the mismatch, for example Duration doesn't match the prompt. Clicking the hint fixes the setting. The hint does not block the launch.
What it gives you
One recording is up to 60 seconds; the transcribed text is appended to the prompt rather than replacing it.
The microphone is in the project generation panel, in the Agent chat and in Studio.
The assistant writes a prompt for the selected model, taking into account its prompt profile, mode, settings and attached materials.
Explicit facts from the idea, such as the character, lines, text in the shot and song lyrics, go into the prompt unchanged.
The Restore previous prompt button restores the text and settings from before the assistant.
A generation that failed because of the prompt can be rewritten with Repair prompt with AI, and a finished image can be turned into a prompt with Extract prompt.
The panel underlines places where the text differs from the mode, duration, format, sound, count or attached files.
When it is useful
An idea by voice from a phone
Tap the microphone in the generation panel, say the scene, correct the transcribed text and start the generation.
One line instead of a full prompt
Write the core of the scene and press the assistant: it adds light, framing, camera and motion for the selected model.
The filter rejected a generation
Repair prompt with AI on the error card rewrites the original prompt and opens it in the panel.
A prompt from a finished image
Extract prompt in the image viewer builds a prompt from its content and style.
Shared project rules
The Prompt improvement field in project settings keeps conditions, such as style or the language of text in the shot, for every assistant run.
FAQ
How long can I speak in one recording?
Up to a minute. After 60 seconds the recording stops and is transcribed on its own. Dictate a long idea in several passes: each part is appended to the end of the prompt.
Do I need to choose a language for dictation?
No, there is no language choice: the recording goes to transcription as is, and the text returns to the prompt field, where you can correct it.
Why do some models have no assistant button?
The assistant appears only when the model's field is a prompt. In voiceover, some avatars and lip-sync, the text goes to the model verbatim, and some models have no prompt field. The microphone stays in these modes.
Why did the assistant write the prompt in English?
The model sets the prompt language. For most models it is English; Nano Banana, Qwen Image 3 and Seedance get the prompt in the interface language.
Can I get my own text back after the assistant?
Yes. The Restore previous prompt button next to the assistant restores the text and settings that were there before it.