Skip to main content
Nano Banana Pro 的 API 参考。Nano Banana Pro(Gemini 3 Pro Image)是 Google Nano Banana 图像生成系列的 Pro 层级,专为复杂场景和清晰可读的文字而设计。

快速开始

在你的 Comfy 工作区中创建一个密钥,并将其导出为 COMFY_API_KEY。Python 和 TypeScript 代码片段使用 Comfy SDK(pip install comfy-sdk 和 npm install @comfyorg/sdk);cURL 代码片段则是通过原始 HTTP 进行的相同调用。 模型 ID: vertexai/gemini-3-pro-image 端点: POST https://api.comfy.org/v2/models/vertexai/gemini-3-pro-image
将同样的请求体发送到 POST https://api.comfy.org/v2/models/vertexai/gemini-3-pro-image/requests。运行一旦被受理,Router 会立即返回 201 和 request_id;结果就绪后,就可以从当前进程或另一个进程中收集它。队列投递 会逐步介绍状态查询、取消和收集。

服务提供商

除非请求指定了其他提供商,否则该模型由 Comfy Router 直接提供服务。以下提供商也在同一端点和相同的模型 ID 下提供该模型,可通过 model_provider 查询参数选择。
  • Comfy(默认):POST https://api.comfy.org/v2/models/vertexai/gemini-3-pro-image
  • fal,以 fal/fal-nano-banana-pro 提供:POST https://api.comfy.org/v2/models/vertexai/gemini-3-pro-image?model_provider=fal
  • Runware,以 runware/runware-nano-banana-pro 提供:POST https://api.comfy.org/v2/models/vertexai/gemini-3-pro-image?model_provider=runware
  • WaveSpeed,以 wavespeed/wavespeed-nano-banana-pro 提供:POST https://api.comfy.org/v2/models/vertexai/gemini-3-pro-image?model_provider=wavespeed
strict_mode 默认为 false,因此 Router 会将本页记录的原始请求体转换为提供商自己的 schema,并将响应转换回来。请参阅 API 参考中的 model_provider、strict_mode 和 fallback_provider,以及 服务提供商 了解所有以此方式路由的模型。

Schema

输入

object[]
必填
与模型当前对话的内容。对于单轮查询,这是一个单独的实例。对于多轮查询,这是一个重复字段,包含对话历史和最新请求。
object[]
必填
object
基于 URI 的数据。
string
URI
string
data 或 fileUri 字段中所指定文件的媒体类型。可接受的值包括以下内容。对于 gemini-2.0-flash-lite 和 gemini-2.0-flash,音频文件的最大长度为 8.4 小时,视频文件(不含音频)的最大长度为一小时。有关更多信息,请参阅 Gemini 音频和视频相关要求。文本文件必须采用 UTF-8 编码。文本文件的内容计入 token 上限。图像分辨率无限制。可能的值:application/pdf、audio/mpeg、audio/mp3、audio/wav、image/png、image/jpeg、image/webp、text/plain、video/mov、video/mpeg、video/mp4、video/mpg、video/avi、video/wmv、video/mpegps、video/flv、image/heic、image/heif、audio/flac、video/webm
object
以原始字节形式提供的行内数据。对于 gemini-2.0-flash-lite 和 gemini-2.0-flash,使用 inlineData 最多可以指定 3000 张图像。
string (byte)
要以内联方式包含在提示中的图像、PDF 或视频的 base64 编码。以内联方式包含媒体时,还必须指定数据的媒体类型(mimeType)。大小限制:20MB格式:byte
string
data 或 fileUri 字段中所指定文件的媒体类型。可接受的值包括以下内容。对于 gemini-2.0-flash-lite 和 gemini-2.0-flash,音频文件的最大长度为 8.4 小时,视频文件(不含音频)的最大长度为一小时。有关更多信息,请参阅 Gemini 音频和视频相关要求。文本文件必须采用 UTF-8 编码。文本文件的内容计入 token 上限。图像分辨率无限制。可能的值:application/pdf、audio/mpeg、audio/mp3、audio/wav、image/png、image/jpeg、image/webp、text/plain、video/mov、video/mpeg、video/mp4、video/mpg、video/avi、video/wmv、video/mpegps、video/flv、image/heic、image/heif、audio/flac、video/webm
string
模型如何读取此部分的视频。设置为 “AGENTIC” 可让模型自行决定要检查哪些片段,而不是采用固定帧率采样。省略则使用默认的固定帧率采样。受支持于 gemini-3.7-flash 及更新的 Flash 模型。
string
文本提示或代码片段。
boolean
表示此部分是模型的思考/推理步骤。
string
可能的值:user、model
object
生成的采样、长度和输出设置。每个字段都是可选的:下面声明了 default 的字段在省略时会应用该默认值,其余字段则回退到模型自身的行为。
object
图像生成的配置
string
生成图像的宽高比
object
可选。生成图像的图像输出格式。
integer
可选。输出图像的压缩质量。
string
可选。在 Vertex AI 路径上,输出图像应保存为的图像格式,这些路径包括:由 Comfy 自有凭证提供服务的请求,以及使用 GCP 服务账号进行身份验证的 BYOK 请求。在这些路径上可接受的值为 image/png 和 image/jpeg,匹配时不区分大小写,并在请求转发之前规范化为小写;任何其他值都会被拒绝,并返回一个指明此字段的 400 错误。省略时默认为 image/png。使用 Google AI Studio API 密钥进行身份验证的 BYOK 请求是个例外:该上游没有此属性,只要它存在就会拒绝整个调用,因此该字段会从请求中移除,而不是被采纳或被拒绝,输出格式则由 AI Studio 自行决定。在所有路径上,都应从你收到的响应部分中读回媒体类型(inlineData.mimeType,如果设置了 uploadImagesToStorage,则为 fileData.mimeType),而不要假设你发送的值。
string
可选。指定生成图像的尺寸。支持的值为 1K、2K、4K。如果未指定,模型将使用默认值 1K。
integer
响应中可以生成的最大 token 数。一个 token 大约相当于 4 个字符。100 个 token 大约对应 60-80 个单词。范围:16 到 65536
`TEXT`, `IMAGE`[]
integer
When seed is fixed to a specific value, the model makes a best effort to provide the same response for repeated requests. Deterministic output isn’t guaranteed. Also, changing the model or parameter settings, such as the temperature, can cause variations in the response even when you use the same seed value. By default, a random seed value is used. Available for the following models:, gemini-2.5-flash, gemini-2.5-pro, gemini-2.5-flash-preview-04-1, gemini-2.5-pro-preview-05-0, gemini-2.0-flash-lite-00, gemini-2.0-flash-001
string[]
number
默认值:"1"
The temperature is used for sampling during response generation, which occurs when topP and topK are applied. Temperature controls the degree of randomness in token selection. Lower temperatures are good for prompts that require a less open-ended or creative response, while higher temperatures can lead to more diverse or creative results. A temperature of 0 means that the highest probability tokens are always selected. In this case, responses for a given prompt are mostly deterministic, but a small amount of variation is still possible. If the model returns a response that’s too generic, too short, or the model gives a fallback response, try increasing the temperatureRange: 0 to 2Format: float
object
Optional. Configuration for thinking features. Thinking is a process where the model breaks down a complex task into smaller steps to generate a higher-quality response.
boolean
Optional. If true, the model will include its thoughts in the response.
integer
Optional. The token budget for the model’s thinking process. The model will make a best effort to stay within this budget.
string
Optional. The thinking level for the model.Possible values: THINKING_LEVEL_UNSPECIFIED, LOW, MEDIUM, HIGH, MINIMAL
integer
默认值:"40"
Top-K changes how the model selects tokens for output. A top-K of 1 means the next selected token is the most probable among all tokens in the model’s vocabulary. A top-K of 3 means that the next token is selected from among the 3 most probable tokens by using temperature.Range: 1 to …
number
默认值:"0.95"
If specified, nucleus sampling is used. Top-P changes how the model selects tokens for output. Tokens are selected from the most (see top-K) to least probable until the sum of their probabilities equals the top-P value. For example, if tokens A, B, and C have a probability of 0.3, 0.2, and 0.1 and the top-P value is 0.5, then the model will select either A or B as the next token by using temperature and excludes C as a candidate. Specify a lower value for less random responses and a higher value for more random responses.Range: 0 to 1Format: float
object[]
Per request settings for blocking unsafe content. Enforced on GenerateContentResponse.candidates.
string
必填
Possible values: HARM_CATEGORY_SEXUALLY_EXPLICIT, HARM_CATEGORY_HATE_SPEECH, HARM_CATEGORY_HARASSMENT, HARM_CATEGORY_DANGEROUS_CONTENT
string
必填
Possible values: OFF, BLOCK_NONE, BLOCK_LOW_AND_ABOVE, BLOCK_MEDIUM_AND_ABOVE, BLOCK_ONLY_HIGH
object
Instructions for the model to steer it toward better performance. For example, “Answer as concisely as possible” or “Don’t use technical terms in your response”. The text strings count toward the token limit. The role field of systemInstruction is ignored and doesn’t affect the performance of the model. Note: Only text should be used in parts and content in each part should be in a separate paragraph.
object[]
必填
A list of ordered parts that make up a single message. Different parts may have different IANA MIME types. For limits on the inputs, such as the maximum number of tokens or the number of images, see the model specifications on the Google models page.
string
A text prompt or code snippet.
string
The identity of the entity that creates the message. The following values are supported: user: This indicates that the message is sent by a real person, typically a user-generated message. model: This indicates that the message is generated by the model. The model value is used to insert messages from the model into the conversation during multi-turn conversations. For non-multi-turn conversations, this field can be left blank or unset.Possible values: user, model
object[]
A piece of code that enables the system to interact with external systems to perform an action, or set of actions, outside of knowledge and scope of the model. See Function calling.
object[]
string
string
必填
object
函数参数的 JSON schema
boolean
如果设为 true,已生成图像将上传到云端存储,并以签名 URL 的形式返回,而不是内联 base64 数据。这些 URL 会在 24 小时后过期。
object
对于视频输入,视频的开始和结束偏移量,采用 Duration 格式。例如,要指定从 1:00 开始的 10 秒片段,请设置 “startOffset”: { “seconds”: 60 } 和 “endOffset”: { “seconds”: 70 }。仅当视频数据以 inlineData 或 fileData 形式提供时,才应指定该元数据。
object
表示视频时间轴位置的时长偏移量。
integer
以纳秒为分辨率的有符号秒的小数部分。带小数的负秒值仍必须具有非负的 nanos 值。范围:0 到 999999999
integer
时间段的有符号秒数。必须在 -315,576,000,000 到 +315,576,000,000 之间(含两端)。范围:-315576000000 到 315576000000
object
表示视频时间轴位置的时长偏移量。
integer
以纳秒为分辨率的有符号秒的小数部分。带小数的负秒值仍必须具有非负的 nanos 值。范围:0 到 999999999
integer
时间段的有符号秒数。必须在 -315,576,000,000 到 +315,576,000,000 之间(含两端)。范围:-315576000000 到 315576000000
根据 Router 在 GET /v2/models/vertexai/gemini-3-pro-image/openapi.json 提供的 schema 生成,这也是它在请求到达提供商之前用于校验调用的同一份文档。

输出

object[]
object
object[]
string[]
integer
string
string (date)
格式:date
integer
string
string
object
与模型进行当前对话的内容。对于单轮查询,这是单个实例。对于多轮查询,这是包含对话历史和最新请求的重复字段。
object[]
必填
object
基于 URI 的数据。
string
URI
string
data 或 fileUri 字段中指定的文件的媒体类型。可接受的值包括以下内容。对于 gemini-2.0-flash-lite 和 gemini-2.0-flash,音频文件的最大长度为 8.4 小时,视频文件(不含音频)的最大长度为一小时。有关更多信息,请参阅 Gemini 音频和视频要求。文本文件必须采用 UTF-8 编码。文本文件的内容计入 token 限制。图像分辨率无限制。可能的值:application/pdf、audio/mpeg、audio/mp3、audio/wav、image/png、image/jpeg、image/webp、text/plain、video/mov、video/mpeg、video/mp4、video/mpg、video/avi、video/wmv、video/mpegps、video/flv、image/heic、image/heif、audio/flac、video/webm
object
以原始字节表示的内联数据。对于 gemini-2.0-flash-lite 和 gemini-2.0-flash,使用 inlineData 最多可以指定 3000 个图像。
string (byte)
要在提示中内联包含的图像、PDF 或视频的 base64 编码。内联包含媒体时,还必须指定数据的媒体类型(mimeType)。大小限制:20MB格式:byte
string
data 或 fileUri 字段中指定的文件的媒体类型。可接受的值包括以下内容。对于 gemini-2.0-flash-lite 和 gemini-2.0-flash,音频文件的最大长度为 8.4 小时,视频文件(不含音频)的最大长度为一小时。有关更多信息,请参阅 Gemini 音频和视频要求。文本文件必须采用 UTF-8 编码。文本文件的内容计入 token 限制。图像分辨率无限制。可能的值:application/pdf、audio/mpeg、audio/mp3、audio/wav、image/png、image/jpeg、image/webp、text/plain、video/mov、video/mpeg、video/mp4、video/mpg、video/avi、video/wmv、video/mpegps、video/flv、image/heic、image/heif、audio/flac、video/webm
string
模型如何读取此部分的视频。将 “AGENTIC” 设置为让模型决定要检查哪些片段,而不是按固定帧率采样。省略则使用默认的固定帧率采样。在 gemini-3.7-flash 及更新的 Flash 模型中受支持。
string
文本提示或代码片段。
boolean
表示此部分是模型的思考/推理步骤。
string
可能的值:user、model
string
object[]
string
可能的值:HARM_CATEGORY_SEXUALLY_EXPLICIT、HARM_CATEGORY_HATE_SPEECH、HARM_CATEGORY_HARASSMENT、HARM_CATEGORY_DANGEROUS_CONTENT
string
内容违反指定安全类别的概率可能的值:NEGLIGIBLE、LOW、MEDIUM、HIGH、UNKNOWN
string
响应创建时的时间戳。
string
用于生成响应的模型版本。
object
string
string
object[]
string
Possible values: HARM_CATEGORY_SEXUALLY_EXPLICIT, HARM_CATEGORY_HATE_SPEECH, HARM_CATEGORY_HARASSMENT, HARM_CATEGORY_DANGEROUS_CONTENT
string
内容违反指定安全类别的概率Possible values: NEGLIGIBLE, LOW, MEDIUM, HIGH, UNKNOWN
string
响应的唯一标识符。
object
integer
仅输出。输入中缓存部分(即缓存内容)的 token 数量。
integer
响应中的 token 数量。
object[]
按模态划分的候选 token 明细。
string
输入或输出内容的模态类型。Possible values: MODALITY_UNSPECIFIED, TEXT, IMAGE, VIDEO, AUDIO, DOCUMENT
integer
给定模态的 token 数量。
integer
请求中的 token 数量。设置 cachedContent 时,这仍然是提示的总有效长度,也就是说其中包含缓存内容中的 token 数量。
object[]
按模态划分的提示 token 明细。
string
输入或输出内容的模态类型。Possible values: MODALITY_UNSPECIFIED, TEXT, IMAGE, VIDEO, AUDIO, DOCUMENT
integer
给定模态的 token 数量。
integer
思考输出中存在的 token 数量。
integer
工具使用提示中存在的 token 数量。
object[]
按模态划分的工具使用提示 token 明细。
string
输入或输出内容的模态类型。Possible values: MODALITY_UNSPECIFIED, TEXT, IMAGE, VIDEO, AUDIO, DOCUMENT
integer
给定模态的 token 数量。
integer
token 总数(提示 + 候选)。
string
请求使用的流量类型(例如 PROVISIONED_THROUGHPUT)。

示例

输入

输出

读取 parts

默认情况下,已生成的图像 part 会在 inlineData.data 中包含 base64 字节,并在 inlineData.mimeType 中包含媒体类型。解码这些字节并将其保存到文件。当设置 uploadImagesToStorage: true 时,上传的图像改用 fileData.fileUri 提供签名 URL,用 fileData.mimeType 提供媒体类型。请在这些 URL 过期之前下载这些图像,它们自创建起 24 小时后过期。 fileData 可能出现在并未请求它的响应中。 这两种形状是按 part 而非按响应决定的:当设置了 uploadImagesToStorage: true 时,上传失败的图像会保留为 inlineData,因此同一个响应中可以混用两者。应依据实际存在的键来分支处理,而不是依据你请求了什么。唯一不可能出现的情况恰好相反:当该字段未设置或为 false 时,每个已生成的图像都会以 inlineData 返回,不会产生任何 fileData 图像 part。也可能出现文本 part,并且图像不保证是第一个 part,因此应根据你需要的字段来选择 part,而不是按索引。上面快速入门示例中的 candidates[0].content.parts[0].inlineData.data 路径读取的是本页示例响应中唯一的内联 part;面对真实响应时,应扫描 parts 查找你想要的键,而不是索引位置 0。 thoughtSignature 也是一个 part 字段,而且它很大。 一个 part 可以携带 thoughtSignature,这是模型推理过程的不透明 base64 签名,其存在是为了让该思维过程能在后续请求中被重放。它未出现在上面的已生成 schema 中,该 schema 遵循 Router 发布的请求/响应文档。在实际响应中实测,它每个 part 大约为 1 到 2 MB,与图像本身相当,因此如果你要记录响应日志、通过无服务器函数转发响应或存储响应,就值得为其预留空间或将其显式丢弃。

imageSize 是一个档位,而不是宽度

generationConfig.imageConfig.imageSize 接受 1K、2K 或 4K,每提升一档都会将两条边都翻倍,而不是设定某个宽度。一个 16:9 的请求在 1K 下实测为 1376x768、约 1.35 MB,在 2K 下为 2752x1536、约 5.6 MB:相同的宽高比,四倍的像素,大约四倍的字节数。因此 2K 并不意味着 2048 像素宽的图像,如果把 2K 当作宽度来规划上传路径、响应体限制或存储桶容量,会导致资源不足约四倍。

参考图串联

将上一个已生成的图像直接传回,可以在同一主体上获得新的相机角度,因此对同一场景的一系列拍摄是一连串调用,而不是一个必须一次性描述所有内容的提示词。架构、材质和光照常常能在串联中延续,但模型并不保证这一点;请把连续性视为一个需要检查的可能结果,而不是可以依赖的属性。 请按上一个响应使用的形状把图像传回。对于 inlineData part,从 inlineData.data 取出 base64 字节,并将其原样作为 inlineData 发送,如下所示。对于在 uploadImagesToStorage: true 下以 fileData 返回的 part,在签名 URL 仍然有效时,改为将其 fileData.fileUri 和 fileData.mimeType 作为 fileData part 发送;把字节重新上传为 inlineData 同样可行,且不会过期。然后提出你想要的改动:
每次调用都是独立的,因此请传入你想在其基础上继续构建的图像,而不要依赖对话历史记录。串联时要注意请求体大小:内联图像会计入 Router 的请求体上限,而 2K 图像的字节数大约是 1K 的四倍。

发布前须知

SDK 会生成 Idempotency-Key 并在自动重试中复用它。手动重试时,请复用原始 key。Router 最长可保持连接 10 分钟。 请求失败时,Router 会发送 X-Comfy-Error-Type 响应头说明原因。422 表示 Router 在调用提供商之前就拒绝了输入,413 表示请求体超出了 Router 可接受的大小。已生成的资源请及时下载,因为结果 URL 会过期。 上文任何字段描述中提到的尺寸限制,都是提供商对该字段自身的限定,引自提供商的规范。Router 会对整个请求体另行设置上限,base64 编码的媒体内容也计入其中:参见请求体大小。 本页记录的是通过 Comfy Router 调用的某一个合作伙伴模型。同一个 comfy-sdk / @comfyorg/sdk 包还提供第二个客户端,用于在 Comfy Cloud 上运行完整的 ComfyUI 工作流图:Comfy(api_key=...) / new Comfy({ apiKey }),并带有 client.workflows、client.assets 和 client.jobs。请参阅 Comfy SDKs。

请求头

身份验证、幂等性、请求 ID、错误分类、重试节奏、消费限额。

使用 Router API

模型发现、验证错误、重试与计费。

限制

Router 目前不支持的功能,以及替代方案。