Baidu's ERNIE Image model for high-quality multilingual text-to-image generation with built-in prompt expansion.
Added Apr 10, 2026
Approx. Price
$0.010 per image
Model Type
text-to-image
Settings
Generation controls available for this model.
Images Per Run
Up to 4
Output images
Input Images
N/A
No reference/edit image input
Output Sizes
Enable Safety Checker
Default
No
Guidance Scale
Default
5
How strongly to follow the prompt (1-20).
Inference Steps
Default
50
Number of denoising steps (1-100).
Negative Prompt
Default
N/A
Describe what to avoid in the image.
Number of Images
Default
1
Output Format
Default
jpeg
Options (2)
JPEG, PNG
Image output format.
Prompt Expansion
Default
Yes
Enhance the prompt automatically for richer outputs.
Resolution
Default
1024x1024
Options (11)
512x512 (Square), 1024x1024 (Square HD), 768x1024 (Portrait (3:4)), 576x1024 (Portrait (9:16)) +7 more
Seed
Default
N/A
Random seed for reproducible outputs.
Benchmarks
Benchmarks
Human preference benchmarks sourced from Artificial Analysis.
Text to Image
#93 / 165
ELO
914.0
Appearances
12,016
95% CI
-8/8
Release Date 2026-04 · Matched as ERNIE Image
Artificial Analysis APIExamples
Loading examples…
Related image models
Compare ERNIE Image with similar models from the same provider or model family.
ERNIE Image Turbo
ernie-image/turboSpeed-optimized ERNIE Image variant with lower default step count for faster multilingual text-to-image generation.
VOSR2 Image Upscaler
wavespeed-ai/vosr2/imageRestore and upscale photos to 2K or 4K in a single pass while preserving composition, color, faces, and fine detail.
Seedream 5.0 Flash
bytedance/seedream-v5.0-flashFast Seedream 5.0 image generation. Upload images to edit them.
Seedream 5.0 Flash Layerize
bytedance/seedream/v5/flash/layerizeSeparate an image into a base image and editable layers.
Recraft V4.1 Flash
recraft-ai/recraft-v4.1-flash/text-to-imageRecraft V4.1 Flash quickly creates design-focused raster images for exploring prompts and compositions.
Qwen Image 2.1 Edit
wavespeed-ai/qwen-image-2.1/editEdit from up to ten reference images with natural-language instructions, strong subject preservation, flexible framing, and native output up to 2K.