Best AI for image generation
Human preference across generation and editing, rated blind, pairwise.
Rumeqo runs these models inside your team rooms. See what each one costs.
| rank | model | vendor | composite | benchmarks | elo | elo | elo |
|---|---|---|---|---|---|---|---|
| 1 | GPT Image 2 (medium)openai/gpt-image-2:medium | OpenAI | 81.6 | 3 of 3 benchmarks | 1380 | 1463 | 1454 |
| 2 | Muse Imagemeta/muse-image | Meta | 79.6 | 3 of 3 benchmarks | 1283 | 1405 | 1401 |
| 3 | Seedream 5.0 Probytedance-seed/seedream-5-0-pro | ByteDance Seed | 78.2 | 3 of 3 benchmarks | 1257 | 1393 | 1414 |
| 4 | Nano Banana 2 (Gemini 3.1 Flash Image)google/gemini-3.1-flash-image | 75.1 | 3 of 3 benchmarks | 1263 | 1385 | 1365 | |
| 5 | Nano Banana Pro (Gemini 3 Pro Image Preview)google/gemini-3-pro-image-preview | 74.1 | 3 of 3 benchmarks | 1232 | 1385 | 1368 | |
| 6 | Nano Banana Pro (Gemini 3 Pro Image) (2K)google/gemini-3-pro-image:2k | 74.0 | 3 of 3 benchmarks | 1246 | 1389 | 1364 | |
| 7 | MAI-Image-2.5microsoft/mai-image-2.5 | Microsoft | 73.5 | 2 of 3 benchmarks | 1256 | 1402 | |
| 8 | Reve 2.0reve/reve-2.0 | Reve | 72.6 | 3 of 3 benchmarks | 1270 | 1358 | 1344 |
| 9 | Reve 2.1reve/reve-2.1 | Reve | 72.0 | 2 of 3 benchmarks | 1302 | 1374 | |
| 10 | Chatgpt Image High Fidelityopenai/chatgpt-image-high-fidelity | OpenAI | 71.0 | 2 of 3 benchmarks | 1390 | 1353 | |
| 11 | GPT Image 1.5 High Fidelityopenai/gpt-image-1.5-high-fidelity | OpenAI | 70.5 | 3 of 3 benchmarks | 1239 | 1370 | 1342 |
| 12 | Grok Imagine Image Qualityx-ai/grok-imagine-image-quality | SpaceXAI | 70.3 | 2 of 3 benchmarks | 1228 | 1390 | |
| 13 | Uni-1.1 (max)luma/uni-1.1:max | Luma AI | 67.8 | 3 of 3 benchmarks | 1188 | 1334 | 1308 |
| 14 | Qwen Image 3.0 Proqwen/qwen-image-3.0-pro | Qwen | 67.0 | 1 of 3 benchmarks | 1258 | ||
| 15 | Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)google/gemini-3.1-flash-lite-image | 66.8 | 3 of 3 benchmarks | 1251 | 1314 | 1286 | |
| 16 | Uni-1.1luma/uni-1.1 | Luma AI | 64.1 | 3 of 3 benchmarks | 1180 | 1312 | 1287 |
| 17 | Qwen Image 2.0 Proqwen/qwen-image-2.0-pro | Qwen | 64.0 | 2 of 3 benchmarks | 1191 | 1303 | |
| 18 | Grok Imagine Imagex-ai/grok-imagine-image | xAI | 63.8 | 2 of 3 benchmarks | 1172 | 1330 | |
| 19 | Ideogram 4.0 Qualityideogram/ideogram-4.0-quality | Ideogram | 63.3 | 1 of 3 benchmarks | 1206 | ||
| 20 | Mai Image 2microsoft/mai-image-2 | Microsoft | 61.9 | 1 of 3 benchmarks | 1182 | ||
| 21 | Cosmos3 Super Text2imagenvidia/cosmos3-super-text2image | NVIDIA | 61.5 | 1 of 3 benchmarks | 1181 | ||
| 22 | Seedream 4.5bytedance-seed/seedream-4.5 | ByteDance Seed | 60.1 | 3 of 3 benchmarks | 1146 | 1301 | 1291 |
| 23 | Recraft V4.1 Utility Prorecraft/recraft-v4.1-utility-pro | Recraft | 60.1 | 1 of 3 benchmarks | 1169 | ||
| 24 | Grok Imagine Image Prox-ai/grok-imagine-image-pro | xAI | 59.1 | 1 of 3 benchmarks | 1161 | ||
| 25 | Hunyuan Image 3.0 Instructtencent/hunyuan-image-3.0-instruct | Tencent | 57.8 | 1 of 3 benchmarks | 1302 | ||
| 26 | Reve v1.5reve/reve-v1.5 | Reve | 57.8 | 1 of 3 benchmarks | 1154 | ||
| 27 | Hunyuan Image 3.0tencent/hunyuan-image-3.0 | Tencent | 57.3 | 1 of 3 benchmarks | 1151 | ||
| 28 | FLUX.2 Maxblack-forest-labs/flux.2-max | Black Forest Labs | 56.5 | 3 of 3 benchmarks | 1162 | 1262 | 1250 |
| 29 | Imagen Ultra 4.0 Generate 001google/imagen-ultra-4.0-generate-001 | 56.4 | 1 of 3 benchmarks | 1148 | |||
| 30 | Seedream 5.0 Litebytedance/seedream-5.0-lite | ByteDance | 55.8 | 3 of 3 benchmarks | 1136 | 1293 | 1270 |
| 31 | Wan2.7 Image Proalibaba/wan2.7-image-pro | Alibaba | 55.4 | 3 of 3 benchmarks | 1103 | 1302 | 1289 |
| 32 | Nano Banana (Gemini 2.5 Flash Image)google/gemini-2.5-flash-image | 55.0 | 3 of 3 benchmarks | 1150 | 1294 | 1234 | |
| 33 | Wan2.6 T2Ialibaba/wan2.6-t2i | Alibaba | 54.0 | 1 of 3 benchmarks | 1136 | ||
| 34 | FLUX.2 Problack-forest-labs/flux.2-pro | Black Forest Labs | 53.9 | 3 of 3 benchmarks | 1155 | 1244 | 1238 |
| 35 | Reve v1.1reve/reve-v1.1 | Reve | 53.7 | 2 of 3 benchmarks | 1261 | 1261 | |
| 36 | Recraft V4.1 Prorecraft/recraft-v4.1-pro | Recraft | 53.6 | 1 of 3 benchmarks | 1130 | ||
| 37 | Wan2.7 Imagealibaba/wan2.7-image | Alibaba | 53.1 | 3 of 3 benchmarks | 1100 | 1301 | 1280 |
| 38 | Imagen 4.0 Generate 001google/imagen-4.0-generate-001 | 53.1 | 1 of 3 benchmarks | 1129 | |||
| 39 | Qwen Image 2512qwen/qwen-image-2512 | Qwen | 52.6 | 1 of 3 benchmarks | 1126 | ||
| 40 | Kling Image O1kwaivgi/kling-image-o1 | Kling AI | 52.5 | 2 of 3 benchmarks | 1251 | 1252 | |
| 41 | Krea 2 (medium)krea/krea-2:medium | Krea | 52.2 | 1 of 3 benchmarks | 1122 | ||
| 42 | Hidream O1 Imagehidream/hidream-o1-image | HiDream | 51.7 | 1 of 3 benchmarks | 1118 | ||
| 43 | Wan2.5 T2I Previewalibaba/wan2.5-t2i-preview | Alibaba | 51.3 | 1 of 3 benchmarks | 1117 | ||
| 44 | FLUX.2 Flexblack-forest-labs/flux.2-flex | Black Forest Labs | 51.1 | 3 of 3 benchmarks | 1156 | 1225 | 1234 |
| 45 | Qwen Image Editqwen/qwen-image-edit | Qwen | 50.3 | 1 of 3 benchmarks | 1241 | ||
| 46 | Recraft V4recraft/recraft-v4 | Recraft | 49.9 | 1 of 3 benchmarks | 1113 | ||
| 47 | Seedream 4 (2K)bytedance/seedream-4:2k | ByteDance | 49.4 | 3 of 3 benchmarks | 1140 | 1271 | 1201 |
| 48 | Krea 2 Turbokrea/krea-2-turbo | Krea | 48.9 | 1 of 3 benchmarks | 1110 | ||
| 49 | Krea 2 Largekrea/krea-2-large | Krea | 48.0 | 1 of 3 benchmarks | 1105 | ||
| 50 | Reve v1reve/reve-v1 | Reve | 47.2 | 2 of 3 benchmarks | 1234 | 1222 | |
| 51 | Mai Image 1microsoft/mai-image-1 | Microsoft | 46.6 | 1 of 3 benchmarks | 1093 | ||
| 52 | Seedream 3bytedance/seedream-3 | ByteDance | 46.2 | 1 of 3 benchmarks | 1082 | ||
| 53 | Wan2.6 Imagealibaba/wan2.6-image | Alibaba | 46.0 | 2 of 3 benchmarks | 1230 | 1219 | |
| 54 | Z Image Turboalibaba/z-image-turbo | Alibaba | 45.7 | 1 of 3 benchmarks | 1081 | ||
| 55 | Flux 2 Devblack-forest-labs/flux-2-dev | Black Forest Labs | 45.7 | 3 of 3 benchmarks | 1146 | 1224 | 1202 |
| 56 | Qwen Image Prompt Extendqwen/qwen-image-prompt-extend | Qwen | 44.3 | 1 of 3 benchmarks | 1060 | ||
| 57 | Reve v1.1 Fastreve/reve-v1.1-fast | Reve | 44.1 | 1 of 3 benchmarks | 1207 | ||
| 58 | Qwen Image Edit 2511qwen/qwen-image-edit-2511 | Qwen | 43.8 | 2 of 3 benchmarks | 1235 | 1173 | |
| 59 | Reve Edit Fastreve/reve-edit-fast | Reve | 43.5 | 1 of 3 benchmarks | 1198 | ||
| 60 | Imagen 3.0 Generate 002google/imagen-3.0-generate-002 | 43.4 | 1 of 3 benchmarks | 1058 | |||
| 61 | Qwen Imageqwen/qwen-image | Qwen | 42.9 | 1 of 3 benchmarks | 1057 | ||
| 62 | Ideogram v3 Qualityideogram/ideogram-v3-quality | Ideogram | 42.5 | 1 of 3 benchmarks | 1049 | ||
| 63 | Seedream 4 High Res Falbytedance/seedream-4-high-res-fal | ByteDance | 42.2 | 3 of 3 benchmarks | 1113 | 1217 | 1209 |
| 64 | Photonluma/photon | Luma AI | 42.0 | 1 of 3 benchmarks | 1035 | ||
| 65 | Runway Gen4runway/runway-gen4 | Runway | 41.1 | 1 of 3 benchmarks | 1025 | ||
| 66 | Flux 2 Klein 9Bblack-forest-labs/flux-2-klein-9b | Black Forest Labs | 40.8 | 3 of 3 benchmarks | 1070 | 1224 | 1213 |
| 67 | Recraft V3recraft/recraft-v3 | Recraft | 40.6 | 1 of 3 benchmarks | 1021 | ||
| 68 | Flux 1.1 Problack-forest-labs/flux-1.1-pro | Black Forest Labs | 40.2 | 1 of 3 benchmarks | 1016 | ||
| 69 | Lucid Originleonardo/lucid-origin | Leonardo AI | 39.7 | 1 of 3 benchmarks | 1013 | ||
| 70 | Seedream 4 Falbytedance/seedream-4-fal | ByteDance | 39.5 | 3 of 3 benchmarks | 1116 | 1210 | 1156 |
| 71 | Ideogram v2ideogram/ideogram-v2 | Ideogram | 39.2 | 1 of 3 benchmarks | 1013 | ||
| 72 | GLM Imagez-ai/glm-image | Z.ai | 38.8 | 1 of 3 benchmarks | 1010 | ||
| 73 | Flux 1 Devblack-forest-labs/flux-1-dev | Black Forest Labs | 37.8 | 1 of 3 benchmarks | 969 | ||
| 74 | DALL E 3openai/dall-e-3 | OpenAI | 37.4 | 1 of 3 benchmarks | 968 | ||
| 75 | Wan2.5 I2I Previewalibaba/wan2.5-i2i-preview | Alibaba | 36.8 | 2 of 3 benchmarks | 1181 | 1160 | |
| 76 | Stable Diffusion v35 Largestability-ai/stable-diffusion-v35-large | Stability AI | 36.4 | 1 of 3 benchmarks | 938 | ||
| 77 | Step1x Editstepfun/step1x-edit | StepFun | 36.0 | 1 of 3 benchmarks | 998 | ||
| 78 | GPT Image 1openai/gpt-image-1 | OpenAI | 35.0 | 3 of 3 benchmarks | 1115 | 1139 | 1126 |
| 79 | FLUX.2 Klein 4Bblack-forest-labs/flux.2-klein-4b | Black Forest Labs | 33.7 | 3 of 3 benchmarks | 1030 | 1188 | 1166 |
| 80 | GPT Image 1 Miniopenai/gpt-image-1-mini | OpenAI | 33.0 | 3 of 3 benchmarks | 1109 | 1124 | 1111 |
| 81 | Flux 1 Kontext (max)black-forest-labs/flux-1-kontext:max | Black Forest Labs | 32.0 | 3 of 3 benchmarks | 1074 | 1181 | 1059 |
| 82 | Flux 1 Kontext Problack-forest-labs/flux-1-kontext-pro | Black Forest Labs | 31.3 | 3 of 3 benchmarks | 1059 | 1176 | 1061 |
| 83 | Seededit 3.0bytedance/seededit-3.0 | ByteDance | 30.2 | 2 of 3 benchmarks | 1139 | 1042 | |
| 84 | Bagelbytedance/bagel | ByteDance | 27.5 | 2 of 3 benchmarks | 898 | 1026 | |
| 85 | Gemini 2.0 Flash Preview Image Generationgoogle/gemini-2.0-flash-preview-image-generation | 24.9 | 3 of 3 benchmarks | 975 | 1081 | 1058 | |
| 86 | Flux 1 Kontext Devblack-forest-labs/flux-1-kontext-dev | Black Forest Labs | 24.6 | 3 of 3 benchmarks | 940 | 1149 | 1034 |
How this ranks
Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.
A model scored on fewer than 2 of the 3 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.
Data sources
Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.
- LMArenaCC BY 4.0
Arena ratings by LMArena, from the public leaderboard dataset.