66 MODELS · 12 PROVIDERS · 4 MODALITIES
Generate with every model.Keep every result.










Made here · From the public gallery · Updated continuously
01 · IMAGE
One line. Four models answer.
GPT Image 2
Gemini 3 Pro
FLUX 2 Pro
Seedream 5.002 · LORA
Find one, load it. Load as many as you like.
STYLE · AnimaLoad
STYLE · AnimaLoad
STYLE · SDXLLoad
POSE · AnimaLoad
DETAIL · AnimaLoad
LIGHT · AnimaLoad
2.0 / 0.1
1.4 / 0.7
0.7 / 1.4
0.1 / 2.003 · AUDIO
Pick a voice, type, and it speaks.
04 · VIDEO
Image, voice, footage, text — all of it can be a reference.
05 · CANVAS
Talk out a script, grow it into nodes, wire it into a cut.
Shot 01 · entrance
Shot 02 · shared umbrella
Shot 03 · arcade06 · VAULT
Everything you keep can go back on stage.
▶ 05 · The Borrowed Umbrella
04 · character anchor
Image
Image
Image
Reference loaded · canvasAny one of these, dragged into the canvas or a workbench, is the reference for the next generation. A result is not the end, it is material.
Start creating →GPT Image 2
The all-rounder: top of the blind-test boards, and the best at reading a complicated brief.
- Text inside the image lands almost every time — poster-grade
- Highest fidelity on long multi-condition prompts; nothing gets dropped
- Chinese, Japanese and Korean render just as cleanly
- Thinking mode adds latency; it is on the slow side
- Expensive in bulk, and transparency support is only fair
Gemini 3 Pro Image
Fast and steady workhorse: 2–5 seconds an image, and the finest editing control here.
- 2–5 seconds an image — editing never breaks your rhythm
- Finest control over inpainting, relighting and camera moves
- Search-grounded: real landmarks and real light have a source
- Holds back on the surreal and the heavily stylised
- Small faces and small type still break
FLUX 2 Pro
The texture one: skin and cloth without the plastic sheen, and rock-steady on long descriptions.
- Real texture: little of the plastic sheen, sharp on objects and materials
- Dual-stream architecture holds long descriptive prompts precisely
- Best value at this quality tier, with no parameter fiddling
- Portraits drift towards a porcelain-doll over-retouch
- Text inside the image is still unreliable
Seedream 5.0
A designer's head: it understands layout and Chinese, and searches the web while it generates.
- Design-grade sense of layout: a magazine cover in one pass
- Native type in 10+ languages, strongest in Chinese
- Searches the web mid-generation, so the facts have a source
- Long strings of text still come out wrong, and it takes 8–15 seconds
- Photographic texture reads AI-ish; weaker than Gemini at realism
Recraft V4 Pro
The designer's tier: the only one that outputs native SVG, and it sets whole paragraphs.
- The only model that generates true vector SVG
- Readable paragraphs where others manage 2–3 words
- Top of the HF text-to-image arena since October 2024
- $0.21 an image is the most expensive in the image section
- Design-specialised: general photorealism is not its ground
NovelAI Diffusion V5
The newest anime specialist: 22 characters in one frame without blending, and a whole comic page in one pass.
- Twice the size plus a 32-channel VAE: detail down to hair strands and trinkets
- Up to 22 characters in one frame — multi-panel comics lay out directly
- Prompts in Japanese, traditional Chinese and more; transparency supported
- Anime-specialised; not for realistic scenes
- Not in the unlimited tier, and billed tighter than V4.5
Illustrious XL
The leading open anime base: enormous character knowledge and the busiest fine-tuning scene.
- Trained on all of Danbooru: the widest character recall here
- Far steadier structure — hands and feet — than its open-source generation
- The busiest fine-tune scene (WAI and others); you will meet it again in LoRA
- Character knowledge stops mid-2024; it will not know a new season
- Realism and design work are not its race
Illustrious XL
The base-model entry: the anime foundation with the liveliest LoRA scene.
- The largest Civitai scene: the fullest set of character and style LoRAs
- Full Danbooru grounding — character terms land straight away
- The API route needs no deployment; the Runner route mounts anything
- Weak towards realism
- New-season characters have to come from a LoRA
WAI-Illustrious
The crowd-favourite tune: a polished anime face straight out of the box.
- A fixture on the community charts — good-looking on default settings
- Takes Illustrious-family LoRAs, so characters reproduce reliably
- Runs straight on the Runner with every parameter open
- Leans sweet; not the one for hard realism
- Queue depends on Runner load
Anima Pencil-XL
Pencil and clean line: sketch-feel illustration in one pass.
- A pencil and clean-line texture nothing else here has
- Works for line art and for finished colour alike
- Runs straight on the Runner with every parameter open
- Narrow subject range: illustration is the point
- Not made for thick-paint realism
Pony Diffusion V6 XL
The grandfather of the style universe: somebody has trained every look on it.
- The widest LoRA compatibility of any base here
- A mature system of style trigger words
- The deepest community back catalogue
- Plain faces out of the box; LoRAs do the lifting
- Prompt syntax has a learning curve (the score terms)
SDXL 1.0
The general-purpose veteran: realism through illustration, it takes it all.
- The open base with the broadest ecosystem coverage
- Realism, design and illustration all run on it
- The most mature tooling and the most documentation
- Raw quality has been left behind by newer models
- Standing out takes a LoRA stack
Anima(DiT)
The new-architecture testbed: an anime newcomer on a DiT foundation.
- A new DiT architecture with a high ceiling for detail
- The v4 pipeline is already deployed here
- A fresh option on the anime side
- The scene is only starting; few LoRAs
- Generation takes longer
Seedance
Reference generation in one go: the video workhorse, on 18 routes.
- Reference endpoint: an image and three voice samples, fed together
- Every tier from 2.0 to 2.5 is live, with four integrations backing each other up
- The steadiest first-frame image-to-video here
- Top-end 2.5 is the priciest per second
- Past 10 seconds you have to stitch
MiniMax H3
The value workhorse: two regions, and reference input on both.
- A reference endpoint even at the low tier
- The China routes are friendly on latency
- Well regarded for natural human motion
- Few 1080p tiers
- Weaker than Kling on stylised camera work
Wan 3.0
Alibaba's workhorse: shipped upstream, and we are still adding routes.
- The lowest unit price on the whole site, at $0.10 a second
- The reference-to-video route is live
- Upstream keeps shipping endpoints and we keep adding them
- Plain camera language
- Only fair consistency over longer takes
HappyHorse 1.1
The fast-tier dark horse: top-five value on the blind-test board.
- Top five on the Artificial Analysis blind-test board (v1.1 and v1.0 share the ranking — our note)
- Quick to finish a clip
- Friendly unit price
- Only fair style control
- Little written about it
可灵
The camera one: the name people trust for movement and texture.
- First rank for camera movement and physical texture
- O3 Pro Omni is the all-round endpoint
- Delicate human performance
- 3.0 Pro does not take reference input (our note)
- Upper-middle unit price
Gemini Omni Flash
The quick all-rounder: get the idea moving, cheaply.
- Text-to-video and image-to-video in one place
- Fast and cheap — right for a rough pass
- Strong image understanding from the Google side
- A lower ceiling than the dedicated video models
- Coarse style control
Fish Audio S2.1 Pro
The workhorse for lines: the name people trust for natural Chinese speech.
- First rank for natural spoken Chinese
- Pick from the in-app voice library, or clone
- S2.1 Pro free and paid tiers connected
- Slightly flat English accent
- Extreme emotion needs the temperature adjusted
ElevenLabs
Effects and score: SFX and Music, two routes.
- SFX: one sentence, one foley take
- Music: the structure of a full cue stays under control
- The industry benchmark for audio quality
- Only fair with Chinese lyrics
- Billed per second, so length has to be watched
Rodin Gen-2.5
Multi-view control: up to 5 references feed a single model.
- Up to 5 reference images (our cap)
- Topology clean enough to keep editing afterwards
- A four-view workflow
- On the slow tier for generation
- Only fair texture ceiling
Hunyuan3D
The polygon king: a 1.5-million-face ceiling for the detail-minded.
- 40k to 1.5M faces, adjustable (500k by default)
- Textured mesh in one pass
- v3.1 Pro sharpens the detail further
- High-poly exports are large files
- Better at organic forms than hard surfaces
Trellis 2
Every dial exposed: resolution, texture and decimation are all yours.
- Three resolution tiers: 512 / 1024 / 1536
- Textures from 1K to 4K
- Decimation from 5k to 2M drops straight into a game pipeline
- Middling speed
- Stylised input trips it up
TripoSR
The seconds tier: one image, one grey model, a few seconds.
- The fastest model here
- Light grey models, right for a mockup
- Open-source roots, friendly ecosystem
- No texture detail
- Complex structures blur
Generating is where it starts.
Generate, archive and reuse in one place. Every result keeps its key settings.
Open the canvas