OpenAI shipped GPT Image 2 yesterday. The benchmarks are loud — a 242-point Elo lead on text-to-image over the next-best model on Arena, which is the kind of gap you don't usually see in a mature category. The demos are mostly magazine covers, manga panels, and product mockups.
None of that is how I use image models at 312. I use them to teach. And teaching is a different stress test from anything the launch reels optimize for.
This is a short post on what actually matters when you're generating images for a lesson, how the current field holds up against those criteria, and what I've switched to as of this morning.
Teaching is a weird use case for image models
Most people judging an image model are looking at a single hero image: does it look good, does the composition work, is the style coherent. Teachers look at images in batches, embedded in explanations, under time pressure, with a specific claim the image has to make. The bar is different.
Four things matter for instructional images, roughly in this order:
Aesthetic quality is fifth. A beautiful image of the wrong thing is worse than a plain image of the right thing, and teachers know this immediately because a student will raise their hand.
What GPT Image 2 actually changed
The headline feature of GPT Image 2 is "thinking" — when you use the thinking variant in ChatGPT, the model can search the web, analyze uploaded materials, reason through layout before rendering, and self-check its own output. For poster design this is a nice-to-have. For teaching, it's the thing.
Take a prompt like "a labeled cross-section of a plant stem showing xylem, phloem, cambium, and epidermis, annotated for a 7th-grade biology worksheet." Every prior model I've used will produce something that looks like a labeled cross-section and gets at least one label wrong. Sometimes the xylem and phloem swap. Sometimes the cambium is in the wrong place. Sometimes the spelling drifts.
GPT Image 2 with thinking checks its own work against something. I don't know exactly what — web content, its own reasoning pass, probably both — but on my test run of ten biology and physics prompts, it caught mistakes that every other model I tried made. Not all of them. It still produced a free-body diagram with the normal force vector pointing the wrong way on an inclined plane. But it flagged uncertainty on two prompts and asked clarifying questions, which none of the other models did.
That's the actual shift: an image model that knows when it doesn't know.
The field, by instructional criteria
Here's how the current contenders hold up when you score them against what teaching actually needs, not what makes a good marketing demo.
| Model | Text | Factual | Consistency | Cost signal |
|---|---|---|---|---|
| GPT Image 2 with thinking | Best in class | Self-checks | Good in session | $0.04 base; more w/ thinking |
| Nano Banana 2 Gemini 3 Pro Image | Very strong | No verification | Good | Free tier usable |
| Seedream 4.5 ByteDance | Strong, incl. non-Latin | Average | Best — 6 ref images | Cheapest paid |
| Midjourney v8 aesthetic lead | Weak on labels | Vibes over facts | Hard to control | $10/mo min |
| Flux 2 Max open-weight | OK | Average | Tunable if hosted | Free self-hosted |
| Ideogram typography specialist | Excellent | Not built for it | Limited | Mid |
Midjourney is out for this use case. It's the model I'd use to design a book cover for a course, not a single figure inside one. Flux is interesting if you want to fine-tune on your own curriculum's visual language, but most teachers won't.
The real decision is between GPT Image 2, Nano Banana 2, and — for specific jobs — Seedream 4.5.
The volume problem
Here's where cost starts mattering in a way that benchmarks ignore.
A single teaching unit is maybe 30–60 images. An 8-week course is 400+. If you're iterating — and you always iterate, because the first image is never the one you use — multiply by three or four. That's how you end up generating a couple thousand images per course.
$62 to build a course on the best model in the world is, let's be honest, nothing. The thinking variant at ~$180 is also fine if the course is going to be used by hundreds of students. But the pattern I've landed on is mixed: thinking mode only for images where correctness matters (labeled diagrams, equations, factual illustrations), base model or Nano Banana for everything else (scene illustrations, characters, covers, decorative elements).
The consistency problem nobody talks about
One thing almost every review I read missed: teaching materials aren't one image, they're a series. If your lesson has a recurring character — a student asking questions, a scientist in a lab, a generic "you" avatar — that character needs to look like the same person across 12 slides. Same hair, same clothes, same style.
This is where Seedream 4.5 quietly wins. It accepts up to six reference images and preserves structural coherence across character, pose, lighting, and environment. I've been using it for character sheets and recurring scenes inside a unit, then layering GPT Image 2 for anything that needs accurate labels.
GPT Image 2 has improved on multi-image consistency — it topped the Arena multi-image edit leaderboard too — but Seedream's reference-image handling is still more predictable in my tests.
What I'm using as of today
Midjourney — aesthetic but uncontrollable for labeled work. I'd use it for a course cover, not a course figure. DALL-E 3 — retiring May 12 anyway. Flux — great if you're fine-tuning on a curriculum's visual language, but that's a different project.
Three days ago this would have been a simpler list — Nano Banana 2 was the default for almost everything and the thinking tier didn't exist for images. The GPT Image 2 launch actually matters because it's the first model with explicit reasoning over visual output, and teaching is the use case where that matters most.
What I'm watching
Two things.
First: whether the thinking premium holds up for teaching specifically. If a school's use case is heavy on factual diagrams, that cost delta compounds fast. If Nano Banana adds reasoning (I'd bet they do within a few months), the calculus changes again.
Second: the aggregators. Freepik and WaveSpeedAI let you run the same prompt across many models without juggling subscriptions. For a school evaluating what to standardize on, that's probably a better first stop than picking a model and committing.
None of this is a finished opinion. The field moved more in the last 48 hours than in the previous two months. I'm writing this down now because that's how it works here — we build, we try, we write down what we learned the same week.