← Back to blog
Ai5 min read

How Cutthroat Have Domestic AI Image Generators Gotten in the Past Month?

Published Sep 27, 2026
How Cutthroat Have Domestic AI Image Generators Gotten in the Past Month?

Over the past month, Chinese and overseas communities have, unusually, been talking about the same thing: domestic AI image generation models.

On September 20, Alibaba's Tongyi team's Qwen Image 2.1 hit the Hacker News front page, with a post earning 739 points and nearly 200 comments. The comment section wasn't just pleasantries. Someone said, "A 7-billion-parameter diffusion model renders Chinese type better than Windows." Others wanted to know how to actually run it locally, and one person came right out and said they were "a little worried"—because this much capability currently comes with no watermark at all.

This isn't an isolated event. That same week, Tencent Hunyuan's new text-to-image model was being tested by creators on Bilibili, and ByteDance's Doubao image-editing tutorials racked up over ten thousand views there. Domestic text-to-image tools are moving from a stage where they "can produce images" to one where "people are actually using them in earnest." Here are a few signals I've noticed.

Why 7 Billion Parameters Is Worth Talking About

Most people don't care about parameters—they only care whether the image looks good. But that 7 billion figure is the key to the whole thing.

Mainstream closed-source image models routinely have tens of billions of parameters and sit behind an API charging per call. Qwen Image 2.1 has open weights, 7 billion parameters, and runs locally on a single consumer-grade graphics card. That means you can generate images on your own computer, without paying per image, and no one logs your prompts.

Tom's Hardware reported that Alibaba claims this model beats Google's Nano Banana 2.0 on benchmarks and trades blows with OpenAI's and Meta's image models. I usually take vendor benchmarks with a grain of salt, but getting the machine learning crowd arguing is itself telling.

What the Chinese-Language Community Is Actually Talking About

I spent a few days scrolling through several platforms, and the direction of discussion in Chinese communities is a bit different from overseas.

Overseas, people argue about open vs. closed source and whether watermarks are needed. Chinese communities are more practical: how to use it, what prompts to write, and whether it can do portrait photography.

On Bilibili, videos like "Doubao AI Image Editing Tutorial" pull in 17,000 views, with danmaku and comments all asking for the exact steps. Demand for things like "AI scientific figure drawing" and "AI custom photos" is also emerging. On Baidu Search, "how to write AI text-to-image prompts" is a high-frequency query. This shows the real user base isn't dabbling—these are people using it for work: e-commerce images, academic figures, personal portraits.

Prompts Aren't That Complicated, Honestly

Break the image in your head into a few concrete elements: what the subject is, what style, what lighting, what composition, and what you don't want.

Don't write "a nice-looking picture." Write "a woman in a beige coat standing by a floor-to-ceiling window, soft backlight, cinematic, shallow depth of field." Models feed on specifics, not piles of adjectives. Want a certain style? Put the style words in. Don't want something? State it explicitly with a negative prompt. Try several versions and save the prompts you like—that's your own template.

Text prompt being turned into an image

A Few Real Use Cases

Scientific figure drawing is an underrated direction. On Bilibili, some creators specifically teach "replicating AI scientific figures," with views in the thousands to tens of thousands. People who make figures and diagrams for papers used to either learn a pile of specialized software or pay to outsource it. Now they can generate a first draft with text-to-image and then refine it—a completely different level of efficiency.

E-commerce images are another big one. Batch-swapping backgrounds, changing poses, and producing showcase images in different styles is far faster with text-to-image and inpainting than with traditional retouching. Then there are personal portraits and "AI custom photos"—demand has always been strong, and the question was never "can it be done" but "does it actually look like me, and how is privacy protected?"

AI-generated e-commerce product images

How Ordinary Users Should Choose

There's no single right answer here, so I'll go by need:

If you just want to mess around and occasionally generate an avatar, a tool like Doubao is enough—no fuss, just open it and go. If you're making e-commerce images or generating in bulk, it's worth spending time on prompts and inpainting; Doubao and Jimeng both have the relevant features. If you're technically inclined and want to self-host, Qwen Image 2.1 is the most worth trying lately: open source, free, runs on a single graphics card, and Chinese text rendering is one of its strengths.

To be fair, domestic text-to-image models have improved genuinely fast over the past two years, but "fast" doesn't mean "strong at everything." The old problems—hands, text, multi-subject consistency—still linger to varying degrees. Don't let benchmarks set your expectations; run a few images through your own real use case, and that'll be more useful than reading a hundred reviews.

A Final Word on Watermarks

Watermarking and provenance is a heated debate overseas, while Chinese communities discuss it less for now—but that doesn't mean it doesn't matter. Once AI image generation is fully ubiquitous, whether a given image was made by a person, and who made it, will eventually become as basic a question as "is this link a scam?"

By then, whoever does open-source capability, ease of use, and honest provenance labeling well together will be the one still standing. Domestic models now have the benchmark scores and the community buzz; what they compete on next are these "boring" but longer-lasting things.

Related articles