Photorealistic images
The model card puts photorealistic generation first: portraits, street scenes, product-style photos.
Run Tongyi's Z-Image Turbo from your browser. Write a prompt and get an image quickly, for 1 credit. Free credits when you sign up, nothing to install.
Z-Image Turbo is text-to-image only and has one output size (about 1 megapixel). To edit a photo or generate at 2K or 4K, use the standard image generator.
Z-Image Turbo is an open-source image generation model from Alibaba's Tongyi-MAI team, released in November 2025 under the Apache 2.0 license. It has 6 billion parameters and is built on what the team calls a Single-Stream Diffusion Transformer.
Turbo is the distilled version of Z-Image: according to its model card it needs only 8 sampling steps per image, which is why it is fast and cheap to run. That is also why an image costs only 1 credit here.
The strengths listed on the Z-Image Turbo model card.
The model card puts photorealistic generation first: portraits, street scenes, product-style photos.
It renders text in both English and Chinese, so signs, posters and labels with Chinese characters come out readable.
The model card describes its instruction following as robust: what you ask for and where you place it tends to hold.
Eight sampling steps instead of the usual dozens. That is why images come back much faster here than on our standard generator.
Each of these was generated on this page. The caption is the exact prompt sent to the model.




What follows from how the model works.
The Turbo model does not support negative prompts. Instead of "no people", describe the empty street you want to see.
Put the exact words that must appear in quotes and say where they go. Chinese works as well as English.
Subject, setting, light, lens, style. Prompts can be up to 1,000 characters, and the model uses the detail you give it.
At 1 credit per image and a short wait, the quickest way to a good result is to generate a few and adjust the prompt.
1 credit per image, the lowest price of any tool on TikTomato. New accounts get free credits, so you can try it without paying.
Much less than on our standard generator. The exact time varies with load, and the first image after a quiet period can take longer.
About 1 megapixel. A 16:9 image is 1280x720 pixels, 4:3 is 1152x864. You can choose 1:1, 4:3, 3:4, 16:9 or 9:16. For something larger, run the result through our image upscaler.
Not with Z-Image Turbo: it is a text-to-image model. Our standard image generator takes up to 4 reference images.
Z-Image Turbo is a much smaller model built for speed and cost. It is a good fit for quick ideas, drafts and photorealistic scenes. For complex compositions, long passages of text or editing photos, GPT Image on our standard generator is the stronger choice.
The model is released under the Apache 2.0 license, which does not restrict commercial use of its output. You are responsible for what you create and how you use it.
Prompts and results are checked for sexual and other sensitive content, and flagged requests are declined. Rephrase the prompt and try again. Credits for declined or failed requests are refunded.