To make a YouTube thumbnail with AI, you give the tool a reference whose idea you admire, your own photo (or no photo), and the words that go on the image, and it designs a new thumbnail from that brief. In WThumb that is one form with seven steps and a wait of about thirty seconds, and new accounts get two free credits (1 credit = 1 thumbnail) with no card needed.
This is the full walk-through across every Studio choice in the AI YouTube thumbnail generator. If you only want the URL route, use the YouTube-link workflow; for the seven decisions under any tool, read how to make a YouTube thumbnail. At each step we say what the form asks for, what WThumb does with it, and what makes a request fail, with a worked example brief near the end.
One rule sits under all of it: the reference is a brief, not a template. WThumb keeps the subject, the visual idea, the energy and the job each element does, and it designs the layout, the framing, the typography and the color treatment itself. The idea carries over; the design is new.
How to make a YouTube thumbnail with AI
The whole process in one list. Each step is expanded in the sections below.
- Pick the format. 16:9 for a normal video, 9:16 for Shorts, or tick both.
- Give it a reference. Paste a link to a video whose thumbnail idea you like, or upload an image.
- Add your photo, or skip it. Your photo supplies identity input for a fresh composition. Without one, WThumb instructs the model to remove people and supplies no replacement photo. Review the finished image before publishing.
- Type the words. Up to three lines of text for the image, plus optional instructions in plain sentences.
- Choose one or two designs. Both come from the same brief.
- Wait. A run takes about thirty seconds.
- Download, or revise it. The pencil on the thumbnail takes a change in plain words.
What you need before you start
Three things, and none of them is a design skill. A WThumb account, because new accounts get two free credits and those are what your first runs use. A reference, meaning a YouTube link or an image file. And the words the thumbnail should say. If you plan to appear on it, have a photo of yourself ready; if you run a faceless channel, you need nothing else.
Steps 1 to 3: format, reference and photo
Step 1: pick the format, or both
The first choice after you pick Recreate a thumbnail is the shape of the frame. 16:9 is the standard YouTube thumbnail and comes back as a high quality Full HD+ JPEG at 2048 by 1152 pixels. 9:16 is for Shorts and comes back at 1152 by 2048. Both are above what YouTube asks for, so there is nothing to resize.
You can tick both in one submission: write the brief once and it is designed for both shapes, so a video and its Short come out of the same form.
Step 2: paste a link, or upload a reference
Paste a YouTube link and WThumb fetches that video's public thumbnail image, and nothing else about the video. Or upload an image of your own: PNG, JPEG or WebP, up to 12 MiB. Either way, this image is the brief.
So choose it for its idea, not its polish. A person reacting to a number. An object with a red cross through it. What carries over is that idea, the subject, the energy and the role each element plays. What does not carry over is the layout, the pose, the framing or the color treatment, because WThumb designs those itself.
If you cannot say in one sentence what the reference is doing, pick another. When the brief is clear, recreate a thumbnail from a reference with your own subject and words.
Step 3: add your photo, or skip it for faceless
The photo field is optional. Add a clear personal photo, up to 8 MiB, to supply your identity for the new design. Skip it to request a faceless version: WThumb supplies no personal or replacement photo and instructs the image model to remove people. Review the finished image for unwanted people or details before publishing.
A good photo is a clear, well lit shot of you, front on or close to it, with an expression that fits the video. You do not need to match the pose in the reference, because WThumb designs the framing itself.


Steps 4 and 5: the words, the instructions and how many designs
Step 4: type the words, then add instructions
There are two text fields and they do different jobs. The first holds the literal words that go on the thumbnail, up to three lines. Type exactly what you want to read on the image, and nothing else. Short is better, because it has to be legible on a phone.
The second field is optional, and it is where you talk to the tool in plain sentences: the color treatment, the object at the center, the mood, a logo. You can attach up to four extra images and point at them as @1, @2, @3 and @4, which is how a specific product, screenshot or channel mark gets into the design instead of an approximation of it.
Write instructions the way you would brief a person. Make it pop is not an instruction, because nobody knows what it means. Swap the blue background for a warm brown and keep the mug as the main object is one. Every sentence should name a thing and say what happens to it.
Before anything is generated, your words go through a moderation service. A refused request is never generated and costs nothing.
Step 5: choose one design or two
The last choice is how many designs you want from this brief: one or two. Two is useful when you want the same idea taken two ways. On a first run, one is enough to see how your brief reads.
Steps 6 and 7: wait, then download or change one thing
Step 6: wait about thirty seconds
Once you submit, the generation runs on an OpenAI image model through WThumb's own instructions, and takes about thirty seconds from link to finished image. Write the title while it works.
Step 7: download it, or revise it
The finished thumbnail gives you two ways out. Download the JPEG and upload it to YouTube, or press the pencil on it: one field where you describe one change in plain words, and WThumb makes a new version with that change.
Three things about a revision are worth knowing. It costs one credit, the same as a fresh run. There is no limit: the new version can be revised again. And it is not only for now; reopen the conversation later and the pencil is still on the latest version.
There is no editor in WThumb. No canvas, no layers, nothing to drag. The way you change a thumbnail is to describe the change in words, in a revision.
A worked example brief
Say the video is a thirty day experiment of giving up coffee. Here is a brief that works, with the reasoning beside each choice.
- Format: 16:9 only, because this one is a full upload. If it were also going out as a Short, we would tick 9:16 as well.
- Reference: a link to a video whose thumbnail shows a creator holding up a mug with a red cross over it and a bold number beside them. We chose it for the idea (person, object, cross, number), not for its colors or layout.
- Photo: a front on shot with a tired but determined face, taken by a window. It does not match the reference pose and does not need to.
- Words: line one 30 DAYS, line two NO COFFEE. Two lines, four words, readable on a phone.
- Instructions: Keep the mug as the main object with the red cross over it. Swap the blue background for a warm brown. Include the wall calendar from @1, small. We attached a photo of our own calendar as @1.
- Designs: two, because we could not decide whether the mug or the number should dominate.
What comes back about thirty seconds later is two thumbnails that both have us, a mug, a red cross, the number and the calendar on a brown ground, and neither looks like the reference. One puts the number huge and the mug small in our hand; the other puts the mug at the center and the words down one side. The idea carried over. The design is new, twice.
We picked the second and revised it for one thing: make the cross thicker. Pasting the link to a file on disk was two waits of about thirty seconds each, plus a minute to fill in the form.
What to do when the result is wrong
Sometimes it is wrong. The reasons fall into a short list, and most are fixed in the brief, not the output.
- The words were vague. Make it look professional, make it exciting, make it clean. None of these names a thing. Name the object, the color and the mood.
- You asked for a copy. Make it exactly like the reference but with my face is the request WThumb is built not to fulfill. It keeps the idea and designs the rest, so that brief always comes back different from what you pictured.
- The reference had no idea in it. A landscape, a blank screenshot, a product on a table. Nothing to keep, so nothing kept. Choose a thumbnail that is doing something.
- The photo did not fit the mood. A neutral passport face under a headline like I RUINED MY KITCHEN reads as odd because it is odd. Reshoot with the expression the words need.
- The instructions contradicted the words. If the words say NO COFFEE and the instructions say make the coffee look delicious, something has to give.
- The request was refused. The moderation check said no. Nothing was generated and nothing was spent. Rewrite the brief without whatever tripped it.
If the result is close and one thing is off, a revision is the tool. If it is far off, do not spend a revision on it; fix the brief and run again. Either way, read the words on the finished image back before you upload it.
What AI thumbnail makers cannot do (yet)
The plain list of what WThumb does not do, and what to use instead.
- It cannot tell you whether the thumbnail will get clicks. There are no analytics, no score, no click reports and no way to run two thumbnails against each other. Nobody can promise a result for a video from the image alone.
- It cannot copy a layout. By design. A pixel accurate copy of a specific layout is a job for an editor and your own hands.
- It cannot be nudged by hand. There is no canvas. The adjustment WThumb makes is a revision, a change described in words. For anything by hand, such as a crop, a border or a logo overlay, take the finished JPEG into Canva, Adobe Express, Figma or Photoshop, which are editors built for that.
- It does not need a reference. This walk-through follows Recreate, which starts from one. The Studio's other start, from a prompt, takes the video's title and a plain description of what happens in it, or of the thumbnail you want, instead.
- The personal-photo step is optional. Skip it to request a faceless result. WThumb supplies no replacement photo and instructs the model to remove people; review the finished image for unwanted people or details.
- It cannot upload for you. You download a JPEG and upload it in YouTube Studio like any other image.
Other AI thumbnail generators such as Pikzels and Thumbmagic work in the same broad family, a brief in and an image out, and each draws its own lines about references and editing. vidIQ is a channel analytics and keyword toolkit that also offers thumbnail tools, so it answers a different question. Row by row, WThumb vs vidIQ for thumbnails and the VisualKit alternative compare two such makers with WThumb. Judge each on whether the image it gives you is one you would upload.
For WThumb, that test is free. New accounts get two free credits, no card needed, and paid plans are there when you need more. Pick a reference with a real idea in it, write the words, and see whether describing a thumbnail gets you further than building one.



