All posts

Thumbnails9 min readUpdated

How to make a YouTube thumbnail with AI, step by step

Seven steps from a YouTube link to a finished 16:9 or 9:16 thumbnail in WThumb, with a worked example brief, the requests that fail before they start, and what an AI thumbnail maker still cannot do.

Seven small circles joined by one rising line, the last circle larger and red, headline: THUMBNAIL WITH AI

To make a YouTube thumbnail with AI, you give the tool a reference whose idea you admire, your own photo (or no photo), and the words that go on the image, and it designs a new thumbnail from that brief. In WThumb that is one form with seven steps and a wait of about thirty seconds, and new accounts get two free credits (1 credit = 1 thumbnail) with no card needed.

This is the full walk-through across every Studio choice in the AI YouTube thumbnail generator. If you only want the URL route, use the YouTube-link workflow; for the seven decisions under any tool, read how to make a YouTube thumbnail. At each step we say what the form asks for, what WThumb does with it, and what makes a request fail, with a worked example brief near the end.

One rule sits under all of it: the reference is a brief, not a template. WThumb keeps the subject, the visual idea, the energy and the job each element does, and it designs the layout, the framing, the typography and the color treatment itself. The idea carries over; the design is new.

How to make a YouTube thumbnail with AI

The whole process in one list. Each step is expanded in the sections below.

  1. Pick the format. 16:9 for a normal video, 9:16 for Shorts, or tick both.
  2. Give it a reference. Paste a link to a video whose thumbnail idea you like, or upload an image.
  3. Add your photo, or skip it. Your photo supplies identity input for a fresh composition. Without one, WThumb instructs the model to remove people and supplies no replacement photo. Review the finished image before publishing.
  4. Type the words. Up to three lines of text for the image, plus optional instructions in plain sentences.
  5. Choose one or two designs. Both come from the same brief.
  6. Wait. A run takes about thirty seconds.
  7. Download, or revise it. The pencil on the thumbnail takes a change in plain words.

What you need before you start

Three things, and none of them is a design skill. A WThumb account, because new accounts get two free credits and those are what your first runs use. A reference, meaning a YouTube link or an image file. And the words the thumbnail should say. If you plan to appear on it, have a photo of yourself ready; if you run a faceless channel, you need nothing else.

Steps 1 to 3: format, reference and photo

Step 1: pick the format, or both

The first choice after you pick Recreate a thumbnail is the shape of the frame. 16:9 is the standard YouTube thumbnail and comes back as a high quality Full HD+ JPEG at 2048 by 1152 pixels. 9:16 is for Shorts and comes back at 1152 by 2048. Both are above what YouTube asks for, so there is nothing to resize.

You can tick both in one submission: write the brief once and it is designed for both shapes, so a video and its Short come out of the same form.

Step 2: paste a link, or upload a reference

Paste a YouTube link and WThumb fetches that video's public thumbnail image, and nothing else about the video. Or upload an image of your own: PNG, JPEG or WebP, up to 12 MiB. Either way, this image is the brief.

So choose it for its idea, not its polish. A person reacting to a number. An object with a red cross through it. What carries over is that idea, the subject, the energy and the role each element plays. What does not carry over is the layout, the pose, the framing or the color treatment, because WThumb designs those itself.

If you cannot say in one sentence what the reference is doing, pick another. When the brief is clear, recreate a thumbnail from a reference with your own subject and words.

Step 3: add your photo, or skip it for faceless

The photo field is optional. Add a clear personal photo, up to 8 MiB, to supply your identity for the new design. Skip it to request a faceless version: WThumb supplies no personal or replacement photo and instructs the image model to remove people. Review the finished image for unwanted people or details before publishing.

A good photo is a clear, well lit shot of you, front on or close to it, with an expression that fits the video. You do not need to match the pose in the reference, because WThumb designs the framing itself.

Reference briefA fishing reference with a golden fish and bold episode text supplied as the visual brief
WThumb resultA new camping thumbnail made by WThumb with a creator holding a lantern at purple dusk
The reference gives WThumb a visual direction; the title, creator photo, and instructions produce a new composition.

Steps 4 and 5: the words, the instructions and how many designs

Step 4: type the words, then add instructions

There are two text fields and they do different jobs. The first holds the literal words that go on the thumbnail, up to three lines. Type exactly what you want to read on the image, and nothing else. Short is better, because it has to be legible on a phone.

The second field is optional, and it is where you talk to the tool in plain sentences: the color treatment, the object at the center, the mood, a logo. You can attach up to four extra images and point at them as @1, @2, @3 and @4, which is how a specific product, screenshot or channel mark gets into the design instead of an approximation of it.

Write instructions the way you would brief a person. Make it pop is not an instruction, because nobody knows what it means. Swap the blue background for a warm brown and keep the mug as the main object is one. Every sentence should name a thing and say what happens to it.

Before anything is generated, your words go through a moderation service. A refused request is never generated and costs nothing.

Step 5: choose one design or two

The last choice is how many designs you want from this brief: one or two. Two is useful when you want the same idea taken two ways. On a first run, one is enough to see how your brief reads.

Steps 6 and 7: wait, then download or change one thing

Step 6: wait about thirty seconds

Once you submit, the generation runs on an OpenAI image model through WThumb's own instructions, and takes about thirty seconds from link to finished image. Write the title while it works.

Step 7: download it, or revise it

The finished thumbnail gives you two ways out. Download the JPEG and upload it to YouTube, or press the pencil on it: one field where you describe one change in plain words, and WThumb makes a new version with that change.

Three things about a revision are worth knowing. It costs one credit, the same as a fresh run. There is no limit: the new version can be revised again. And it is not only for now; reopen the conversation later and the pencil is still on the latest version.

There is no editor in WThumb. No canvas, no layers, nothing to drag. The way you change a thumbnail is to describe the change in words, in a revision.

A worked example brief

Say the video is a thirty day experiment of giving up coffee. Here is a brief that works, with the reasoning beside each choice.

  • Format: 16:9 only, because this one is a full upload. If it were also going out as a Short, we would tick 9:16 as well.
  • Reference: a link to a video whose thumbnail shows a creator holding up a mug with a red cross over it and a bold number beside them. We chose it for the idea (person, object, cross, number), not for its colors or layout.
  • Photo: a front on shot with a tired but determined face, taken by a window. It does not match the reference pose and does not need to.
  • Words: line one 30 DAYS, line two NO COFFEE. Two lines, four words, readable on a phone.
  • Instructions: Keep the mug as the main object with the red cross over it. Swap the blue background for a warm brown. Include the wall calendar from @1, small. We attached a photo of our own calendar as @1.
  • Designs: two, because we could not decide whether the mug or the number should dominate.

What comes back about thirty seconds later is two thumbnails that both have us, a mug, a red cross, the number and the calendar on a brown ground, and neither looks like the reference. One puts the number huge and the mug small in our hand; the other puts the mug at the center and the words down one side. The idea carried over. The design is new, twice.

We picked the second and revised it for one thing: make the cross thicker. Pasting the link to a file on disk was two waits of about thirty seconds each, plus a minute to fill in the form.

What to do when the result is wrong

Sometimes it is wrong. The reasons fall into a short list, and most are fixed in the brief, not the output.

  • The words were vague. Make it look professional, make it exciting, make it clean. None of these names a thing. Name the object, the color and the mood.
  • You asked for a copy. Make it exactly like the reference but with my face is the request WThumb is built not to fulfill. It keeps the idea and designs the rest, so that brief always comes back different from what you pictured.
  • The reference had no idea in it. A landscape, a blank screenshot, a product on a table. Nothing to keep, so nothing kept. Choose a thumbnail that is doing something.
  • The photo did not fit the mood. A neutral passport face under a headline like I RUINED MY KITCHEN reads as odd because it is odd. Reshoot with the expression the words need.
  • The instructions contradicted the words. If the words say NO COFFEE and the instructions say make the coffee look delicious, something has to give.
  • The request was refused. The moderation check said no. Nothing was generated and nothing was spent. Rewrite the brief without whatever tripped it.

If the result is close and one thing is off, a revision is the tool. If it is far off, do not spend a revision on it; fix the brief and run again. Either way, read the words on the finished image back before you upload it.

What AI thumbnail makers cannot do (yet)

The plain list of what WThumb does not do, and what to use instead.

  • It cannot tell you whether the thumbnail will get clicks. There are no analytics, no score, no click reports and no way to run two thumbnails against each other. Nobody can promise a result for a video from the image alone.
  • It cannot copy a layout. By design. A pixel accurate copy of a specific layout is a job for an editor and your own hands.
  • It cannot be nudged by hand. There is no canvas. The adjustment WThumb makes is a revision, a change described in words. For anything by hand, such as a crop, a border or a logo overlay, take the finished JPEG into Canva, Adobe Express, Figma or Photoshop, which are editors built for that.
  • It does not need a reference. This walk-through follows Recreate, which starts from one. The Studio's other start, from a prompt, takes the video's title and a plain description of what happens in it, or of the thumbnail you want, instead.
  • The personal-photo step is optional. Skip it to request a faceless result. WThumb supplies no replacement photo and instructs the model to remove people; review the finished image for unwanted people or details.
  • It cannot upload for you. You download a JPEG and upload it in YouTube Studio like any other image.

Other AI thumbnail generators such as Pikzels and Thumbmagic work in the same broad family, a brief in and an image out, and each draws its own lines about references and editing. vidIQ is a channel analytics and keyword toolkit that also offers thumbnail tools, so it answers a different question. Row by row, WThumb vs vidIQ for thumbnails and the VisualKit alternative compare two such makers with WThumb. Judge each on whether the image it gives you is one you would upload.

For WThumb, that test is free. New accounts get two free credits, no card needed, and paid plans are there when you need more. Pick a reference with a real idea in it, write the words, and see whether describing a thumbnail gets you further than building one.

Common questions

Can I make a YouTube thumbnail with AI for free?
Yes. In WThumb, new accounts get two free credits and no card is needed. That is enough to run one real brief and one revision, or two separate briefs, and see whether the result is something you would upload. Paid plans are there for people who want more than that.
How long does it take to make a thumbnail with AI?
In WThumb a run takes about thirty seconds from pasting the link to a finished image. Filling in the form takes a minute or two on top of that. A revision makes a new version, which is another generation, so expect a similar wait for it.
Do I need a photo of myself to use an AI thumbnail maker?
Not in WThumb. The photo step is optional. Add a personal photo as identity input, or skip it to request a faceless composition. Without a photo, WThumb supplies no replacement photo and instructs the model to remove people. Review the finished image before publishing, because model output can miss an instruction.
Will the AI copy the thumbnail I paste as a reference?
No. WThumb treats the reference as a brief, not a template. It keeps the subject, the visual idea, the energy and the role each element plays, and it designs the layout, framing, typography and colors itself. The idea carries over; the design is new.
Can I edit the thumbnail after the AI makes it?
There is no editor in WThumb, so there is no canvas or layers to adjust. Press the pencil on a finished thumbnail, describe a single change in plain words and get a new version; each revision costs one credit, and there is no limit on revisions. For hand adjustments, open the downloaded JPEG in an editor such as Canva or Photoshop.
What size is a thumbnail made with AI in WThumb?
A 16:9 thumbnail comes back as a high quality Full HD+ JPEG at 2048 by 1152 pixels, and a 9:16 Shorts thumbnail at 1152 by 2048. Both are larger than YouTube's minimum, so you can upload them as they are. You can tick both formats in one submission and get the same brief in both shapes.
ShareX