Many YouTube thumbnail tools give you a canvas. You generate something rough, then you nudge the text down four pixels, swap the arrow, make the jacket a bit redder. The description barely matters, because the editor is where the work actually happens.
WThumb does not work that way. One run produces exactly one image. There are no layers, no canvas, no drag-to-nudge. When a thumbnail finishes you can describe a change in plain words and get a new version, and each revision costs a credit too (1 credit = 1 thumbnail); a result that is wrong at the root means describing it again and running again, and each run takes about thirty seconds. So the description is not a hint you throw at the machine. It is the job.
The good news is that describing a thumbnail is a learnable skill with a fixed order to it. Format, reference, who is in it, what is next to them, the background, then the exact words. Write it in that order every time and you stop leaving the decisions that matter most to chance. The same logic runs through how to make a YouTube thumbnail from scratch, where the shape is decided first too.
Why the order of a thumbnail description matters
WThumb is a reference-based editor, not a generator that invents a composition out of a sentence. You do not type "make me a thumbnail about crypto" and hope. You hand it a thumbnail that already works (a YouTube link, or an image you upload) and describe how to change it into yours.
That changes what a good description looks like. You are not painting from nothing. You are giving change instructions against something concrete that already exists. Call it a thumbnail prompt if that is the word you are used to, but it behaves less like a wish typed into a box and more like a list of edits handed to someone working from a photograph. And change instructions have a natural order: the decisions that constrain everything else go first.
Frame comes before layout, because a 16:9 layout and a 9:16 layout put things in different places. The reference comes before the person, because the reference decides the pose. The person comes before the props, because the props sit around them. The background comes near the end because it is what is left over. And the words go last, in their own field, because they are the one thing that must come out exactly as typed.
Skip the order and you get a specific failure: you write three careful sentences about the lighting on a face, then discover the frame you picked put that face somewhere else entirely.
Every run uses up one of your credits, and there is no editor to patch the result afterwards. That is why the first try is the one worth spending time on.
The six-step order to describe a thumbnail
1. Format
Pick the output frame first: YouTube 16:9 or Shorts 9:16. This is a choice you make in the interface, not something you write. It matters more than it sounds. A horizontal reference can be rebuilt into a vertical frame and the other way round: the tool re-stages the layout rather than cropping it, so a wide two-person shot becomes a genuinely vertical composition rather than a squashed one. But the frame decides where everything lands, so decide it before you write a word.
2. Reference
Paste a YouTube link and the thumbnail from that video is pulled in, or upload your own image. PNG, JPEG or WebP, up to 12 MiB for a reference.
Choose the reference for its idea, not its subject. What carries over is the subject, the visual idea and the role each element plays: a face reacting, a product held up, a number doing the work. WThumb designs the layout, the framing, the typography and the colors fresh, so you are not borrowing a composition, you are borrowing a way of thinking. A cooking thumbnail can be an excellent reference for a software tutorial if the idea is right. A reference picked for its niche rather than its idea often gives you a thumbnail that says the wrong thing well.
3. Who is in it
Upload one clear portrait and you are the person in the design. WThumb keeps the energy of the reference rather than its exact pose, so a reference where somebody is shouting and pointing gives you a design with that kind of energy, built around your photo rather than theirs. Photos go up to 8 MiB, and saved photos stay in your library so you can reuse the same one across runs.
Skip the photo and you get a faceless version. The person in the reference is removed entirely and the space they occupied is filled with supporting graphics (arrows, icons, badges) drawn in the reference's own style. This is a deliberate mode, not a gap. If you run a faceless channel, pick a reference with a strong person in it anyway; the space they leave becomes your graphics.
4. What is next to them
Objects, logos, badges, arrows. This is where the "anything else" field starts doing real work, and where extra images come in. Be concrete about placement: left of the face, bottom right corner, behind the shoulder. "Add a laptop" leaves the position to chance. "A laptop in the bottom left corner, screen facing the camera" does not.
5. Background
Say what the background should be, not what it should not be. "Flat dark blue background" is a description. "Remove the messy background" is a deletion request, and deletion requests are the weakest kind. More on why below.
6. The exact words
These go in their own field, and they are the last thing you write because they are the only thing in the whole request that has one correct output. Everything above is interpretation. The text is not.
The two description fields, and why they are separate
WThumb splits your description into two boxes, and the split is worth understanding before anything else.
- "The words on it" is the literal text that gets rendered onto the thumbnail. Up to three lines. Whatever you type here is what appears.
- "Anything else" is everything else described in plain words: colors, clothing, backgrounds, logos, objects, placement, mood.
The split exists because of an ambiguity that ruins thumbnails in single-box tools. If you write "make the text say I QUIT" into one combined field, something has to decide whether "make the text say" is an instruction or part of the headline. Two fields remove the guess. The words field is quoted; the other field is interpreted.
The words field is quoted. The other field is interpreted. Everything you type in the first one is a promise about the output.
Inside the "anything else" field there are still capitalisation rules worth knowing, because you will sometimes need to talk about text there too. A line typed in CAPITALS is treated as literal text to render. A sentence like "write BROKE instead of poor" is read as an instruction about the text, not as something to print. And one word on its own is always treated as literal. So if you write just the word BROKE, expect to see BROKE on the thumbnail, not a mood.
Practical consequence: keep your headline in the words field, and keep your sentences in the other field lowercase unless you genuinely mean "render this".
Attaching extra images with @1 and @2
You can attach up to four extra images alongside the reference and your photo, then point at them in the "anything else" field with @1, @2, @3, @4. Same formats: PNG, JPEG or WebP. File size is rarely the issue here: a logo or a screenshot is normally a small fraction of the 8 MiB allowed for a photo or the 12 MiB allowed for a reference.
Attachments cover what words cannot pin down. Your channel logo. A specific product box. A sponsor badge. A screenshot from the video. A particular car. You can describe "a red sports car" and get some red sports car; you cannot describe your red sports car, so attach it.
Point at attachments with placement, not just presence:
- Weak. "add @1"
- Better. "put @1 in the top right corner, about the size of the face"
- Better still. "put @1 in the top right corner at roughly a quarter of the frame height, with a thin white outline so it separates from the background"
The scale reference is the part people forget. Images do not carry an obvious intended size, so if you do not say how big, you are rolling dice on how big.
A worked example: a thumbnail description built line by line
A video about quitting a software job to go full-time on a channel. Start with the sloppy version most people would type, then build the real one.
The sloppy version, all in one blob: "a thumbnail about me quitting my job, make it look shocking and professional, put my face in it and my logo, text says I QUIT." Every word of that is a wish rather than an instruction. Shocking how. Professional how. Face where. Now the real thing.
Format. YouTube 16:9, because it is a long-form video.
Reference. A YouTube link to a thumbnail with the structure I want: person on the right at about half the frame height, big two-line text stacked on the left, a hard color block behind them. I am not borrowing the topic. I am borrowing that staging.
Photo. One clear portrait of me. It will take the reference person's pose and lighting, which is the point: I picked a reference where the person looks genuinely alarmed.
Extra images. @1 is my channel logo. @2 is a photo of my old office badge.
The words on it, typed into the words field, two lines:
- I QUIT
- AFTER 9 YEARS
"Anything else", built up one clause at a time, in order:
- Clothing. "navy hoodie instead of the shirt": settles what I am wearing rather than leaving it to the reference.
- Props and placement. "put @2 in my right hand, held up towards the camera at about the size of a phone": a real object, a real position, a real scale.
- Branding. "put @1 in the bottom right corner, small, roughly a tenth of the frame height": small and specific beats "add my logo".
- Background. "flat deep red background with a soft dark vignette at the edges": what it should be, not what to strip out.
- Text treatment. "the text in heavy white with a thin black outline, left aligned": this is a description of the words, so it belongs here, in lowercase, not in the words field.
Read the finished request end to end and it is boring, which is the sign it is right. Every sentence names a thing, a place and a size. Nothing is left for the tool to invent except the parts I genuinely do not care about.
What makes a thumbnail request fail
Vague adjectives
"Clean", "professional", "eye-catching", "modern", "epic", "aesthetic". These words feel like direction and carry none. They describe your reaction to a finished image, not any property of it. Every one of them can be replaced with something concrete: "clean" usually means one background color and no more than two objects; "professional" usually means restrained colors and no outline glow; "epic" usually means a low camera angle and strong side lighting. Write the concrete version.
The test: if two people read your sentence and would produce visibly different images, the sentence is not finished.
Two conflicting requests in one line
"Minimal but with lots of energy." "Dark background but bright and cheerful." "Big text but don't cover the face." These are not descriptions, they are unresolved arguments, and handing an argument to the tool means one side wins at random.
Resolve it yourself before you run. If the text must be big and the face must stay clear, say where each one lives: "text across the top third, face in the lower right, no overlap." Now there is no conflict, just a layout.
Describing what to remove
This is the subtlest one. "Remove the clutter." "Get rid of the second person." "No text at the bottom." "Delete the arrows."
The problem is that removal leaves a hole and says nothing about what fills it. Describe the end state instead. Not "remove the clutter" but "plain gray background behind me". Not "get rid of the second person" but "just me in frame, centered". Not "delete the arrows" but "nothing between the text and the face". You will notice the fixed versions are all positive statements about what exists in the final image, which is exactly what a description should be.
Loading a single run with a redesign
If your "anything else" field runs to a dozen unrelated changes, you are not describing a thumbnail any more, you are describing three of them. A run produces one image, and the more instructions compete for attention, the more likely one of the ones you actually cared about gets treated as optional. Cut it back to the changes that decide whether the thumbnail works. If the list is still long, the problem is the starting image: a closer reference does more than a longer request.
Not looking at the reference while you write
Keep the reference open in another window while you write. It is easy to start describing changes to a thumbnail you are remembering rather than the one you actually attached. The reference is the ground your instructions stand on. Every sentence you write should be checkable against the image you actually attached.
A checklist before you run a thumbnail
New accounts get two free credits, no card required, and they do not renew. So the first run is worth spending five minutes on. A quick pass before you commit:
- Frame chosen early. 16:9 or 9:16, decided before the reference.
- Reference chosen for structure. Pose, text position and color weighting are what you are borrowing.
- Photo decision made on purpose. Either a clear portrait, or deliberately none for the faceless version.
- Words field holds only the words. No instructions, no "make it say", three lines maximum.
- Every object has a place and a size. Including anything you attached as @1 or @2.
- No adjective left that two people could read differently. Replace it or delete it.
- Nothing phrased as a removal. Every line describes what is in the final image.
Then run it, and give it the thirty seconds or so it takes. If it comes back wrong, change the one clause that produced the wrong thing rather than starting the description over. Make that clause more specific, and run again. The description is the work, but it is work that gets faster every time, because the order stays the same.



