Midjourney for YouTube thumbnails works, with three extra jobs the model does not do for you: the shape, your face and the words. Its own documentation covers an aspect ratio parameter, a way to bring a person from a reference image into the result, and text inside double quotation marks. What it does not know is what YouTube wants back. As of September 2026, YouTube's Add custom thumbnails page says to try for 16:9 for videos and 9:16 for Shorts, and that the image should be as large as possible.
This post walks the workflow from prompt to upload, then shows where a tool built for thumbnails takes a different route. It is a fork, not a verdict.
The short answer: yes, with three extra jobs
Yes, you can use Midjourney for YouTube thumbnails, and if you already pay for it, the image half of the job is covered. The question is the other half. A thumbnail is a picture with a promise on it, at a fixed shape, often with your face. Midjourney gives you the picture. It leaves you three jobs.
- The shape. Images start square. You set 16:9 or 9:16 with a parameter, then export at a size YouTube accepts.
- Your face. Midjourney's reference features can bring a person from one image into the result. Whether it is recognizably you, every upload, you check by eye.
- The words. Midjourney's documentation says it renders text from a prompt in versions 6 and later. Whether it renders your two words cleanly, you test.
Each job is a round of prompts, and a round of prompts is GPU time on the plan you pay for. The sibling question, can ChatGPT make YouTube thumbnails, lands in the same place: a general image model is doing a different job from a thumbnail tool.
The short fork: Midjourney gives you more knobs and more work; a thumbnail generator gives you fewer knobs and a finished file. Every Midjourney detail below comes from Midjourney's own documentation, read in September 2026; every YouTube detail is dated and linked. For how thumbnail generators compare with each other, see the best AI thumbnail maker roundup and AI thumbnail makers compared on price and size.
Getting the shape: aspect ratio and export size
This section covers the Midjourney setting that matters most and the sizes to export against. Midjourney's Aspect Ratio page says images start as squares and that you change this with the aspect ratio parameter, --ar or --aspect, added to the end of the prompt: --ar 16:9 for a video thumbnail, --ar 9:16 for a Short. The same page says you can set a default aspect ratio in the settings panel; set it once.
The parameter sets proportions, not pixels: that page says aspect ratio is not the same as image dimensions, and the final size depends on the version and upscaler. As of September 2026, YouTube's Add custom thumbnails page recommends a resolution of 3840 x 2160 pixels for videos and 2160 x 3840 for Shorts, with a minimum width of 640 pixels for videos and a minimum height of 640 pixels for Shorts, in formats such as JPG or PNG. The same page also says the file size limit is 2 MB for video thumbnails on mobile (10 MB for podcasts) and 50 MB on desktop, and that your account must be verified.
The export check
Before you upload, read two numbers off the file: the width (or height, for a Short) against the 640 minimum, and the file size against your device's limit. If it falls short, upscale and export again.
Getting yourself in: Omni Reference and the Edit Model
This section covers what Midjourney's own page says about putting a person into an image, and what that means for a channel. The Omni Reference page says the feature allows you to put characters, objects, vehicles, or non-human creatures from a reference image into your Midjourney creations. It is compatible with version 7, takes one image only, and costs 2x more GPU time than regular V7 images. On version 8 the page points you to the Edit Model, which it describes as supporting up to four reference images.
Two lines on that page matter. Under best practices: intricate details like specific freckles or logos on clothing may not perfectly match your reference. A viewer who has watched you for a year knows your face better than a weight slider does, so check the eyes and the jaw, not the lighting. And the weight itself: --ow runs from 1 to 1,000, default 100, and the page advises keeping it below 400 unless you use a very high stylize value, or results may be unpredictable. More weight means more of the reference and less room for the prompt.
What this means across a channel
One thumbnail with a close likeness is a good result. Twenty where the likeness drifts a little each time is a different problem: your row of videos stops looking like one person. Keep the same reference photo, weight and version for a series, and compare each export against the last. A faceless channel skips this section entirely.
Getting the words: text on a Midjourney image
This section covers text: asking for it, testing it, and when to add it elsewhere. Midjourney's Text Generation page says that in versions 6 and later you can get words or phrases to show up by putting them inside double quotation marks. Its best practices say shorter words have a better chance of appearing just right, and that phrases like with the words or written help. If the text comes out wrong, the page suggests Raw mode, a lower Stylize value, or Midjourney's web Editor and Vary Region for minor glitches.
So test it on your own words. Run a prompt with two CAPITAL words, WORTH IT? is a good pair, on your version, and look for three things: every letter present, the letters in order, and a shape you could read at phone size. As of September 2026, YouTube's Thumbnail & title tips page says if you add text, make sure to use a font that's easy to read, and the same page also says to try not to make the design too complex. A word in a stylized script can have every letter and still fail at phone size.
When to add the words in an editor instead
If the test fails twice, stop spending GPU time on letters. Prompt for empty space (the recipe below shows where) and set two or three words in Canva, Photoshop or any image editor with a plain heavy font. A heavily stylized render can also read as AI at a glance; are AI thumbnails a turnoff covers what that costs.
A brief to paste
- Words. WORTH IT? / I USED IT FOR 30 DAYS
- Instructions. The words large and plain, high contrast against the background, nothing behind the letters that fights them.
A Midjourney thumbnail prompt recipe, with a worked example
This section is a prompt structure, a worked example for a review video, and the rounds to expect. Midjourney's Prompt Basics page says short and simple prompts typically generate the best images, to describe what you do want instead of what you don't, and to be clear about subject, medium, environment, lighting, color, mood and composition. For a thumbnail, add one item that page does not list: empty space for the words.
- Subject. One thing, with a number: one small robot vacuum on a wooden floor.
- Framing. Close and low, filling the left two thirds.
- Lighting. One named source: warm window light from the left.
- Palette. Two named colors: teal shadows, orange highlights.
- Empty space. A dark, uncluttered right third for the words.
- Parameters. --ar 16:9 at the very end, or --ar 9:16 for a Short.
The worked example: a budget robot vacuum review
Prompt: photo of one small robot vacuum on a wooden floor, close, low angle, warm window light from the left, teal shadows and orange highlights, the right third dark and empty --ar 16:9. To let Midjourney try the words, add before the parameter: with the words "WORTH IT?" in bold white letters on the right. Expect rounds: one to see whether the subject reads, one to fix the empty space, one for the text test, and an upscale of the keeper.
One caution. As of September 2026, YouTube's Thumbnails policy says not to post a thumbnail that misleads viewers to think they're about to view something that's not in the video, and the same page says a thumbnail that is not appropriate for all audiences but does not violate the Community Guidelines may be removed or age-restrict the video, without a strike. A generated vacuum on fire is a misleading thumbnail if nothing burns in the video. The full route from link to upload is in make a YouTube thumbnail with AI, step by step.
What it costs in time and plan
This section covers the plans as Midjourney's own page listed them when we read it in September 2026, and why GPU time is the budget that matters. The Comparing Midjourney Plans page lists four subscription tiers: Basic at $10 a month, Standard at $30, Pro at $60 and Mega at twice the Pro price, with a 20% discount for an annual plan paid upfront. Each tier carries a monthly allowance of Fast GPU time: 3.3 hours on Basic, 15 on Standard, 30 on Pro and 60 on Mega. Relax mode, with unlimited image generations, starts at Standard. The page says plan availability and features are subject to change, so re-read it before you decide.
GPU time is the real budget because every step in this post spends it: Omni Reference at 2x a regular V7 image, then the text test, the framing rounds and the upscale. Count the recipe's rounds against the Fast hours on your tier; for daily uploads in two formats, that arithmetic decides the plan.
Commercial use
Every tier lists General Commercial Terms, and the footnote says a subscriber is free to use images in just about any way, except that a company making more than $1,000,000 USD in gross revenue per year must purchase Pro or Mega. Read the Terms of Service the page points to; this is a summary, not legal advice.
How to do this in WThumb
Here a thumbnail generator takes a different route. In WThumb you start from a YouTube link or a reference image plus your words, and get back a finished 16:9 or 9:16 thumbnail with the words placed, and you in the design if you added a photo. The reference is a brief, not a template: the idea carries over; the design is new.
- Paste the link of a YouTube video whose idea you like (WThumb fetches only its public thumbnail), or upload a reference image as PNG, JPEG or WebP.
- Type the words, up to three lines, and one line of instructions in plain words: WORTH IT? / I USED IT FOR 30 DAYS, then warm light, teal and orange, the vacuum close and low. See how to describe a thumbnail so you get it right in one generation.
- Add a photo of yourself and you are placed in the design. Leave it out for a faceless thumbnail: the person in the reference is removed and the space is filled in the same style.
- Tick 16:9, 9:16 or both, and ask for one or two designs.
- Wait about thirty seconds for a Full HD+ JPEG.
The @1 trick for a product shot
Attach the product photo as an extra image and point at it: the vacuum from @1, front and center. Up to four extra images work this way, @1 to @4; they are briefs too, not templates.
Two designs, then revise
What WThumb lacks is also clear: no parameters, no style references, no editor. Ask for two designs from one brief and keep the better one. After a thumbnail finishes, the pencil on it takes one change in plain words and WThumb makes a new version. Each revision costs one credit (1 credit = 1 thumbnail), and the new version can be revised again. The comparison is cheap: new accounts get two free credits, no card needed, so make one video's thumbnail both ways.
Midjourney gives you a picture. A thumbnail is a picture with a promise on it, at the right shape, with your face.



