
How to direct a Midjourney image in 2026 (no magic prompts)
Midjourney will not guess your campaign brief. Ask for «a cinematic coffee ad, 8k, award winning» and you get generic coffee, steam and wood. In 22 minutes you learn to write a shot, read a grid and iterate without burning the GPU allowance.
The problem it solves
«Magic» prompts (40-adjective lists from social) push the model toward the same catalog look. The model is not the problem: you have not decided subject, camera, light or format. Without that, Midjourney fills in its default aesthetic.
What you will learn
- Write a shot prompt (subject + camera + light + exclusions)
- Pick --ar for the real destination of the image
- Use stylize and references without losing the subject
- Vary one axis (light or camera) instead of rewriting everything
- Know when to stop and when to move to retouch or Runway
Before you start
- A paid Midjourney account (Basic is enough for this exercise)
- The Midjourney web app or Discord — we will use the web app
- A real job: a product, a space or a character. Not «something pretty»
- The output format (9:16 stories, 4:5 feed, 16:9 hero, A4 print)
- 22 minutes and the discipline not to smash Vary All
Steps 1
Lock the job in a sentence a DP would understand
On a pad, not in Midjourney: «250 g coffee tin, matte aluminum, on a steel table, three-quarter shot at tin height, window light from the left, background out of focus, no extra logos, no hands.» If you cannot say it out loud, the prompt will become an adjective list. There is one subject. «Premium lifestyle with happy people» is not a shot.
What you should see: A 20–40 word sentence with subject, surface, angle, light and at least one exclusion. You have not generated anything yet.
Steps 2
Build the prompt in this order (then stop stacking style)
Write in this order — English or Spanish, be consistent:
[concrete subject] , [shot and angle] , [light] , [background] , [material / lens if it matters] --ar 4:5
Example: «matte aluminum 250g coffee tin, three-quarter shot at tin height, soft window light from the left, steel table, background falloff, 50mm, no extra logos --ar 4:5»
Do not add cinematic, ultra detailed, masterpiece, 8k, Unreal Engine. Those tokens shove the default contest look. If you want film, name grain, contrast and color temperature — or a reference you actually have rights to.
What you should see: A 1–3 line prompt + --ar. No 15-adjective list. You can still wait before hitting Generate.
Steps 3
Generate a grid and pick a direction, not «the prettiest»
Run one generation. You get four stills. Do not pick the shiniest: pick the one that keeps subject, angle and light. If all four are stock coffee with steam, your prompt is still a mood, not a shot — go back to step 1 and lock the angle («from above, tin centered, lid facing camera»). Write half a line on why you picked it («#2: you can read the left-hand light»).
What you should see: A 2×2 grid. One image marked (U or a web click) and a note from you, not a «I like it».
Steps 4
Iterate one axis: light or camera, not both
From the winning image, run a variation or remix changing one thing. Example: «same shot, light now from behind, rim on the tin edge» or «same light, camera 20 cm higher, more table in frame». Do not change product, light, camera and style at once. If you use a reference (--sref or image prompt), use it for color/texture, not to clone someone else’s ad.
What you should see: A second grid that still has the same subject and only changes the axis you asked for. If everything changed, the remix was too loose.
Steps 5
Check the usual failures before you export
Zoom to 100% on: product edges, impossible shadows, ghost type, hands if any, reflections that clone a logo. If the still is for an ad, drop it into a mock of the real format (a stories frame, a web hero). A still that «paints well» on a 400 px grid often dies at real size. If the problem is local (a dent, a corner), use regional edit / inpaint. If the problem is the idea, do not patch: go back to the shot.
What you should see: The winning image viewed large, with 1–3 faults noted or an inpaint already done on one zone. Real format (9:16 etc.) respected.
Steps 6
Close: an 8-image moodboard, not 80, and pick the next job
Keep 6–8 stills that tell the same job (same subject, different light or angle). That is a presentable moodboard. Not an 80-file «just in case» folder. If the next step is motion, take ONE working still and ONE action into Runway. If the next step is copy, do not ask Midjourney for the headline: go back to ChatGPT or Claude with the still as visual context, not as a source of facts.
What you should see: A folder or board with 6–8 images and the prompts pasted. Zero «cinematic 8k» in those prompts.
Real use cases
A product moodboard in an afternoon
One SKU, three lights, two angles, feed --ar. 12 generations, 8 keepers. The client argues about light, not whether it «looks premium».
A space concept (hotel, store, set)
Architecture + time of day + one human figure from behind for scale. No recognizable faces. Present the shot in 16:9.
A base for a 5-second clip
The still already has camera direction. In Runway you ask for one gesture (steam, a 20° turn). Not «a full ad».
Common mistakes
The 80-word novel prompt
Every extra adjective is a vote for the default look. If it does not change the shot, cut it.
Shooting 1:1 «and we’ll crop»
Composition is born in --ar. Cropping 1:1 to 9:16 eats the subject.
Vary All as a method
It is a noise generator. One variable per round is slower on day one and cheaper the rest of the month.
Asking the model for sharp type
The pack, the headline and the legal line are design. Midjourney supplies the still, not the final art.
Close
Directing in Midjourney is deciding a shot and defending it on the grid. Magic prompts only speed the path to the same stock coffee. If you can explain why image 2 won (light, angle, exclusion), you are no longer «rolling generations». You are doing art direction.
Next steps
- Repeat the same subject tomorrow with a different light. Compare the two prompts, not just the JPGs.
- If you need motion, open Runway with one still and one action.
- If the campaign brief is still «modern and warm», go back to the ChatGPT tutorial and translate that into a shot before you spend GPU.
Free guide
Want the 15-tools guide?
We send it free. We will also ping you when the next guide in this path goes out.
Unsubscribe whenever you want. We do not sell the list.
Part of the path Create images with AI