MEMO · LOCAL · AUGUST 2026

Make AI videos on your own computer

9 min read · both machines · every number sourced below

The short version: the same kind of model behind the AI video you are paying a monthly fee for can be downloaded and run on your own machine, for nothing. The software is WanGP, it is free and open source, and once it is installed it works with the Wi-Fi switched off. You need either a Windows PC with an NVIDIA card of 8 GB VRAM or more (12 GB is the sensible entry point) or an Apple Silicon Mac with 16 GB or more of unified memory. Installing it is genuinely horrible, so you do not: you paste the prompt on this page into Claude Code and let it do the install. That subscription is the only thing here that costs money, and only until the install finishes.

two dogs, a fairground
a lift full of elephants
a man in a gallery
FREE · OFFLINE
Three clips, no camera, no subscription. The connection is off the whole time.

What you actually do, once it is working

Worth getting the shape of this clear before any of the hardware talk, because the day-to-day is much smaller than the setup suggests.

You double-click a launcher. A black window opens and stays open, because it is the engine. After a moment a page opens in your browser at localhost:7860. On that page there is a model dropdown, a big empty text box, a couple of number fields and a Generate button.

You type one sentence into the box. Something like two cats playing basketball on a city court, sunny, shallow depth of field. You set the frame count to 49, which is about two seconds, and the step count to 20. You click Generate.

Some minutes later a video file appears in the outputs folder inside the Wan2GP directory, and in the “Generated videos” panel on the page. That is the whole loop. No credits, no queue, no monthly cap, and nothing uploaded anywhere.

LOCALHOST IS NOT A WEBSITE
The interface looks like a web app, which trips people up. localhost literally means this computer here. You can turn your Wi-Fi off and keep generating. The only moments that genuinely need the internet are the install, and the first time you pick a model, when it downloads that model to your disk.

Can your machine do it?

On both platforms the answer comes down to almost exactly one number: how much memory your graphics processor can reach. On a PC that is VRAM, soldered onto the graphics card, and it is small and expensive. On an Apple Silicon Mac there is one shared pool, which is what “unified memory” means, so a 48 GB MacBook has more usable graphics memory than any consumer NVIDIA card. The problem on the Mac is not size, it is speed.

Windows PC with an NVIDIA card

Press Ctrl + Shift + Esc, go to Performance, click GPU, and read “Dedicated GPU memory”. That number is your VRAM.

Your VRAMVerdictWhat you can actually do
Under 6 GBNoNot worth the pain. Even the small models spill out of the card and you will spend your life reading out-of-memory errors.
8 GB (RTX 3060 Ti, 4060, 2070)Just aboutWan 2.1 T2V 1.3B at 480p, plus quantised medium models. Use --profile 5 and keep clips short.
12 GB (RTX 3060 12 GB, 4070, 5070)FineMedium 5B-8B models at 480p comfortably, 720p on short clips. The sensible entry point.
16 GB (RTX 4080, 5080, 4060 Ti 16 GB)GoodThe 14B models in quantised form, 720p, clips of several seconds. Enough for real work.
24 GB and up (3090, 4090, 5090)ComfortableNearly everything, including the big 22B models at full precision. Memory stops being the limit.

AMD and Intel cards are not realistically supported. And system RAM matters more than people say: WanGP’s whole trick is keeping part of the model in system RAM and streaming it into the card as needed, which is how a 16 GB card runs a model that does not fit in 16 GB. That makes 16 GB of system RAM a real bottleneck and 32 GB a genuine upgrade, separately from your graphics card.

Mac with Apple Silicon

Apple menu, About This Mac. Note the chip, the memory and the macOS version. You need M1 or newer and macOS 13 or later; Intel Macs are out.

Unified memoryVerdictWhat you can actually do
8 GBNoIt will not fit. The smallest useful video model already takes nearly all of it, and macOS has to live somewhere.
16 GBJust aboutWan 2.1 T2V 1.3B, 480p, 2 to 3 seconds. Close every other app. It works, but it is handicraft.
24 GBFineThe 1.3B comfortably, plus Wan 2.2 TI2V 5B or HunyuanVideo 1.5 at 480p. Clips of 3 to 5 seconds.
36 to 48 GBGood8B to 14B models and 720p. The first tier where results start looking like what you see online.
64 GB and upComfortableAlmost everything, including the big 22B models. Memory is no longer the limit. Patience is.

The chip affects speed, not what is possible. An M4 Max is roughly four to six times faster than a base M1 for this work. One technical wrinkle: M1 and M2 do not handle a number format called bfloat16, so the software falls back to another one. On M3 and M4 that does not arise.

Budget disk space too. A single video model weighs between 2 GB and 25 GB, and WanGP downloads the text encoder and video decoder alongside it. 50 GB free is the minimum for trying one model, 100 to 150 GB for two or three, 250 GB and up to explore properly. Keep 25 GB of headroom on top. On Windows, keep the models off any path with spaces or accents in it and out of OneDrive-synced folders, because both cause strange failures later.

PC versus Mac: the honest gap

Plenty of tutorials say “it runs on a Mac” without mentioning the price of that. The thing at the centre of it is called CUDA, an NVIDIA technology that all modern AI was built on, and which researchers write their most aggressive optimisations in. Apple’s equivalent is MPS. It works, it is decent, and it does not have ten years of optimisation from the entire planet behind it.

AccelerationNVIDIA PCApple Silicon
SageAttention / Flash Attention (the biggest win, 2-3x)YesDoes not exist
Triton GPU kernel compilerYesDoes not exist
torch.compileYesForce-disabled
NVFP4 / INT8 CUDA quantisationYesIncompatible
Pinned memory and async transfersYesDisabled
SDPA (PyTorch standard attention)YesYes, and this is the one you use
Lots of available memory8-16 GB, 24 GB at the topUp to 128 GB
THE NUMBER TO REMEMBER
At identical settings, expect an Apple Silicon Mac to be 5 to 15 times slower than a PC with a recent graphics card. A generation that takes one minute on a gaming PC takes 15 to 30 minutes on a MacBook. It is not a bug and no setting fixes it.

The Mac support is not a hack, though, which is worth saying. It is built into WanGP: a dedicated Apple Silicon module detects your chip, presents itself as an NVIDIA card to the rest of the program, disables everything that does not exist on a Mac, and routes the maths to Apple’s GPU. The dependency list contains lines written specifically for macOS.

So: use the PC if you have one with a decent NVIDIA card, because it is the right tool by a wide margin. Use the Mac if that is what you own and you want to learn and make short clips without paying anyone or uploading anything. Rent a cloud GPU by the hour if you need ten videos a day and do not own the hardware.

How a computer invents a video

You do not need this to install anything, but every setting further down stops being a mystery button once you have it.

You know the grey fizzing snow of an old untuned television? In computing that is called noise: pure randomness, with no picture in it at all. The AI starts exactly there. It fills the screen with random snow, then asks itself a strange question: if this image were really a cat running on a beach hidden under noise, what would it look like with a tiny bit of that noise removed?

It removes a little. It asks again. It removes a little more. Twenty or thirty times. Each pass the picture gets less blurry and more “cat on a beach”, and at the end there is no noise left, only the image. A sculptor looks at a block of marble and says the statue is already in there, and they are removing what is not part of it. Same thing, with noise instead of marble. Your sentence is what tells it which statue to look for.

Video makes that much more expensive. A video is not one image, it is 49 images in a row, and if the AI made them one at a time the cat would change colour every frame. So it denoises all 49 frames at once, constantly looking from one frame to the next to stay consistent. That is why video is not “49 times slower” than an image but often considerably worse, and it is why video eats memory the way it does: every frame has to be held at the same time.

Which teaches you four things you will use in a minute:

  • More denoising steps means cleaner output and proportionally longer waits.
  • More frames means longer waits and more memory, at the same time.
  • Higher resolution makes every single step cost more.
  • Your prompt is the only compass the model has for the whole process.

The install: one prompt, both machines

Here is the part everyone gets stuck on. The software is free and the models are free, but installing them means stacking a lot of tools that all have to agree with each other, and the internet is full of contradictory advice about which versions those should be.

So do not do it yourself. Claude Code is an AI that lives in your terminal: you tell it in plain English what you want, it reads your files, writes the commands, runs them with your permission, and fixes its own mistakes. If you have not met it before, the ECC config memo is a gentler first introduction to what it is and how it behaves.

It needs a Claude Pro, Max, Team or Enterprise account. The free plan does not include Claude Code. That subscription is the only money in this entire page, and you only need it for the install.

1

Open a terminal

PC: press the Windows key, type “PowerShell”, press Enter. The normal one, not the administrator one.

Mac: press Cmd + Space, type “Terminal”, press Enter.

A terminal is just a window where you type orders instead of clicking. You will use it about four times in this whole process. The universal stop button is Ctrl + C on both platforms, and if the cursor blinks and nothing appears to happen, it is usually working silently. Some steps here take twenty minutes.

2

Install Claude Code

Windows, into PowerShell:

PowerShell
irm https://claude.ai/install.ps1 | iex

If Windows blocks the script with an execution policy error, run this first. It affects only your own account, only this session:

PowerShell, only if blocked
Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass

Mac, into Terminal:

Terminal
curl -fsSL https://claude.ai/install.sh | bash

Then check it landed with claude --version. If you get “command not found” or “not recognized as a cmdlet”, close the terminal window completely and open a fresh one. That fixes it about 90% of the time.

3

Make a working folder and start it

Windows:

PowerShell
mkdir $HOME\AI; cd $HOME\AI; claude

Mac:

Terminal
mkdir -p ~/AI && cd ~/AI && claude

On the first start it opens your browser to sign in to your Claude account. Approve, come back to the terminal, done for good.

4

Paste the prompt for your machine

Set aside one to two hours, most of it downloading. On a Mac, plug into the mains first. The two prompts are not interchangeable — use the one for your platform.

Windows + NVIDIA

Paste this whole thing into Claude Code
I am on Windows with an NVIDIA graphics card. I want to install WanGP
(the Wan2GP project) to generate AI videos locally on my GPU via CUDA.
I am a beginner: explain each step in simple English before you do it.

Do the following steps in order:

1. DIAGNOSIS
   Check my exact GPU model, my VRAM, my driver version, my CUDA version,
   my system RAM and my free disk space. Run nvidia-smi and show me the
   output. Tell me honestly whether my machine can run this software and
   which models are realistic for me. If I have less than 6 GB of VRAM or
   less than 50 GB of free disk, stop and warn me.

2. PREREQUISITES
   Install only what is missing:
   - Python 3.10 or 3.11 (NOT 3.12 or newer - some dependencies break)
   - git
   - ffmpeg
   Add them to PATH. Do not touch any Python that another program depends on.

3. DOWNLOAD THE SOFTWARE
   Clone https://github.com/deepbeepmeep/Wan2GP into a folder with no
   spaces and no accents in its path, and NOT inside OneDrive.

4. PYTHON ENVIRONMENT
   Create a Python 3.10/3.11 virtual environment called .venv inside the
   Wan2GP folder and activate it. Everything below installs in there,
   never globally.

5. PYTORCH - THIS IS THE STEP THAT GOES WRONG
   Look up my GPU architecture first. If I have an RTX 50-series card
   (Blackwell, sm_120), I need a PyTorch build for CUDA 12.8 or newer -
   older builds fail with "no kernel image is available for execution on
   the device". For older cards, a CUDA 12.4+ build is fine.
   Install the matching torch, torchvision and torchaudio from the correct
   PyTorch index URL. Then verify torch.cuda.is_available() returns True
   AND that a small tensor operation actually runs on the GPU. Show me
   both results. If either fails, stop and explain why.

6. DEPENDENCIES
   Install requirements.txt. If a package fails to build:
   - do not abandon the whole install;
   - tell me which one, what it is for, and whether WanGP works without it;
   - continue with the rest.

7. SPEED EXTRAS (optional, do this only after everything already works)
   Try to install SageAttention and Triton for Windows - they give the
   biggest single speed-up (roughly 2x). They need matching CUDA and
   compiler versions and they often fail to build. If they do not install
   cleanly in a reasonable time, skip them and tell me: WanGP runs fine
   without them using --attention sdpa.

8. LAUNCH SCRIPT
   Create start-wangp.bat next to the Wan2GP folder that:
   - changes into the Wan2GP folder
   - activates .venv
   - runs: python wgp.py --attention sdpa --profile 4 --open-browser
   If SageAttention installed successfully in step 7, use
   --attention sage2 instead of sdpa and tell me you did.
   Choose the profile based on MY VRAM: profile 5 for 8 GB, profile 4 for
   12-16 GB, profile 3 for 24 GB or more. Explain your choice.

9. FIRST RUN
   Run the script with me and fix errors one at a time until the web
   interface opens on http://localhost:7860 . Show me the error messages
   in plain text and explain what you are fixing.

10. SUMMARY
    When everything works, write me a HOW-TO-LAUNCH.md file, max 10 lines,
    explaining how to start the software later and which settings suit MY
    VRAM specifically.

General rules:
- Explain each command in one sentence before running it.
- If you have to choose between "fast" and "it works", choose "it works".
- Do not modify anything outside my AI working folder.
- If you are stuck more than three attempts on the same error, stop and
  explain the problem in simple English instead of trying again.

Mac with Apple Silicon

Paste this whole thing into Claude Code
I am on a Mac with an Apple Silicon chip. I want to install WanGP
(the Wan2GP project) to generate AI videos locally, on the Mac GPU via
MPS. I am a beginner: explain each step in simple English before you
do it.

Do the following steps in order:

1. DIAGNOSIS
   Check my chip, my amount of unified memory, my macOS version and my
   free disk space. Tell me honestly whether my machine can run this
   software and which models are realistic for me. If I have less than
   16 GB of memory or less than 50 GB of free disk, stop and warn me.

2. PREREQUISITES
   Install only what is missing:
   - the Xcode Command Line Tools
   - Homebrew
   - git, ffmpeg
   - Python 3.11 (via Homebrew)
   Do not touch the Python that ships with macOS.

3. DOWNLOAD THE SOFTWARE
   Clone https://github.com/deepbeepmeep/Wan2GP into ~/AI/Wan2GP

4. PYTHON ENVIRONMENT
   Create a Python 3.11 virtual environment called .venv inside
   ~/AI/Wan2GP and activate it. Everything below installs in there,
   never globally.

5. MAC BUILD OF PYTORCH
   Install PyTorch for macOS:
   pip install torch torchvision torchaudio
   Use NO CUDA index. Then verify torch.backends.mps.is_available()
   returns True and show me the result. If it is False, stop and explain
   why.

6. DEPENDENCIES
   Install requirements.txt. WARNING: this file is written for
   Windows/NVIDIA first. Its first line is a CUDA index, and it contains
   packages that do not exist on Mac. Most lines already carry
   platform_system == "Darwin" markers, but if a package still refuses
   to install:
   - do not abandon the whole install;
   - tell me which one, what it is for, and offer to skip it;
   - continue with the rest.
   NEVER install: sageattention, flash-attn, triton, xformers,
   onnxruntime-gpu, decord, or any package whose name contains cuda,
   cu12, cu13 or nvidia. They do not exist on Mac and will fail.

7. LAUNCH SCRIPT
   Create ~/AI/Wan2GP/start-mac.sh that:
   - changes into the script's folder
   - activates .venv
   - runs: python wgp.py --attention sdpa --profile 4 --open-browser
   Make it executable with chmod +x.
   Do NOT use --compile, and do NOT use --attention sage/sage2/flash:
   they do not exist on Mac and will crash.

8. FIRST RUN
   Run the script with me and fix errors one at a time until the web
   interface opens on http://localhost:7860 . Show me the error messages
   in plain text and explain what you are fixing.

9. SUMMARY
   When everything works, write me a ~/AI/HOW-TO-LAUNCH.md file, max
   10 lines, explaining how to start the software later and which
   settings suit MY amount of memory specifically.

General rules:
- Explain each command in one sentence before running it.
- If you have to choose between "fast" and "it works", choose "it works".
- Do not modify anything outside ~/AI.
- If you are stuck more than three attempts on the same error, stop and
  explain the problem in simple English instead of trying again.
RED IS NOT DISASTER
The terminal prints an enormous amount of text during this, including yellow and sometimes red warnings that block nothing at all. Do not panic and do not close anything. Letting Claude Code read and sort that output is precisely the job you are giving it. If something is genuinely broken it will tell you in plain English.

Roughly what happens, and how long each phase takes:

PhaseTypical durationWhat you see
Diagnosis1 minuteWhat you have, and whether it is viable
Prerequisites5-25 minutesLots of scrolling text. On Mac the Xcode Command Line Tools may open a system window: accept it. On PC, Python or git installers may raise a UAC prompt: accept it.
Downloading the code1-3 minutesWanGP lands in your AI folder
PyTorch and dependencies10-40 minutesThe longest and noisiest phase. Yellow warnings are normal.
Speed extras (PC only)0-30 minutesSageAttention and Triton either build or they do not. Either outcome is fine.
First run2-10 minutesPossibly a few rounds of fixing. That is expected.

Three rescue prompts

These cover the large majority of ways it goes wrong.

If a package refuses to install
The install is stuck on this package. Explain in simple English what it is for, whether WanGP can work without it on my platform, and if so, skip it and carry on with the installation.
If you want a clean restart
I want to start over cleanly. Delete only the Wan2GP folder and its virtual environment, without touching Python, git, my GPU driver, or anything else on this computer. Then resume the installation from step 3.
If none of it makes sense any more
Stop. Explain to me in very simple English, as if to someone who knows nothing about this: where we are, what just happened, whether there is a problem, and what the next thing I need to do is.

Launching it, every time after that

From here you do not need Claude Code day to day. On Windows, double-click start-wangp.bat. A black window opens — leave it open, it is the engine — and the browser follows on its own. On a Mac:

Terminal
cd ~/AI/Wan2GP && ./start-mac.sh

What the launcher options mean:

OptionIn plain English
--attention sdpaUse the standard maths. On a Mac this is the only one available. On a PC, if SageAttention installed, use --attention sage2 instead: roughly twice as fast, for free.
--profile 4Save memory even if it is slower. A good default on both.
--open-browserOpen the browser when the engine is ready.
--preload 8000 (PC)Keep 8 GB of the model resident in VRAM instead of streaming it every step. A big speed win if the card has room; lower it on out-of-memory errors.
--perc-reserved-mem-max 0.3 (Mac)Add this if your Mac stutters or freezes.
--advancedShow every advanced setting in the interface. Add later.

The profiles are strategies for where the model lives while it works, and lower means more of it in fast memory: --profile 3 for 24 GB VRAM or a 64 GB+ Mac, --profile 4 for 12-16 GB VRAM and most Macs, --profile 5 for 8 GB or anything that keeps crashing.

STOP IT PROPERLY
Closing the browser tab is not enough. Go back to the terminal or the black window and press Ctrl + C, or the program keeps running in the background holding your memory and your graphics card.

One warning about the first start: the first time you select a model, WanGP downloads it, and depending on the model that is 2 to 25 GB. The interface can look frozen the entire time. Watch the terminal window instead, where the download progress actually is.

The settings for your memory

Four dials decide 90% of the result and 100% of the waiting: the model, the resolution, the frame count and the step count. Landmarks worth memorising: 25 frames is about 1 second, 49 is 2 seconds, 73 is 3 seconds. For steps, 15 is fast, 20 is balanced, 30 is careful, and past 30 the gain is nearly invisible. Resolution is the expensive one, because doubling the width and height multiplies the work by four, not two.

There is a fifth, less discussed: Guidance, which decides how literally the model obeys your prompt. 3 to 5 interprets freely, 7 to 10 is the default and the right answer in almost every case, and 12 and up gives literal obedience that can look oversaturated and stiff.

Starting settings, Windows PC

VRAMModelResolutionFramesStepsProfile
8 GBWan 2.1 T2V 1.3B480p49205
12 GBWan 2.2 TI2V 5B480p49-73204
16 GBWan 2.2 14B (quantised), HunyuanVideo 1.5480p to 720p73-9720-254
24 GB+Almost anything, including LTX-2.3720p97+25-303

Starting settings, Apple Silicon Mac

MemoryModelResolutionFramesStepsClip length
16 GBWan 2.1 T2V 1.3B480p25-4915-201-2 s
24 GBWan 2.1 1.3B, then Wan 2.2 TI2V 5B480p49-73202-3 s
36-48 GBWan 2.2 T2V/I2V, HunyuanVideo 1.5480p to 720p73-9720-253-4 s
64 GB+Almost anything, including LTX-2.3 distilled720p97+25-304 s and up
THE RULE THAT SAVES YOUR EVENING
Find your idea small, then redo it big. Test the prompt at 480p, 25 frames, 15 steps: within a couple of minutes you know whether the composition works. When it does, keep the same seed and re-run with everything turned up. Never hunt for an idea at high quality, because you will spend hours producing something you throw away.

Two more habits that pay for themselves immediately. Note your seeds, because the seed is the number the randomness starts from and the same prompt plus the same seed gives the same video, which is how you reproduce a good result or change exactly one thing. And change one setting at a time, or you will never know what improved it.

Which model to pick

WanGP gives you access to more than two hundred models. About a dozen matter. The three families: Wan 2.1 / 2.2 (1.3B to 14B, the generalist family and the most mature, where you start), HunyuanVideo 1.5 (8.3B, a notch better image quality, with a lighter 480p version), and LTX-2 / LTX-2.3 (22B, generates picture and sound together and handles long clips, but 22 billion parameters is not for modest hardware).

You will also meet words like Lightning, FusioniX, FastWan, Distilled, Self-Forcing, GGUF, NVFP4 and Nunchaku. These are not new families, they are accelerated or compressed versions of an existing model. On a PC they are your friends: a quantised 14B is how a 16 GB card runs a model that should not fit. On a Mac, anything called NVFP4 or Nunchaku will not work, because those formats need NVIDIA cards. When in doubt take the plain or bf16 version.

Quantisation, by the way, is fitting a model into less memory by writing its numbers with less precision. It is going from a RAW photo to a JPEG: much lighter, very slightly less beautiful, and in 95% of cases you cannot see the difference.

Or just ask, from inside the Wan2GP folder:

Claude Code
Look at my available video memory and at the list of models in WanGP's defaults/ folder. Recommend the three video models best suited to MY machine, from lightest to heaviest, explaining for each: its approximate size, what it can do, and the realistic generation time I should expect. Rule out anything my hardware cannot run.

Writing a prompt that works

An average model with an excellent prompt beats an excellent model with a sloppy one, and this is the only part of the whole business where you improve enormously without buying hardware. Build it like a brief for a camera operator, in this order: the subject (an old fisherman, not “a person”), the action (mending a net, slowly — with no verb you get an animated photo), the setting (on a wooden pier at dawn), the light (soft golden light, light fog — the most powerful lever you have), the camera (slow dolly-in, shallow depth of field), and the style (35mm film, cinematic, muted colors).

All six parts, one sentence
An old fisherman mending a net, slowly, on a wooden pier at dawn, soft golden light, light fog over the water, slow dolly-in, shallow depth of field, 35 mm film, cinematic, muted colors.

And a negative prompt that works in almost every case:

Negative prompt
blurry, low quality, distorted face, deformed hands, extra fingers, watermark, text, jittery motion, flickering

Three mistakes to skip past. The one-word prompt (“a dragon”) makes the model invent everything, and what it invents is the average of everything it has seen, which is bland. The shopping list of twenty adjectives (“beautiful, amazing, stunning, 8k, ultra realistic, masterpiece”) dilutes the words that do matter. And asking for five actions in a two-second clip (“he opens the door, walks in, sits down, turns on the lamp and smiles”) is impossible in 49 frames, so the model blends it into mush. One shot, one action. To tell a story, chain clips.

Write prompts in English, because these models were trained almost entirely on English descriptions and other languages give noticeably less faithful results. If your English is shaky, ask Claude to translate and improve the description first.

How long it actually takes

Tutorial videos show generations finishing in 45 seconds. Here is what to expect, as orders of magnitude rather than promises, for a 2-second clip at 480p and 20 steps:

MachineSmall model (1.3B)Medium model (5B-8B)
PC · RTX 3060 12 GB60-120 s4-8 min
PC · RTX 4070 / 507040-70 s2-5 min
PC · RTX 4080 / 5080~30 s1-3 min
PC · RTX 4090 / 509015-25 s45 s - 2 min
Mac · Air M1/M2, 16 GB20-40 minOut of reach
Mac · M3 Pro, 18-24 GB10-20 min45-90 min
Mac · M4 Pro, 24-48 GB5-12 min25-50 min
Mac · M4 Max, 48-128 GB3-8 min12-30 min

Those PC numbers assume plain --attention sdpa. If SageAttention installed and you run --attention sage2, roughly halve them.

And what multiplies the wait, which is the part people miss:

If you changeTime is multiplied by
480p to 720pabout 2.5
2 s to 4 s of videoabout 2 (and memory explodes)
15 to 30 stepsexactly 2
1.3B model to 14B model5 to 10

These multiply together. Going from “480p, 2 s, 15 steps, 1.3B” to “720p, 4 s, 30 steps, 14B” is not a bit slower, it is on the order of fifty to a hundred times slower.

THE STRATEGY THAT MAKES A SLOW MACHINE PLEASANT
Work in batches, not live. Prepare five prompts during the day, stack them with the “Add” button in the queue, start the lot before you go to bed, and look at the results over breakfast. Seen that way a Mac’s slowness stops mattering: you do not experience it, you make it work while you sleep.

When it breaks

It will break. That is not bad luck, it is the norm in this field. The universal reflex: select the error message in the terminal, copy it, start Claude Code and paste it with “Here is the error I get when running WanGP. Explain in simple English what is wrong and fix it.” That resolves the large majority of cases without you needing to understand anything.

SymptomLikely causeWhat you do
“out of memory” / “CUDA out of memory”You asked for more than your memory holdsReduce in this order: frame count, then resolution, then a smaller model. Close other apps. On PC try --profile 5 or lower --preload.
“no kernel image is available for execution on the device” (PC)Your PyTorch build does not know your card’s architecture. Classic on RTX 50-series.Reinstall PyTorch from a CUDA 12.8+ index. Paste the error to Claude Code and say which card you have.
torch.cuda.is_available() is False (PC)CPU-only PyTorch build, or the driver is too oldAsk Claude Code to check the driver with nvidia-smi and reinstall the correct CUDA build.
“CUDA not available” (Mac)Some code thinks it is talking to an NVIDIA cardGive the error to Claude Code. It is a known case with a known Mac-side fix.
“sageattention not found” / flash / tritonYou launched with an attention mode you do not haveChange the launcher to --attention sdpa. On a Mac those three never exist.
localhost:7860 will not openStill starting, or it crashed on launchRead the terminal window, not the browser. The real message is there.

One safety rule that outlives this guide: never paste a command into a terminal when you do not know where it came from. A command can delete files. Everything here comes from official documentation and is explained, and from the install prompt onward it is Claude Code writing the commands and asking your permission before each one. If you are working this way often, the habits that keep a Claude Code session cheap are worth twenty minutes.

What it actually costs

WanGP is free. The models are free. Electricity and disk space are real but small. The one line item is a Claude subscription for the install assistant, and only for as long as the install takes. After that you generate forever, paying nothing, with no internet connection at all.

The honest counterweight, so nobody arrives disappointed: a 16 GB card runs out of memory constantly, Windows has its own creative ways of breaking, and a MacBook is 5 to 15 times slower than a gaming PC and no setting changes that. Both platforms genuinely work, for free, and both are more than enough to learn, experiment and make short clips. Neither is a free replacement for a professional cloud pipeline, and anyone telling you otherwise is selling something.

And your first video will probably be disappointing. That is expected. The 1.3B model is the smallest in the family: faces are wrong, hands are nightmarish, motion is mushy. It is not your machine and it is not your installation. You have just proved the whole chain works, which was the only goal. From there you move up.

Questions people actually ask

Can I really generate AI video on my own computer for free?

Yes. WanGP (also written Wan2GP) is free and open source, and the video models it runs are free to download. Once installed it works with no internet connection at all, because the interface runs at localhost:7860 on your own machine. The only money involved is a Claude Pro or Max subscription if you use Claude Code to do the installation for you, and you only need that during the install itself.

How much VRAM do I need for AI video generation?

On a Windows PC with an NVIDIA card: under 6 GB is not worth attempting, 8 GB runs the small Wan 2.1 1.3B model at 480p, 12 GB is the sensible entry point for medium 5B-8B models, 16 GB runs quantised 14B models at 720p, and 24 GB or more removes memory as a limit. AMD and Intel cards are not realistically supported. System RAM matters separately: 16 GB is a real bottleneck and 32 GB is a genuine upgrade, because WanGP streams part of the model from system RAM into the card.

Does this work on a Mac?

On Apple Silicon, yes, and the support is built into WanGP rather than being a community patch. There is a dedicated Apple Silicon module that detects your chip, disables the NVIDIA-only features and routes the maths to Apple's GPU through MPS. You need an M1 or newer, macOS 13 or later, and at least 16 GB of unified memory. 8 GB will not fit. Intel Macs are out.

How much slower is a Mac than a gaming PC for AI video?

At identical settings, expect 5 to 15 times slower. A generation that takes one minute on a PC with a recent NVIDIA card takes 15 to 30 minutes on a MacBook. This is not a bug and no setting fixes it: the accelerations that give the biggest speed-ups, SageAttention, Triton, torch.compile and NVIDIA-only compressed model formats, do not exist on Apple Silicon. The workaround is to batch overnight rather than iterate live.

How long does one AI video take to generate?

For a 2-second clip at 480p and 20 steps: roughly 60 to 120 seconds on an RTX 3060 12 GB, 40 to 70 seconds on a 4070 or 5070, and 15 to 25 seconds on a 4090 or 5090, using the small 1.3B model. Medium 5B-8B models take 45 seconds to 8 minutes on the same cards. On a 16 GB M1 or M2 Air the small model takes 20 to 40 minutes. Doubling to 720p multiplies by about 2.5, doubling the length by about 2, and going from 1.3B to 14B by 5 to 10, and those factors multiply together.

Does anything I generate leave my computer?

No. The interface looks like a website but the address is localhost:7860, which literally means this computer here. You can switch your Wi-Fi off and it keeps working. The only time it needs the internet is the first time you select a model, when it downloads that model from the internet, and during the install itself.

Do I need to know how to code to install this?

No, and that is the point of the install prompt on this page. Claude Code reads your hardware, installs the prerequisites, creates the Python environment, works out which build of PyTorch matches your graphics card, installs the dependencies, writes you a launcher and then fixes the errors from the first run. You approve each step. You never type a command yourself, though every command is shown to you.

Why did my first video come out looking terrible?

Because the first model you should try is the smallest one in the family. Wan 2.1 T2V 1.3B gets faces wrong, hands are a nightmare and motion is mushy. That is expected and it is not your installation. The point of the first generation is to prove the whole chain works. Quality comes from moving up to a 5B or 14B model, writing a proper six-part prompt, and only then raising resolution and step count.

Want this built for you?

We write these memos because we build this stuff every day. If you want it working in your business instead of sitting on your reading list, that is literally our job.