Make AI videos on your own computer
The short version: the same kind of model behind the AI video you are paying a monthly fee for can be downloaded and run on your own machine, for nothing. The software is WanGP, it is free and open source, and once it is installed it works with the Wi-Fi switched off. You need either a Windows PC with an NVIDIA card of 8 GB VRAM or more (12 GB is the sensible entry point) or an Apple Silicon Mac with 16 GB or more of unified memory. Installing it is genuinely horrible, so you do not: you paste the prompt on this page into Claude Code and let it do the install. That subscription is the only thing here that costs money, and only until the install finishes.
What you actually do, once it is working
Worth getting the shape of this clear before any of the hardware talk, because the day-to-day is much smaller than the setup suggests.
You double-click a launcher. A black window opens and stays open, because it is the engine. After a moment a page opens in your browser at localhost:7860. On that page there is a model dropdown, a big empty text box, a couple of number fields and a Generate button.
You type one sentence into the box. Something like two cats playing basketball on a city court, sunny, shallow depth of field. You set the frame count to 49, which is about two seconds, and the step count to 20. You click Generate.
Some minutes later a video file appears in the outputs folder inside the Wan2GP directory, and in the “Generated videos” panel on the page. That is the whole loop. No credits, no queue, no monthly cap, and nothing uploaded anywhere.
localhost literally means this computer here. You can turn your Wi-Fi off and keep generating. The only moments that genuinely need the internet are the install, and the first time you pick a model, when it downloads that model to your disk.Can your machine do it?
On both platforms the answer comes down to almost exactly one number: how much memory your graphics processor can reach. On a PC that is VRAM, soldered onto the graphics card, and it is small and expensive. On an Apple Silicon Mac there is one shared pool, which is what “unified memory” means, so a 48 GB MacBook has more usable graphics memory than any consumer NVIDIA card. The problem on the Mac is not size, it is speed.
Windows PC with an NVIDIA card
Press Ctrl + Shift + Esc, go to Performance, click GPU, and read “Dedicated GPU memory”. That number is your VRAM.
| Your VRAM | Verdict | What you can actually do |
|---|---|---|
| Under 6 GB | No | Not worth the pain. Even the small models spill out of the card and you will spend your life reading out-of-memory errors. |
| 8 GB (RTX 3060 Ti, 4060, 2070) | Just about | Wan 2.1 T2V 1.3B at 480p, plus quantised medium models. Use --profile 5 and keep clips short. |
| 12 GB (RTX 3060 12 GB, 4070, 5070) | Fine | Medium 5B-8B models at 480p comfortably, 720p on short clips. The sensible entry point. |
| 16 GB (RTX 4080, 5080, 4060 Ti 16 GB) | Good | The 14B models in quantised form, 720p, clips of several seconds. Enough for real work. |
| 24 GB and up (3090, 4090, 5090) | Comfortable | Nearly everything, including the big 22B models at full precision. Memory stops being the limit. |
AMD and Intel cards are not realistically supported. And system RAM matters more than people say: WanGP’s whole trick is keeping part of the model in system RAM and streaming it into the card as needed, which is how a 16 GB card runs a model that does not fit in 16 GB. That makes 16 GB of system RAM a real bottleneck and 32 GB a genuine upgrade, separately from your graphics card.
Mac with Apple Silicon
Apple menu, About This Mac. Note the chip, the memory and the macOS version. You need M1 or newer and macOS 13 or later; Intel Macs are out.
| Unified memory | Verdict | What you can actually do |
|---|---|---|
| 8 GB | No | It will not fit. The smallest useful video model already takes nearly all of it, and macOS has to live somewhere. |
| 16 GB | Just about | Wan 2.1 T2V 1.3B, 480p, 2 to 3 seconds. Close every other app. It works, but it is handicraft. |
| 24 GB | Fine | The 1.3B comfortably, plus Wan 2.2 TI2V 5B or HunyuanVideo 1.5 at 480p. Clips of 3 to 5 seconds. |
| 36 to 48 GB | Good | 8B to 14B models and 720p. The first tier where results start looking like what you see online. |
| 64 GB and up | Comfortable | Almost everything, including the big 22B models. Memory is no longer the limit. Patience is. |
The chip affects speed, not what is possible. An M4 Max is roughly four to six times faster than a base M1 for this work. One technical wrinkle: M1 and M2 do not handle a number format called bfloat16, so the software falls back to another one. On M3 and M4 that does not arise.
Budget disk space too. A single video model weighs between 2 GB and 25 GB, and WanGP downloads the text encoder and video decoder alongside it. 50 GB free is the minimum for trying one model, 100 to 150 GB for two or three, 250 GB and up to explore properly. Keep 25 GB of headroom on top. On Windows, keep the models off any path with spaces or accents in it and out of OneDrive-synced folders, because both cause strange failures later.
PC versus Mac: the honest gap
Plenty of tutorials say “it runs on a Mac” without mentioning the price of that. The thing at the centre of it is called CUDA, an NVIDIA technology that all modern AI was built on, and which researchers write their most aggressive optimisations in. Apple’s equivalent is MPS. It works, it is decent, and it does not have ten years of optimisation from the entire planet behind it.
| Acceleration | NVIDIA PC | Apple Silicon |
|---|---|---|
| SageAttention / Flash Attention (the biggest win, 2-3x) | Yes | Does not exist |
| Triton GPU kernel compiler | Yes | Does not exist |
| torch.compile | Yes | Force-disabled |
| NVFP4 / INT8 CUDA quantisation | Yes | Incompatible |
| Pinned memory and async transfers | Yes | Disabled |
| SDPA (PyTorch standard attention) | Yes | Yes, and this is the one you use |
| Lots of available memory | 8-16 GB, 24 GB at the top | Up to 128 GB |
The Mac support is not a hack, though, which is worth saying. It is built into WanGP: a dedicated Apple Silicon module detects your chip, presents itself as an NVIDIA card to the rest of the program, disables everything that does not exist on a Mac, and routes the maths to Apple’s GPU. The dependency list contains lines written specifically for macOS.
So: use the PC if you have one with a decent NVIDIA card, because it is the right tool by a wide margin. Use the Mac if that is what you own and you want to learn and make short clips without paying anyone or uploading anything. Rent a cloud GPU by the hour if you need ten videos a day and do not own the hardware.
How a computer invents a video
You do not need this to install anything, but every setting further down stops being a mystery button once you have it.
You know the grey fizzing snow of an old untuned television? In computing that is called noise: pure randomness, with no picture in it at all. The AI starts exactly there. It fills the screen with random snow, then asks itself a strange question: if this image were really a cat running on a beach hidden under noise, what would it look like with a tiny bit of that noise removed?
It removes a little. It asks again. It removes a little more. Twenty or thirty times. Each pass the picture gets less blurry and more “cat on a beach”, and at the end there is no noise left, only the image. A sculptor looks at a block of marble and says the statue is already in there, and they are removing what is not part of it. Same thing, with noise instead of marble. Your sentence is what tells it which statue to look for.
Video makes that much more expensive. A video is not one image, it is 49 images in a row, and if the AI made them one at a time the cat would change colour every frame. So it denoises all 49 frames at once, constantly looking from one frame to the next to stay consistent. That is why video is not “49 times slower” than an image but often considerably worse, and it is why video eats memory the way it does: every frame has to be held at the same time.
Which teaches you four things you will use in a minute:
- More denoising steps means cleaner output and proportionally longer waits.
- More frames means longer waits and more memory, at the same time.
- Higher resolution makes every single step cost more.
- Your prompt is the only compass the model has for the whole process.
The install: one prompt, both machines
Here is the part everyone gets stuck on. The software is free and the models are free, but installing them means stacking a lot of tools that all have to agree with each other, and the internet is full of contradictory advice about which versions those should be.
So do not do it yourself. Claude Code is an AI that lives in your terminal: you tell it in plain English what you want, it reads your files, writes the commands, runs them with your permission, and fixes its own mistakes. If you have not met it before, the ECC config memo is a gentler first introduction to what it is and how it behaves.
It needs a Claude Pro, Max, Team or Enterprise account. The free plan does not include Claude Code. That subscription is the only money in this entire page, and you only need it for the install.
Open a terminal
PC: press the Windows key, type “PowerShell”, press Enter. The normal one, not the administrator one.
Mac: press Cmd + Space, type “Terminal”, press Enter.
A terminal is just a window where you type orders instead of clicking. You will use it about four times in this whole process. The universal stop button is Ctrl + C on both platforms, and if the cursor blinks and nothing appears to happen, it is usually working silently. Some steps here take twenty minutes.
Install Claude Code
Windows, into PowerShell:
irm https://claude.ai/install.ps1 | iex
If Windows blocks the script with an execution policy error, run this first. It affects only your own account, only this session:
Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
Mac, into Terminal:
curl -fsSL https://claude.ai/install.sh | bash
Then check it landed with claude --version. If you get “command not found” or “not recognized as a cmdlet”, close the terminal window completely and open a fresh one. That fixes it about 90% of the time.
Make a working folder and start it
Windows:
mkdir $HOME\AI; cd $HOME\AI; claude
Mac:
mkdir -p ~/AI && cd ~/AI && claude
On the first start it opens your browser to sign in to your Claude account. Approve, come back to the terminal, done for good.
Paste the prompt for your machine
Set aside one to two hours, most of it downloading. On a Mac, plug into the mains first. The two prompts are not interchangeable — use the one for your platform.
Windows + NVIDIA
I am on Windows with an NVIDIA graphics card. I want to install WanGP
(the Wan2GP project) to generate AI videos locally on my GPU via CUDA.
I am a beginner: explain each step in simple English before you do it.
Do the following steps in order:
1. DIAGNOSIS
Check my exact GPU model, my VRAM, my driver version, my CUDA version,
my system RAM and my free disk space. Run nvidia-smi and show me the
output. Tell me honestly whether my machine can run this software and
which models are realistic for me. If I have less than 6 GB of VRAM or
less than 50 GB of free disk, stop and warn me.
2. PREREQUISITES
Install only what is missing:
- Python 3.10 or 3.11 (NOT 3.12 or newer - some dependencies break)
- git
- ffmpeg
Add them to PATH. Do not touch any Python that another program depends on.
3. DOWNLOAD THE SOFTWARE
Clone https://github.com/deepbeepmeep/Wan2GP into a folder with no
spaces and no accents in its path, and NOT inside OneDrive.
4. PYTHON ENVIRONMENT
Create a Python 3.10/3.11 virtual environment called .venv inside the
Wan2GP folder and activate it. Everything below installs in there,
never globally.
5. PYTORCH - THIS IS THE STEP THAT GOES WRONG
Look up my GPU architecture first. If I have an RTX 50-series card
(Blackwell, sm_120), I need a PyTorch build for CUDA 12.8 or newer -
older builds fail with "no kernel image is available for execution on
the device". For older cards, a CUDA 12.4+ build is fine.
Install the matching torch, torchvision and torchaudio from the correct
PyTorch index URL. Then verify torch.cuda.is_available() returns True
AND that a small tensor operation actually runs on the GPU. Show me
both results. If either fails, stop and explain why.
6. DEPENDENCIES
Install requirements.txt. If a package fails to build:
- do not abandon the whole install;
- tell me which one, what it is for, and whether WanGP works without it;
- continue with the rest.
7. SPEED EXTRAS (optional, do this only after everything already works)
Try to install SageAttention and Triton for Windows - they give the
biggest single speed-up (roughly 2x). They need matching CUDA and
compiler versions and they often fail to build. If they do not install
cleanly in a reasonable time, skip them and tell me: WanGP runs fine
without them using --attention sdpa.
8. LAUNCH SCRIPT
Create start-wangp.bat next to the Wan2GP folder that:
- changes into the Wan2GP folder
- activates .venv
- runs: python wgp.py --attention sdpa --profile 4 --open-browser
If SageAttention installed successfully in step 7, use
--attention sage2 instead of sdpa and tell me you did.
Choose the profile based on MY VRAM: profile 5 for 8 GB, profile 4 for
12-16 GB, profile 3 for 24 GB or more. Explain your choice.
9. FIRST RUN
Run the script with me and fix errors one at a time until the web
interface opens on http://localhost:7860 . Show me the error messages
in plain text and explain what you are fixing.
10. SUMMARY
When everything works, write me a HOW-TO-LAUNCH.md file, max 10 lines,
explaining how to start the software later and which settings suit MY
VRAM specifically.
General rules:
- Explain each command in one sentence before running it.
- If you have to choose between "fast" and "it works", choose "it works".
- Do not modify anything outside my AI working folder.
- If you are stuck more than three attempts on the same error, stop and
explain the problem in simple English instead of trying again.Mac with Apple Silicon
I am on a Mac with an Apple Silicon chip. I want to install WanGP (the Wan2GP project) to generate AI videos locally, on the Mac GPU via MPS. I am a beginner: explain each step in simple English before you do it. Do the following steps in order: 1. DIAGNOSIS Check my chip, my amount of unified memory, my macOS version and my free disk space. Tell me honestly whether my machine can run this software and which models are realistic for me. If I have less than 16 GB of memory or less than 50 GB of free disk, stop and warn me. 2. PREREQUISITES Install only what is missing: - the Xcode Command Line Tools - Homebrew - git, ffmpeg - Python 3.11 (via Homebrew) Do not touch the Python that ships with macOS. 3. DOWNLOAD THE SOFTWARE Clone https://github.com/deepbeepmeep/Wan2GP into ~/AI/Wan2GP 4. PYTHON ENVIRONMENT Create a Python 3.11 virtual environment called .venv inside ~/AI/Wan2GP and activate it. Everything below installs in there, never globally. 5. MAC BUILD OF PYTORCH Install PyTorch for macOS: pip install torch torchvision torchaudio Use NO CUDA index. Then verify torch.backends.mps.is_available() returns True and show me the result. If it is False, stop and explain why. 6. DEPENDENCIES Install requirements.txt. WARNING: this file is written for Windows/NVIDIA first. Its first line is a CUDA index, and it contains packages that do not exist on Mac. Most lines already carry platform_system == "Darwin" markers, but if a package still refuses to install: - do not abandon the whole install; - tell me which one, what it is for, and offer to skip it; - continue with the rest. NEVER install: sageattention, flash-attn, triton, xformers, onnxruntime-gpu, decord, or any package whose name contains cuda, cu12, cu13 or nvidia. They do not exist on Mac and will fail. 7. LAUNCH SCRIPT Create ~/AI/Wan2GP/start-mac.sh that: - changes into the script's folder - activates .venv - runs: python wgp.py --attention sdpa --profile 4 --open-browser Make it executable with chmod +x. Do NOT use --compile, and do NOT use --attention sage/sage2/flash: they do not exist on Mac and will crash. 8. FIRST RUN Run the script with me and fix errors one at a time until the web interface opens on http://localhost:7860 . Show me the error messages in plain text and explain what you are fixing. 9. SUMMARY When everything works, write me a ~/AI/HOW-TO-LAUNCH.md file, max 10 lines, explaining how to start the software later and which settings suit MY amount of memory specifically. General rules: - Explain each command in one sentence before running it. - If you have to choose between "fast" and "it works", choose "it works". - Do not modify anything outside ~/AI. - If you are stuck more than three attempts on the same error, stop and explain the problem in simple English instead of trying again.
Roughly what happens, and how long each phase takes:
| Phase | Typical duration | What you see |
|---|---|---|
| Diagnosis | 1 minute | What you have, and whether it is viable |
| Prerequisites | 5-25 minutes | Lots of scrolling text. On Mac the Xcode Command Line Tools may open a system window: accept it. On PC, Python or git installers may raise a UAC prompt: accept it. |
| Downloading the code | 1-3 minutes | WanGP lands in your AI folder |
| PyTorch and dependencies | 10-40 minutes | The longest and noisiest phase. Yellow warnings are normal. |
| Speed extras (PC only) | 0-30 minutes | SageAttention and Triton either build or they do not. Either outcome is fine. |
| First run | 2-10 minutes | Possibly a few rounds of fixing. That is expected. |
Three rescue prompts
These cover the large majority of ways it goes wrong.
The install is stuck on this package. Explain in simple English what it is for, whether WanGP can work without it on my platform, and if so, skip it and carry on with the installation.
I want to start over cleanly. Delete only the Wan2GP folder and its virtual environment, without touching Python, git, my GPU driver, or anything else on this computer. Then resume the installation from step 3.
Stop. Explain to me in very simple English, as if to someone who knows nothing about this: where we are, what just happened, whether there is a problem, and what the next thing I need to do is.
Launching it, every time after that
From here you do not need Claude Code day to day. On Windows, double-click start-wangp.bat. A black window opens — leave it open, it is the engine — and the browser follows on its own. On a Mac:
cd ~/AI/Wan2GP && ./start-mac.sh
What the launcher options mean:
| Option | In plain English |
|---|---|
--attention sdpa | Use the standard maths. On a Mac this is the only one available. On a PC, if SageAttention installed, use --attention sage2 instead: roughly twice as fast, for free. |
--profile 4 | Save memory even if it is slower. A good default on both. |
--open-browser | Open the browser when the engine is ready. |
--preload 8000 (PC) | Keep 8 GB of the model resident in VRAM instead of streaming it every step. A big speed win if the card has room; lower it on out-of-memory errors. |
--perc-reserved-mem-max 0.3 (Mac) | Add this if your Mac stutters or freezes. |
--advanced | Show every advanced setting in the interface. Add later. |
The profiles are strategies for where the model lives while it works, and lower means more of it in fast memory: --profile 3 for 24 GB VRAM or a 64 GB+ Mac, --profile 4 for 12-16 GB VRAM and most Macs, --profile 5 for 8 GB or anything that keeps crashing.
One warning about the first start: the first time you select a model, WanGP downloads it, and depending on the model that is 2 to 25 GB. The interface can look frozen the entire time. Watch the terminal window instead, where the download progress actually is.
The settings for your memory
Four dials decide 90% of the result and 100% of the waiting: the model, the resolution, the frame count and the step count. Landmarks worth memorising: 25 frames is about 1 second, 49 is 2 seconds, 73 is 3 seconds. For steps, 15 is fast, 20 is balanced, 30 is careful, and past 30 the gain is nearly invisible. Resolution is the expensive one, because doubling the width and height multiplies the work by four, not two.
There is a fifth, less discussed: Guidance, which decides how literally the model obeys your prompt. 3 to 5 interprets freely, 7 to 10 is the default and the right answer in almost every case, and 12 and up gives literal obedience that can look oversaturated and stiff.
Starting settings, Windows PC
| VRAM | Model | Resolution | Frames | Steps | Profile |
|---|---|---|---|---|---|
| 8 GB | Wan 2.1 T2V 1.3B | 480p | 49 | 20 | 5 |
| 12 GB | Wan 2.2 TI2V 5B | 480p | 49-73 | 20 | 4 |
| 16 GB | Wan 2.2 14B (quantised), HunyuanVideo 1.5 | 480p to 720p | 73-97 | 20-25 | 4 |
| 24 GB+ | Almost anything, including LTX-2.3 | 720p | 97+ | 25-30 | 3 |
Starting settings, Apple Silicon Mac
| Memory | Model | Resolution | Frames | Steps | Clip length |
|---|---|---|---|---|---|
| 16 GB | Wan 2.1 T2V 1.3B | 480p | 25-49 | 15-20 | 1-2 s |
| 24 GB | Wan 2.1 1.3B, then Wan 2.2 TI2V 5B | 480p | 49-73 | 20 | 2-3 s |
| 36-48 GB | Wan 2.2 T2V/I2V, HunyuanVideo 1.5 | 480p to 720p | 73-97 | 20-25 | 3-4 s |
| 64 GB+ | Almost anything, including LTX-2.3 distilled | 720p | 97+ | 25-30 | 4 s and up |
Two more habits that pay for themselves immediately. Note your seeds, because the seed is the number the randomness starts from and the same prompt plus the same seed gives the same video, which is how you reproduce a good result or change exactly one thing. And change one setting at a time, or you will never know what improved it.
Which model to pick
WanGP gives you access to more than two hundred models. About a dozen matter. The three families: Wan 2.1 / 2.2 (1.3B to 14B, the generalist family and the most mature, where you start), HunyuanVideo 1.5 (8.3B, a notch better image quality, with a lighter 480p version), and LTX-2 / LTX-2.3 (22B, generates picture and sound together and handles long clips, but 22 billion parameters is not for modest hardware).
You will also meet words like Lightning, FusioniX, FastWan, Distilled, Self-Forcing, GGUF, NVFP4 and Nunchaku. These are not new families, they are accelerated or compressed versions of an existing model. On a PC they are your friends: a quantised 14B is how a 16 GB card runs a model that should not fit. On a Mac, anything called NVFP4 or Nunchaku will not work, because those formats need NVIDIA cards. When in doubt take the plain or bf16 version.
Quantisation, by the way, is fitting a model into less memory by writing its numbers with less precision. It is going from a RAW photo to a JPEG: much lighter, very slightly less beautiful, and in 95% of cases you cannot see the difference.
Or just ask, from inside the Wan2GP folder:
Look at my available video memory and at the list of models in WanGP's defaults/ folder. Recommend the three video models best suited to MY machine, from lightest to heaviest, explaining for each: its approximate size, what it can do, and the realistic generation time I should expect. Rule out anything my hardware cannot run.
Writing a prompt that works
An average model with an excellent prompt beats an excellent model with a sloppy one, and this is the only part of the whole business where you improve enormously without buying hardware. Build it like a brief for a camera operator, in this order: the subject (an old fisherman, not “a person”), the action (mending a net, slowly — with no verb you get an animated photo), the setting (on a wooden pier at dawn), the light (soft golden light, light fog — the most powerful lever you have), the camera (slow dolly-in, shallow depth of field), and the style (35mm film, cinematic, muted colors).
An old fisherman mending a net, slowly, on a wooden pier at dawn, soft golden light, light fog over the water, slow dolly-in, shallow depth of field, 35 mm film, cinematic, muted colors.
And a negative prompt that works in almost every case:
blurry, low quality, distorted face, deformed hands, extra fingers, watermark, text, jittery motion, flickering
Three mistakes to skip past. The one-word prompt (“a dragon”) makes the model invent everything, and what it invents is the average of everything it has seen, which is bland. The shopping list of twenty adjectives (“beautiful, amazing, stunning, 8k, ultra realistic, masterpiece”) dilutes the words that do matter. And asking for five actions in a two-second clip (“he opens the door, walks in, sits down, turns on the lamp and smiles”) is impossible in 49 frames, so the model blends it into mush. One shot, one action. To tell a story, chain clips.
Write prompts in English, because these models were trained almost entirely on English descriptions and other languages give noticeably less faithful results. If your English is shaky, ask Claude to translate and improve the description first.
How long it actually takes
Tutorial videos show generations finishing in 45 seconds. Here is what to expect, as orders of magnitude rather than promises, for a 2-second clip at 480p and 20 steps:
| Machine | Small model (1.3B) | Medium model (5B-8B) |
|---|---|---|
| PC · RTX 3060 12 GB | 60-120 s | 4-8 min |
| PC · RTX 4070 / 5070 | 40-70 s | 2-5 min |
| PC · RTX 4080 / 5080 | ~30 s | 1-3 min |
| PC · RTX 4090 / 5090 | 15-25 s | 45 s - 2 min |
| Mac · Air M1/M2, 16 GB | 20-40 min | Out of reach |
| Mac · M3 Pro, 18-24 GB | 10-20 min | 45-90 min |
| Mac · M4 Pro, 24-48 GB | 5-12 min | 25-50 min |
| Mac · M4 Max, 48-128 GB | 3-8 min | 12-30 min |
Those PC numbers assume plain --attention sdpa. If SageAttention installed and you run --attention sage2, roughly halve them.
And what multiplies the wait, which is the part people miss:
| If you change | Time is multiplied by |
|---|---|
| 480p to 720p | about 2.5 |
| 2 s to 4 s of video | about 2 (and memory explodes) |
| 15 to 30 steps | exactly 2 |
| 1.3B model to 14B model | 5 to 10 |
These multiply together. Going from “480p, 2 s, 15 steps, 1.3B” to “720p, 4 s, 30 steps, 14B” is not a bit slower, it is on the order of fifty to a hundred times slower.
When it breaks
It will break. That is not bad luck, it is the norm in this field. The universal reflex: select the error message in the terminal, copy it, start Claude Code and paste it with “Here is the error I get when running WanGP. Explain in simple English what is wrong and fix it.” That resolves the large majority of cases without you needing to understand anything.
| Symptom | Likely cause | What you do |
|---|---|---|
| “out of memory” / “CUDA out of memory” | You asked for more than your memory holds | Reduce in this order: frame count, then resolution, then a smaller model. Close other apps. On PC try --profile 5 or lower --preload. |
| “no kernel image is available for execution on the device” (PC) | Your PyTorch build does not know your card’s architecture. Classic on RTX 50-series. | Reinstall PyTorch from a CUDA 12.8+ index. Paste the error to Claude Code and say which card you have. |
| torch.cuda.is_available() is False (PC) | CPU-only PyTorch build, or the driver is too old | Ask Claude Code to check the driver with nvidia-smi and reinstall the correct CUDA build. |
| “CUDA not available” (Mac) | Some code thinks it is talking to an NVIDIA card | Give the error to Claude Code. It is a known case with a known Mac-side fix. |
| “sageattention not found” / flash / triton | You launched with an attention mode you do not have | Change the launcher to --attention sdpa. On a Mac those three never exist. |
| localhost:7860 will not open | Still starting, or it crashed on launch | Read the terminal window, not the browser. The real message is there. |
One safety rule that outlives this guide: never paste a command into a terminal when you do not know where it came from. A command can delete files. Everything here comes from official documentation and is explained, and from the install prompt onward it is Claude Code writing the commands and asking your permission before each one. If you are working this way often, the habits that keep a Claude Code session cheap are worth twenty minutes.
What it actually costs
WanGP is free. The models are free. Electricity and disk space are real but small. The one line item is a Claude subscription for the install assistant, and only for as long as the install takes. After that you generate forever, paying nothing, with no internet connection at all.
The honest counterweight, so nobody arrives disappointed: a 16 GB card runs out of memory constantly, Windows has its own creative ways of breaking, and a MacBook is 5 to 15 times slower than a gaming PC and no setting changes that. Both platforms genuinely work, for free, and both are more than enough to learn, experiment and make short clips. Neither is a free replacement for a professional cloud pipeline, and anyone telling you otherwise is selling something.
And your first video will probably be disappointing. That is expected. The 1.3B model is the smallest in the family: faces are wrong, hands are nightmarish, motion is mushy. It is not your machine and it is not your installation. You have just proved the whole chain works, which was the only goal. From there you move up.
Questions people actually ask
Can I really generate AI video on my own computer for free?
Yes. WanGP (also written Wan2GP) is free and open source, and the video models it runs are free to download. Once installed it works with no internet connection at all, because the interface runs at localhost:7860 on your own machine. The only money involved is a Claude Pro or Max subscription if you use Claude Code to do the installation for you, and you only need that during the install itself.
How much VRAM do I need for AI video generation?
On a Windows PC with an NVIDIA card: under 6 GB is not worth attempting, 8 GB runs the small Wan 2.1 1.3B model at 480p, 12 GB is the sensible entry point for medium 5B-8B models, 16 GB runs quantised 14B models at 720p, and 24 GB or more removes memory as a limit. AMD and Intel cards are not realistically supported. System RAM matters separately: 16 GB is a real bottleneck and 32 GB is a genuine upgrade, because WanGP streams part of the model from system RAM into the card.
Does this work on a Mac?
On Apple Silicon, yes, and the support is built into WanGP rather than being a community patch. There is a dedicated Apple Silicon module that detects your chip, disables the NVIDIA-only features and routes the maths to Apple's GPU through MPS. You need an M1 or newer, macOS 13 or later, and at least 16 GB of unified memory. 8 GB will not fit. Intel Macs are out.
How much slower is a Mac than a gaming PC for AI video?
At identical settings, expect 5 to 15 times slower. A generation that takes one minute on a PC with a recent NVIDIA card takes 15 to 30 minutes on a MacBook. This is not a bug and no setting fixes it: the accelerations that give the biggest speed-ups, SageAttention, Triton, torch.compile and NVIDIA-only compressed model formats, do not exist on Apple Silicon. The workaround is to batch overnight rather than iterate live.
How long does one AI video take to generate?
For a 2-second clip at 480p and 20 steps: roughly 60 to 120 seconds on an RTX 3060 12 GB, 40 to 70 seconds on a 4070 or 5070, and 15 to 25 seconds on a 4090 or 5090, using the small 1.3B model. Medium 5B-8B models take 45 seconds to 8 minutes on the same cards. On a 16 GB M1 or M2 Air the small model takes 20 to 40 minutes. Doubling to 720p multiplies by about 2.5, doubling the length by about 2, and going from 1.3B to 14B by 5 to 10, and those factors multiply together.
Does anything I generate leave my computer?
No. The interface looks like a website but the address is localhost:7860, which literally means this computer here. You can switch your Wi-Fi off and it keeps working. The only time it needs the internet is the first time you select a model, when it downloads that model from the internet, and during the install itself.
Do I need to know how to code to install this?
No, and that is the point of the install prompt on this page. Claude Code reads your hardware, installs the prerequisites, creates the Python environment, works out which build of PyTorch matches your graphics card, installs the dependencies, writes you a launcher and then fixes the errors from the first run. You approve each step. You never type a command yourself, though every command is shown to you.
Why did my first video come out looking terrible?
Because the first model you should try is the smallest one in the family. Wan 2.1 T2V 1.3B gets faces wrong, hands are a nightmare and motion is mushy. That is expected and it is not your installation. The point of the first generation is to prove the whole chain works. Quality comes from moving up to a 5B or 14B model, writing a proper six-part prompt, and only then raising resolution and step count.
Sources
- WanGP / Wan2GP on GitHub — the free, open-source generation software this memo installs
- Anthropic — Install Claude Code (the official installers used above)
- PyTorch — Get Started locally (CUDA index URLs, and the macOS build with no CUDA index)
- NVIDIA — CUDA GPUs and compute capability (why RTX 50-series needs CUDA 12.8+)
- Apple — Metal Performance Shaders, the MPS backend PyTorch uses on Apple Silicon
Want this built for you?
We write these memos because we build this stuff every day. If you want it working in your business instead of sitting on your reading list, that is literally our job.