Pixel & Oak/Free Tools/Local AI VRAM Calculator

Runs Locally · bandwidth arithmetic

What your card can actually hold and read

Decoding one token means reading every weight, so the speed limit is just bandwidth ÷ bytes read per token. Pick a machine and a build below. Every figure here is a manufacturer specification or a file size read off Hugging Face — these are ceilings, not benchmarks.

Overrides the build above.

Where the conversation runs out

Memory needed climbs as the chat grows — weights are fixed, the cache is not. Where the climbing line meets your machine's flat limit is the whole answer.

Every machine, same build

Sorted by reading speed. Your selection is highlighted.

What this does not claim

Sources