Maybe not a mobile device, but running decent local models is definitely within reach for the average consumer today by simply adding a GPU with lots of VRAM to their existing PC.
I mean yes but those gpus are not necessarily within reach of the average consumer anymore. Even if you got one for a grand which is a lot of money you would have cooling and enclosure costs, possibly additional hardware.
24 is ideal, 16 works fine especially with some MoE offloading. If you use llama-swap, a fast PCIe bus and fast storage is ideal. Takes about 20 seconds to swap between two models on my X570 board.
Maybe not a mobile device, but running decent local models is definitely within reach for the average consumer today by simply adding a GPU with lots of VRAM to their existing PC.
I mean yes but those gpus are not necessarily within reach of the average consumer anymore. Even if you got one for a grand which is a lot of money you would have cooling and enclosure costs, possibly additional hardware.
Nah I got an RX 6800 XT for $300 with 16 GB VRAM, perfectly cromulent for a 27B model.
Thats is pretty decent alright, I had heard 24 was what was needed for decent local AI.
24 is ideal, 16 works fine especially with some MoE offloading. If you use llama-swap, a fast PCIe bus and fast storage is ideal. Takes about 20 seconds to swap between two models on my X570 board.