Localhaven

Chapter 2 · free · 20 min

Plan it - hardware, memory, power, and what one small box can really do

The good news first: you do not need a big graphics card or a server rack. Everything in this course runs on one quiet mini PC. This chapter helps you choose it - or check whether something you already own is enough.

The one number that matters: memory

A language model has to fit into memory together with its "context" (the conversation it is looking at). Speed comes later; if it does not fit, it does not run well.

Rough sizes of open models at the common 4-bit quality level ("Q4"):

So: 32 GB of RAM is the minimum we recommend, 64 GB is comfortable. RAM is cheap compared to a graphics card.

Integrated graphics are enough (if they can use the memory)

Modern mini PCs with a recent AMD Ryzen chip have integrated graphics that share the main memory. In the BIOS you can give the graphics part a large share (for example 16 GB of 64 GB), and a model runs on it with the free Vulkan backend of llama.cpp - no special drivers.

Real numbers from our own build (Ryzen 7 8845HS mini PC, 64 GB RAM, 16 GB assigned to the integrated graphics, a 14B model at Q4 with a 32K context):

A dedicated graphics card is faster, but uses much more power and money. Start without one; you can always add a bigger machine later without changing the design.

The shopping list

The server (always on)

Voice at home

Phone control

Safety and backups

Power and noise

A mini PC like this is quiet and draws little power at idle; it uses more while answering. Measure yours with a cheap plug-in power meter for a day - that tells you the real yearly cost at your electricity price better than any spec sheet.

Where it lives

Your checklist before chapter 3

Next: Chapter 3 - Linux + Docker, done safely (members).