Chapter 2 · free · 20 min
Plan it - hardware, memory, power, and what one small box can really do
The good news first: you do not need a big graphics card or a server rack. Everything in this course runs on one quiet mini PC. This chapter helps you choose it - or check whether something you already own is enough.
The one number that matters: memory
A language model has to fit into memory together with its "context" (the conversation it is looking at). Speed comes later; if it does not fit, it does not run well.
Rough sizes of open models at the common 4-bit quality level ("Q4"):
- a 3-4B model: about 2.5 GB - fast, simple answers, fine for commands and short questions
- a 7-8B model: about 5 GB - good everyday assistant
- a 14B model: about 9 GB - noticeably better at reasoning and writing; our choice
- plus 1-3 GB for a long context (the conversation), and a few GB for everything else on the server
So: 32 GB of RAM is the minimum we recommend, 64 GB is comfortable. RAM is cheap compared to a graphics card.
Integrated graphics are enough (if they can use the memory)
Modern mini PCs with a recent AMD Ryzen chip have integrated graphics that share the main memory. In the BIOS you can give the graphics part a large share (for example 16 GB of 64 GB), and a model runs on it with the free Vulkan backend of llama.cpp - no special drivers.
Real numbers from our own build (Ryzen 7 8845HS mini PC, 64 GB RAM, 16 GB assigned to the integrated graphics, a 14B model at Q4 with a 32K context):
- about 8 tokens (roughly 6 words) per second while writing an answer - about reading speed
- when answers are streamed, the first words appear after 1-2.5 seconds, a short spoken answer is complete after 6-8 seconds
- one job at a time: a second question waits until the first is done. Chapter 4 shows how a small queue and streaming make this feel fast anyway.
A dedicated graphics card is faster, but uses much more power and money. Start without one; you can always add a bigger machine later without changing the design.
The shopping list
The server (always on)
- A mini PC with a recent Ryzen 7 (or similar), 32-64 GB RAM, a 1 TB NVMe SSD.
- Check before buying: the BIOS lets you set the graphics memory (often called "UMA frame buffer size"), two memory slots (so you can go to 64 GB), and a wired network port.
- Or reuse: an old PC or laptop works for learning; plan the upgrade when you know what you use it for.
Voice at home
- Any USB microphone works for a start (even a webcam microphone). A USB speakerphone with echo cancellation is better if the assistant should hear you while music plays.
- Any speaker (USB or the audio jack).
Phone control
- Your phone and a chat app with a bot feature (chapter 8 uses a chat bot; nothing has to be installed on the phone).
Safety and backups
- A second storage place for backups: a USB disk, a NAS, or another computer. One copy on the server itself is not a backup (chapter 10).
- Optional: a small UPS so a power cut does not interrupt writes.
Power and noise
A mini PC like this is quiet and draws little power at idle; it uses more while answering. Measure yours with a cheap plug-in power meter for a day - that tells you the real yearly cost at your electricity price better than any spec sheet.
Where it lives
- Wired network (Ethernet), not Wi-Fi, for the server.
- Somewhere with air flow, not in a closed cupboard.
- The microphone and speaker where you talk - they can be in another room, connected over the network later.
Your checklist before chapter 3
- [ ] Mini PC (or a reused computer) with at least 32 GB RAM and an SSD
- [ ] USB microphone and a speaker
- [ ] A USB stick (8 GB+) for installing Linux
- [ ] A second place for backups
- [ ] An evening for chapter 3: installing Linux and Docker the safe way
Next: Chapter 3 - Linux + Docker, done safely (members).