FERVYNDownload
← Back

Installation guide: what to expect, band by band

Fervyn auto-detects your GPU and configures itself — but the install is honest work: a small installer and a one-time download sized to your card. There is nothing else to install: the app provisions its own runtime stack (chat engine, image engine, and the runtime they run on) on first launch. This page walks the whole flow so nothing surprises you. It was written by actually running the installer, start to finish, and timing it.

Before you start: what you need

The installer is one file (~165 MB). Everything else — chat engine, image engine, models — is fetched by the app itself on first run. You need exactly two things:

No Node.js, no Ollama, no ComfyUI, no Python, no terminal. The app ships its own private copies and runs them out of its own folder — they won't touch or conflict with versions you may already have installed. (Already run your own Ollama or ComfyUI? A settings toggle points Fervyn at yours instead.)

What your card gets — and what it downloads

One installer for everyone; your VRAM picks the configuration. The one-time download happens inside the app on first run, shows its progress, and resumes where it left off if your connection drops or you close the app. The figures below include everything: the runtime stack (~4.3 GB, same for every band) plus the models for your band.

Your GPUChat modelImage engineOne-time downloadNotes
8 GB (RTX 3060 8GB / 4060 / 5060, most gaming laptops) 12B model, compressed to fit Compressed (Q8) image engine — the full picture-book experience, tuned to fit ~28 GB Budget config: 768px renders, tighter batching. Everything works; it's the efficiency build.
12 GB (RTX 3060 12GB / 4070 / 5070) 12B model — noticeably richer roleplay Full-quality (fp16) image engine — same engine the top bands use ~32 GB The biggest single step up in the ladder: full-quality images plus a much larger chat model.
16 GB (RTX 4080 / 5080 / 5060 Ti 16GB) 12B model Full-quality (fp16) ~32 GB Long-term memory is on by default at 16 GB and up.
24 GB (RTX 3090 / 4090) 24B model Full-quality (fp16) ~38 GB More headroom: smoother swaps between chat and image duty. The download step up from 16 GB is the larger 24B chat model.
32 GB (RTX 5090) 24B model (optional 31B) Full-quality (fp16) ~38 GB Chat and image models stay loaded together — no swap pause between a reply and its picture. Advanced settings can switch on a larger 31B model for the richest prose (which then takes turns with the image model).

Disk to budget: ~34 GB free for the 8 GB config, ~38 GB free at 12–16 GB, ~45 GB free at 24 GB and up (app + runtime stack + models + support packs, plus room to unpack them — several arrive as archives that are expanded and then deleted). The installer itself checks for considerably less than that at install time — trust this table, not the installer's minimum, or you can pass the install and stall mid-download.

The install, step by step

  1. Download the installer (~165 MB) and verify the SHA-256.
  2. Expect Windows SmartScreen. The installer isn't code-signed (we say so on the front page); Windows shows "unknown publisher". Verify the hash, then More info → Run anyway.
  3. Run the installer — under a minute. It lays down the app (its own runtime included) and creates the engine + model folders next to it. No admin prompt.
  4. The download screen fetches the runtime stack (chat, image, and voice engines, ~4.3 GB) and then the models for your band (table above), in order, with progress. Interrupted? Relaunch — it picks up where it stopped. On a 100 Mbit connection the downloading takes roughly 25–55 minutes depending on band. Unpacking is separate, and on a slow machine it takes longer than the download — the image engine alone expands to about 37,000 files. On a low-end 4-core box we measured the whole first run at around five hours end to end; on a modern desktop it's a good deal quicker. Either way it happens once.
  5. First session: pick a scenario from the cover grid (or import your TavernCard / Chub roster). The first reply streams in seconds; the first image takes a little longer the very first time while the engine warms — after that, panels land while you read. A one-time context optimization also runs in the background so your first real chat turn is fast, not two minutes.

If something looks stuck

Everything above runs on your machine. After the one-time downloads, the companion chat path makes zero outbound network calls — the same guarantee the privacy receipt documents.

Download Fervyn — free