System requirements
Fervyn runs locally, so the experience scales with your computer — mainly your GPU VRAM. Everything below the Lite line still works; it just renders images faster and at higher quality as you go up. For how the writing scales with your card (and which tier is the sweet spot), see writing quality by GPU tier.
Minimum (Lite tier)
| OS | Windows 10/11 (64-bit) |
|---|---|
| GPU | NVIDIA with 8 GB VRAM (e.g. RTX 3060 8GB / 4060 / 5060; RTX 20-series or newer recommended — most 8 GB gaming laptops qualify) |
| RAM | 16 GB |
| Disk | ~34 GB free for the 8 GB config; ~38 GB free at 12–16 GB; ~45 GB free at 24 GB and up (app + one-time download of the runtime stack and models: ~28 GB at 8 GB, ~32 GB at 12–16 GB, ~38 GB at 24 GB+, plus room to unpack) |
| Image engine | Fast single-pass SDXL-class engine, tuned to run in the 8 GB floor |
The companion chat is light; the GPU is for image generation. On Lite you get the full picture-book experience at the Lite image quality.
Recommended (higher quality)
- 12 GB VRAM (RTX 3060 12GB / 4070 / 5070) → the full-quality (fp16) image engine plus the 12B chat model — the biggest single step up in the ladder for the money.
- 16 GB VRAM (RTX 4080 / 5080 / 5060 Ti 16GB) → the same 12B chat model. On the Speed preset, chat and images run at the same time (no waiting between them); the Quality preset swaps between them so it can run the full-fidelity image engine. Long-term memory is on by default. The sweet spot below enthusiast hardware.
- 24 GB VRAM (RTX 3090 / 4090) → a larger 24B chat model — the first tier where the writing itself gets noticeably richer — plus the top image workflows and heavier multi-character scenes.
- 32 GB VRAM (RTX 5090) → the 24B chat model running alongside the full-quality image engine with no waiting between a reply and its picture — the only tier with no trade-off between full-quality writing and pictures — plus the best long-term memory. Enthusiasts can switch on a larger 31B model in Advanced settings for the richest prose (it then takes a short pause per picture as the models take turns).
The companion writer is chosen automatically for your card. Curious how much the writing actually changes tier to tier? We blind-tested it — see writing quality by GPU tier.
Fervyn auto-detects your VRAM and picks a sensible default — a bigger card unlocks higher tiers automatically. You don't have to configure anything, but you can choose how it uses your card (below).
Speed or quality — your choice
Above the Lite floor, Fervyn lets you pick how it spends your VRAM. You can switch any time in Settings.
- Speed keeps the chat model and the image model loaded at the same time, so replies and pictures come back with no waiting between them — the most fluid back-and-forth.
- Quality uses a larger chat model and a higher-fidelity image engine. On most cards those don't both fit at once, so Fervyn swaps between them — a few extra seconds per picture in exchange for richer writing and renders.
- A 32 GB card (RTX 5090) runs the 24B chat model and the full-quality image engine together — top quality with no waiting. That's the clearest reason to go all the way to 32 GB: it removes the swap pause for full-quality pictures entirely. (On 8–24 GB cards, full-quality pictures take a few seconds each as chat and image take turns; the Speed preset gives instant pictures on any card by using a lighter image engine.)
On the image side, drawing the picture is fast on any capable GPU — the time you feel is the model swap, not the rendering. Speed avoids the swap; Quality accepts it for a better result.
What you do NOT need
- No subscription, no online account, no API keys.
- No ID upload.
- No constant internet connection — after the one-time download (~28 GB for the 8 GB config, ~32 GB at 12–16 GB, ~38 GB at 24 GB+; resumable if interrupted), the companion experience runs offline. Full flow: installation guide.
Not sure if your card qualifies?
If you have an 8 GB+ NVIDIA GPU on Windows, start with the Lite installer. Below the Lite floor, image generation won't run well — the chat will, but the image-first experience is the point.
AMD / Intel GPU and macOS support are not part of the Lite tier today.