FERVYNDownload
← Back

Writing quality by GPU tier

Fervyn runs entirely on your computer, so the writing — how well the companion tracks the scene, stays in character, and follows what you ask — scales with your GPU, just like the pictures do. This page is the honest picture of what changes as you go up, and which card is the sweet spot. It's backed by our own blind side-by-side testing, not marketing.

The short version

What runs on your card

Your GPUCompanion writerWhat you get versus the tier below
8 GB (RTX 3060 8GB / 4060 / 5060)12B modelThe full picture-book experience with a capable 12B writer. Chat and pictures take turns on the card, so a picture adds a few seconds.
12 GB (RTX 3060 12GB / 4070 / 5070)same 12B, higher precisionFull-fidelity image engine. The biggest value step — much better pictures, same strong writing.
16 GB (RTX 4080 / 5080 / 5060 Ti 16GB)same 12BSpeed and memory. On the Speed preset, chat and pictures stay loaded together (no waiting); Quality swaps between them so it can run the full-fidelity image engine. Long-term memory is on by default.
24 GB (RTX 3090 / 4090)24B modelRicher writing. A larger model — noticeably better prose and scene-handling. Chat and full-quality pictures take turns, so a picture adds a few seconds.
32 GB (RTX 5090)24B model (optional 31B)The no-compromise tier. The same rich 24B writing as 24 GB, but pictures generate alongside the chat with no waiting. Prefer the richest possible prose? Switch on the larger 31B model in Advanced settings — it trades the instant pictures for a short pause each time.

Fervyn auto-detects your VRAM and picks the right writer for your card — you don't configure anything. A tinkerer can swap in a different model from the Advanced settings.

Which card should I get?

Bottom line: 8 GB is a real, recommended experience. 12 GB is the value pick. 16 GB is the sweet spot. 24 GB+ is for people who want the richest writing. See the full system requirements for the details on each tier.

How we tested this

We don't want to hand-wave "bigger is better," so we measured it. Every candidate model was driven through the same fixed set of role-play conversations on the real app — multi-character scenes that stress name-tracking, long conversations that stress memory, and format requests that stress instruction-following. The transcripts were then scored blind (the judges couldn't see which model produced which transcript) by multiple independent reviewers, plus direct head-to-head comparisons.

A few honest notes about the scores, so you can read them fairly:

This reflects the models shipping in the current release; we re-test and update as better local models come out.

Download the Lite installer