Writing quality by GPU tier
Fervyn runs entirely on your computer, so the writing — how well the companion tracks the scene, stays in character, and follows what you ask — scales with your GPU, just like the pictures do. This page is the honest picture of what changes as you go up, and which card is the sweet spot. It's backed by our own blind side-by-side testing, not marketing.
The short version
- 8 GB is the floor, and it now runs a proper 12-billion-parameter writer. This is the single biggest jump we've made — the 8 GB card went from a small 3B model to a real 12B one, roughly 2.5× the writing quality in blind tests. Every 8 GB user gets this.
- 8 → 12 → 16 GB: the writing is about the same (all three run the same 12B writer). What you gain going up is picture quality (full-fidelity image engine at 12 GB+), speed (no waiting between chat and pictures at 16 GB), and long-term memory (16 GB+).
- 24 GB is where the writing itself gets richer — it runs a bigger 24B model, noticeably better prose. 32 GB runs that same 24B writer, but with the pictures generating alongside the chat — no waiting between a reply and its picture, the smoothest experience. Enthusiasts can also switch a 32 GB card to an even larger 31B model for the richest prose.
What runs on your card
| Your GPU | Companion writer | What you get versus the tier below |
|---|---|---|
| 8 GB (RTX 3060 8GB / 4060 / 5060) | 12B model | The full picture-book experience with a capable 12B writer. Chat and pictures take turns on the card, so a picture adds a few seconds. |
| 12 GB (RTX 3060 12GB / 4070 / 5070) | same 12B, higher precision | Full-fidelity image engine. The biggest value step — much better pictures, same strong writing. |
| 16 GB (RTX 4080 / 5080 / 5060 Ti 16GB) | same 12B | Speed and memory. On the Speed preset, chat and pictures stay loaded together (no waiting); Quality swaps between them so it can run the full-fidelity image engine. Long-term memory is on by default. |
| 24 GB (RTX 3090 / 4090) | 24B model | Richer writing. A larger model — noticeably better prose and scene-handling. Chat and full-quality pictures take turns, so a picture adds a few seconds. |
| 32 GB (RTX 5090) | 24B model (optional 31B) | The no-compromise tier. The same rich 24B writing as 24 GB, but pictures generate alongside the chat with no waiting. Prefer the richest possible prose? Switch on the larger 31B model in Advanced settings — it trades the instant pictures for a short pause each time. |
Fervyn auto-detects your VRAM and picks the right writer for your card — you don't configure anything. A tinkerer can swap in a different model from the Advanced settings.
Which card should I get?
- On a budget / already have an 8 GB card: start here. The writing is genuinely good now, and it's the tier we design and test against first. This is the recommended minimum.
- Best value — 12 GB: the same strong writing, plus the full-quality image engine. If you're buying a card for Fervyn, this is the step that gives the most for the money.
- The sweet spot — 16 GB: everything 12 GB gives, plus no waiting between chat and pictures and the long-term memory feature. The most fluid overall experience below enthusiast hardware.
- For richer writing — 24 GB: a bigger 24B model, where the prose itself meaningfully improves over the 12B tiers.
- For the smoothest experience — 32 GB (RTX 5090): the same rich 24B writing, but the pictures appear alongside the chat with no waiting between them — the only tier that runs full-quality writing and images together with no pause. Enthusiasts can switch on a 31B model for the richest prose (with a short pause per picture).
Bottom line: 8 GB is a real, recommended experience. 12 GB is the value pick. 16 GB is the sweet spot. 24 GB+ is for people who want the richest writing. See the full system requirements for the details on each tier.
How we tested this
We don't want to hand-wave "bigger is better," so we measured it. Every candidate model was driven through the same fixed set of role-play conversations on the real app — multi-character scenes that stress name-tracking, long conversations that stress memory, and format requests that stress instruction-following. The transcripts were then scored blind (the judges couldn't see which model produced which transcript) by multiple independent reviewers, plus direct head-to-head comparisons.
A few honest notes about the scores, so you can read them fairly:
- We scored on a deliberately harsh stress-test scale — a score isn't a school grade, it's how far a model is from flawless on hard, adversarial prompts. Everyday chats go far more smoothly than these tests.
- The relative ranking between tiers is the reliable result; the exact number for any single run varies, so we lean on head-to-head comparisons and the overall ladder, not decimals.
- Two things stay hard for every local model at this size — keeping a large cast's names perfectly straight, and obeying exact-format requests. Bigger models are better at both but none are perfect; we keep improving the prompts that help.
This reflects the models shipping in the current release; we re-test and update as better local models come out.