How many samples you actually need, what makes a great sample, and the most common mistakes that quietly poison your voice training.
Jump to module

The honest answer: more than you think, less than you fear.
Below 10 samples, the system can't separate signal from noise. Your "voice" looks like 5 specific posts on 5 specific topics, and the agent generalizes badly.
Around the 20-sample mark, the system starts seeing the patterns across samples — repeated hooks, repeated rhythms, repeated avoidances. Below that, it sees individual posts. Above that, it sees the writer.
The medium matters less than whether the writing actually sounds like you.
Not all samples are equal. The right 20 samples beat the wrong 50.
Aim for a mix:
Pick samples that you would post again today. Not your old voice. Not posts you cringe at. The system trains on what you give it — give it what you want it to produce.
Mix topics. If all 25 samples are about productivity, your agent will gravitate toward productivity even when you want it to write about something else. Voice is mechanics, but if topics are constant in training, the model lazily ties them to voice.
Most weak voice profiles aren't underfed. They're trained on the wrong things.
You upload your social posts (casual founder voice), your customer emails (slightly more formal), and your About page (corporate). The agent averages them. Output is muddy.
Fix: train per voice. Upload only the samples that match the voice you want this profile to produce.
You used another AI tool last quarter, and some of those posts are in your "samples" pile. The agent now trains on AI-flavored samples and produces more AI-flavored content.
Fix: only include posts you wrote yourself, by hand. AI-edited is fine. AI-written is not.
You sound different now than you did 2 years ago. If your samples are from 2024, your voice profile is preserving an out-of-date version of you.
Fix: prefer recent samples. If you must include older ones, weight recent samples 70%+.
A blog post that quotes a tweet you wrote, then a customer testimonial, then your own writing. The system can't tell whose voice it's training on.
Fix: strip quotations. Train on raw your-writing only.
You upload your top 25 highest-engagement posts. They're all hot takes. Your agent now writes only hot takes.
Fix: include some posts that didn't go viral. Include some that are thoughtful, slow, contemplative. The agent should learn your range, not just your peaks.
You retrain voice every two weeks because output drifted. The agent never gets stable. Output is inconsistent.
Fix: voice should be stable for at least 60-90 days. Drift is usually a persona problem, not a voice problem.
Sample-trained voice is the highest-fidelity option you have. It just requires picking the right samples. Next lesson: training from connected socials — when it's the right call.