plate 03 · the study
A voice measured against itself
73 raw transcripts. 358,847 words, recorded between 2023 and 2026. Counted rather than characterised, because a style guide written from impressions is a style guide about whoever wrote it.
method
Raw only
Each recording folder holds a raw transcript, a cleaned one, and an extracted concept file. Only the raw one is speech.
The cleaned and extracted files have been through a machine. Counting them would have measured model phrasing and reported it as his, which is the exact failure a voice profile exists to prevent. Fifty-eight raw files went in. Twenty-seven cleaned and fifteen concept files were excluded.
Diarization metadata was stripped before counting, because speaker labels from multi-person recordings would otherwise have registered as vocabulary.
vocabulary
Short verbs, plain nouns
Rates per ten thousand words. The verbs are physical and most of them are one syllable.
| Verb | Per 10k |
|---|---|
| get | 38.5 |
| need | 28.4 |
| want | 22 |
| make | 17.8 |
| build | 13.2 |
| work | 15.9 |
| use | 12.8 |
| help | 11.7 |
| take | 9.5 |
| talk | 7.7 |
build, building, and built together run 28.7 per ten thousand. Nothing in the corpus gets facilitated or optimised into being. It gets built, or set up, or figured out.
The words he reaches for, and the ones he does not
| His word | Count | The formal alternative | Count |
|---|---|---|---|
| stuff | 835 | material | 16 |
| thing | 2183 | element | 17 |
| people | 1030 | individuals | 4 |
| tool | 371 | solution | 26 |
| way | 555 | methodology | 7 |
| problem | 92 | challenge | 24 |
the negative result
Sixteen words that never appear
Counts across the whole 358,847 words. The house ban on this vocabulary was written before the corpus was measured, and the corpus agrees with it.
- beacon0
- circle back0
- delve0
- intricate0
- meticulous0
- paramount0
- robust0
- synergy0
- tapestry0
- testament0
- underscore0
- vibrant0
- facilitate1
- lean into1
- pivotal1
- unpack1
transitions
He moves by conjunction, never by announcement
A third of his sentences, 34.6 percent, open with a conjunction. and appears first 2956 times, so 2011 times, but 660.
"and then" is the engine at 74.8 per ten thousand. It sequences everything. "anyway" closes a tangent 291 times and does the work a paragraph break does on the page.
What is missing is signposting. "the thing is" appears 6 times in358,847 words. "at the end of the day" 7. He does not announce a turn, he takes it.
The questions are practical rather than rhetorical. "how do I" 58, "how do we" 65, "how might" zero.
Take away the conjunction openers and it stops sounding like him.
Vocabulary can be matched word for word and the result still reads wrong if every sentence starts with its own subject. The habit is structural rather than lexical, and it is the first thing generated copy loses.
rhythm
Median 12 words
Mean 22.1, median 12, quartiles at 6 and 21. The distribution matters more than the average, because it is not a bell curve.
| Sentence length | Share |
|---|---|
| 1 to 5 words | 21.3% |
| 6 to 15 | 40.1% |
| 16 to 30 | 25.3% |
| 31 to 60 | 10.7% |
| past 60 | 2.6% |
Nearly a fifth land at five words or fewer, and one in forty runs past sixty. Copy sitting at a uniform eighteen words reads as nothing like him even when every word is drawn from his vocabulary.
the trap
The two most common things in the corpus are filler
"like" runs at 270.3 per ten thousand. "you know" at 129.7. Both dwarf every content word he uses, and a naive frequency match would produce a parody.
They are worth measuring anyway, because they mark the line between his spoken voice and his written one. Vocabulary and sentence shape transfer directly. The filler and the sixty-word comma splice do not, and a style guide that fails to separate them is worse than no style guide.
upstream
Where these transcripts come from
Every file counted here is a stage 01 artifact from the ingest pipeline, preserved by digest and immutable once written.
Stage 01 raw transcripts only. Frontmatter, pipeline headings and diarization stamps removed before counting. Sentence-initial counts are first-word occurrences, not totals. Counted across 73 recordings, 2023-09-09-01 to 2026-09-02-02, 16,278 sentences in total. Regenerated 2026-09-02.
These are numbers, not quotations. Nothing on this page is something he said about himself. It is what the recordings do, measured, and it is reproducible: the same script over the same files returns the same figures, and returns different ones as the corpus grows.