Corpus Insights
What we learned from 17 years of NHR work
We scanned every proposal, interview guide, and populated data-grid Bert has produced and reverse-engineered the conventions, phrasings, and structural patterns that define the NHR house style across three artifacts.
1,123
proposals analyzed
1,000
interview guides analyzed
56
paired interview-to-grid engagements
Proposals
1,123 proposals across 397 engagements
The anatomy of an NHR proposal
Top 10 section headers across the corpus. Eight of these appear in 80%+ of proposals — the load-bearing skeleton of every NHR document.
- BRIEF CONTEXT83.3% (935)
- GENERAL PROVISIONS81.2% (912)
- DELIVERABLES80.9% (908)
- BUDGET77.1% (866)
- PROJECT APPROACH AND SCOPE74.1% (832)
- WORKING WITH ‘NHR’59.4% (667)
- APPROXIMATE TIMING42.1% (473)
- APPROXIMATE TIMING AND RESOURCES40.4% (454)
- RESEARCH OBJECTIVES39.9% (448)
- PROJECT OBJECTIVES33.4% (375)
The canonical section order
115 proposals follow this exact 8-section sequence verbatim. Every NHR proposal we generate honors it.
- 1BRIEF CONTEXT
- 2RESEARCH OBJECTIVES
- 3PROJECT APPROACH AND SCOPE
- 4DELIVERABLES
- 5APPROXIMATE TIMING
- 6BUDGET
- 7WORKING WITH NHR
- 8GENERAL PROVISIONS
How Bert opens a proposal
The BRIEF CONTEXT section follows a remarkably consistent pattern. These three openings each appear 38+ times verbatim.
“The Riverside Company (“Riverside” or “you”) is contemplating a potential invest…”
appears in 38 documents
“High Street Capital (“HSC” or “you”) is contemplating a potential investment in …”
appears in 32 documents
“BV Investment Partners, LP (“BV” or “you”) is contemplating a potential investme…”
appears in 31 documents
The model treats this as a template: "{Sponsor full name}" ("{Short}" or "you") is contemplating a potential investment in {Target}…
Sample-size phrasings
The exact ways Bert describes interview and survey sample sizes. Surfaced to the extractor so it picks the right shape when a sponsor mentions a loose range.
3-5 interviews88×~15-20 interviews42×10-15 interviews41×~10-15 interviews21×10 interviews21×15 complete19×~20-25 interviews17×8-10 interviews16×Boilerplate fidelity
Two clauses are near-universal — confidentiality and indemnification — and our generator template includes both by default.
87.6%
of proposals include a Confidentiality clause
68.7%
of proposals include Indemnification language
Interview Guides
1,000 guides scanned
The skeleton of an NHR interview guide
Top section headers across past guides — the same blocks recur on every project Bert runs, in roughly the same order.
- CALL INTRODUCTION37.9% (379)
- INTERVIEW INFORMATION36.7% (367)
- CLOSING30.1% (301)
- KEY INTERVIEW OBJECTIVES27.3% (273)
- WRAP-UP25.9% (259)
- BACKGROUND INFO20.3% (203)
- DEMAND DRIVERS, TRENDS, & OUTLOOK14.2% (142)
- START THE INTERVIEW11% (110)
- DEMAND DRIVERS & TRENDS7.7% (77)
- BACKGROUND INFORMATION ON INTERVIEWEE & COMPANY7.4% (74)
The canonical NHR phone intro
65.1% of guides include an explicit "Hello, I'm calling from New Heights Research…" introduction. The top three variants:
“Hello, I’m calling from New Heights Research, an independent market research organization.…”
appears in 32 documents
“Hello, my name is ______ and I’m with New Heights Research, a market research firm based in Cleveland, OH.…”
appears in 22 documents
“Hello, my name is ______ and I’m with New Heights Research, an independent market research firm.…”
appears in 22 documents
Guide length, in questions
Most guides have between 25 and 75 questions — substantive enough to drive a 30-minute interview, lean enough to actually finish.
- 0-9 questions5.6% (56)
- 10-24 questions17.2% (172)
- 25-49 questions30.5% (305)
- 50-74 questions31.6% (316)
- 75-99 questions11.9% (119)
- 100-149 questions3.2% (32)
- 150+ questions0% (0)
Rating-scale instrumentation
13% of guides include at least one explicit rating-scale question ("on a scale of 1 to 10…", "please rate…"). When NHR needs to quantify perceptions, this is the shape it takes.
13%
of guides instrument at least one rating scale
Interview-to-Grid Extractor
56 engagements with paired transcripts + populated Excel grids; 160 grids sampled
The paired training corpus
Engagements where NHR has BOTH the raw transcripts AND the populated Excel grids that resulted from them. These pairings are the ground truth the i2g extractor learns from.
56
engagements with paired transcript + grid data
160
individual response grids analyzed
Pairing integrity
Confirms the transcripts are real phone interviews, not survey exports. Surveys never produce audio — but 45 of 56 paired engagements carry raw recordings (mp3 / m4a / wav). Written transcripts and audio recordings both appear in Speaker 1 / Speaker 2 conversational format on spot-check.
301
raw audio recordings
358
written transcripts (.doc / .docx / .txt)
45/56
engagements with audio in folder 5
57 non-conversational files (.pdf / .rtf) — typically survey export reports — are filtered out for i2g training and eval.
What's actually captured per interview
Filled cells fall into four shape buckets. The distribution tells us what the model has to produce per row — predominantly short phrases (names, titles, codes, single-digit ratings), with substantial sentence-length and paragraph-length analytical responses where Bert summarizes what an interviewee actually said.
- Code / rating (≤4 chars, numeric or Y/N)20.3%
- Short phrase (≤30 chars)48.4%
- Sentence (≤150 chars)15.6%
- Paragraph (>150 chars)15.6%
The row labels that recur
Most-common column-B labels across sampled grids. Identity rows (Interviewer, Company Name, Title) lead, followed by the standardized rating sub-rows that expand under every factor (Importance / Satisfaction rating + paired Comments), and analytical buckets (KEY TAKEAWAYS, Strengths, Weaknesses).
Comments1489Importance rating456Satisfaction rating331Importance comments294Satisfaction comments287NHR ID (from Interview Log)158Interviewer158Company Name158Interviewee Name158Interviewee Title158State158# absolute157How many rows per grid?
Counts every non-empty row in column B — a mix of three structurally different things: (1) section headers (BACKGROUND, RATINGS, WRAP-UP), (2) top-level interview questions, and (3) per-factor rating sub-rows. A 25-question interview with a 20-factor rating block expands to ~25 + 20×4 = ~105 rows from those structures alone, which is why most grids land in the 120+ bucket. The i2g extractor reads the engagement's blank template to learn which rows mean what.
- 0-19 rows0% (0)
- 20-49 rows1.9% (3)
- 50-79 rows5.6% (9)
- 80-119 rows8.1% (13)
- 120+ rows84.4% (135)
How we use this corpus
- System prompt grounding. Each extractor's system prompt encodes canonical section names, intro phrasings, and structural conventions so the model writes in NHR style by default.
- Few-shot exemplars. Real (input → golden output) pairs from past engagements (Lauxera NEMT, Oaktree / Sembi) are embedded as multi-turn examples in every extraction call.
- Conformance tests. A pytest harness checks every generated artifact against the corpus patterns — canonical section sequence, boilerplate phrases, three-tier objective shape, duration-phrase timing — catching regressions before they ship.
- Template alignment. Word templates use the dominant section names from the corpus (BUDGET, APPROXIMATE TIMING AND RESOURCES, WORKING WITH ‘NHR’ with curly quotes) rather than generic placeholders.
- Phase 2 i2g extractor. The paired transcript-to-grid corpus is the ground truth we'll evaluate the upcoming Interview-to-Grid extractor against — a real eval set, not synthetic data.
Last refreshed — proposals: May 21, 2026 · guides: May 22, 2026 · interview-to-grid: May 22, 2026. Re-runnable via scripts/analyze_*_corpus.py