Corpus Insights

What we learned from 17 years of NHR work

We scanned every proposal, interview guide, and populated data-grid Bert has produced and reverse-engineered the conventions, phrasings, and structural patterns that define the NHR house style across three artifacts.

1,123

proposals analyzed

1,000

interview guides analyzed

56

paired interview-to-grid engagements

Proposals

1,123 proposals across 397 engagements

The anatomy of an NHR proposal

Top 10 section headers across the corpus. Eight of these appear in 80%+ of proposals — the load-bearing skeleton of every NHR document.

  • BRIEF CONTEXT
    83.3% (935)
  • GENERAL PROVISIONS
    81.2% (912)
  • DELIVERABLES
    80.9% (908)
  • BUDGET
    77.1% (866)
  • PROJECT APPROACH AND SCOPE
    74.1% (832)
  • WORKING WITH ‘NHR’
    59.4% (667)
  • APPROXIMATE TIMING
    42.1% (473)
  • APPROXIMATE TIMING AND RESOURCES
    40.4% (454)
  • RESEARCH OBJECTIVES
    39.9% (448)
  • PROJECT OBJECTIVES
    33.4% (375)

The canonical section order

115 proposals follow this exact 8-section sequence verbatim. Every NHR proposal we generate honors it.

  1. 1BRIEF CONTEXT
  2. 2RESEARCH OBJECTIVES
  3. 3PROJECT APPROACH AND SCOPE
  4. 4DELIVERABLES
  5. 5APPROXIMATE TIMING
  6. 6BUDGET
  7. 7WORKING WITH NHR
  8. 8GENERAL PROVISIONS

How Bert opens a proposal

The BRIEF CONTEXT section follows a remarkably consistent pattern. These three openings each appear 38+ times verbatim.

“The Riverside Company (“Riverside” or “you”) is contemplating a potential invest…”

appears in 38 documents

“High Street Capital (“HSC” or “you”) is contemplating a potential investment in …”

appears in 32 documents

“BV Investment Partners, LP (“BV” or “you”) is contemplating a potential investme…”

appears in 31 documents

The model treats this as a template: "{Sponsor full name}" ("{Short}" or "you") is contemplating a potential investment in {Target}…

Sample-size phrasings

The exact ways Bert describes interview and survey sample sizes. Surfaced to the extractor so it picks the right shape when a sponsor mentions a loose range.

3-5 interviews88×~15-20 interviews42×10-15 interviews41×~10-15 interviews21×10 interviews21×15 complete19×~20-25 interviews17×8-10 interviews16×

Boilerplate fidelity

Two clauses are near-universal — confidentiality and indemnification — and our generator template includes both by default.

87.6%

of proposals include a Confidentiality clause

68.7%

of proposals include Indemnification language

Interview Guides

1,000 guides scanned

The skeleton of an NHR interview guide

Top section headers across past guides — the same blocks recur on every project Bert runs, in roughly the same order.

  • CALL INTRODUCTION
    37.9% (379)
  • INTERVIEW INFORMATION
    36.7% (367)
  • CLOSING
    30.1% (301)
  • KEY INTERVIEW OBJECTIVES
    27.3% (273)
  • WRAP-UP
    25.9% (259)
  • BACKGROUND INFO
    20.3% (203)
  • DEMAND DRIVERS, TRENDS, & OUTLOOK
    14.2% (142)
  • START THE INTERVIEW
    11% (110)
  • DEMAND DRIVERS & TRENDS
    7.7% (77)
  • BACKGROUND INFORMATION ON INTERVIEWEE & COMPANY
    7.4% (74)

The canonical NHR phone intro

65.1% of guides include an explicit "Hello, I'm calling from New Heights Research…" introduction. The top three variants:

“Hello, I’m calling from New Heights Research, an independent market research organization.…”

appears in 32 documents

“Hello, my name is ______ and I’m with New Heights Research, a market research firm based in Cleveland, OH.…”

appears in 22 documents

“Hello, my name is ______ and I’m with New Heights Research, an independent market research firm.…”

appears in 22 documents

Guide length, in questions

Most guides have between 25 and 75 questions — substantive enough to drive a 30-minute interview, lean enough to actually finish.

  • 0-9 questions
    5.6% (56)
  • 10-24 questions
    17.2% (172)
  • 25-49 questions
    30.5% (305)
  • 50-74 questions
    31.6% (316)
  • 75-99 questions
    11.9% (119)
  • 100-149 questions
    3.2% (32)
  • 150+ questions
    0% (0)

Rating-scale instrumentation

13% of guides include at least one explicit rating-scale question ("on a scale of 1 to 10…", "please rate…"). When NHR needs to quantify perceptions, this is the shape it takes.

13%

of guides instrument at least one rating scale

Interview-to-Grid Extractor

56 engagements with paired transcripts + populated Excel grids; 160 grids sampled

The paired training corpus

Engagements where NHR has BOTH the raw transcripts AND the populated Excel grids that resulted from them. These pairings are the ground truth the i2g extractor learns from.

56

engagements with paired transcript + grid data

160

individual response grids analyzed

Pairing integrity

Confirms the transcripts are real phone interviews, not survey exports. Surveys never produce audio — but 45 of 56 paired engagements carry raw recordings (mp3 / m4a / wav). Written transcripts and audio recordings both appear in Speaker 1 / Speaker 2 conversational format on spot-check.

301

raw audio recordings

358

written transcripts (.doc / .docx / .txt)

45/56

engagements with audio in folder 5

57 non-conversational files (.pdf / .rtf) — typically survey export reports — are filtered out for i2g training and eval.

What's actually captured per interview

Filled cells fall into four shape buckets. The distribution tells us what the model has to produce per row — predominantly short phrases (names, titles, codes, single-digit ratings), with substantial sentence-length and paragraph-length analytical responses where Bert summarizes what an interviewee actually said.

  • Code / rating (≤4 chars, numeric or Y/N)
    20.3%
  • Short phrase (≤30 chars)
    48.4%
  • Sentence (≤150 chars)
    15.6%
  • Paragraph (>150 chars)
    15.6%

The row labels that recur

Most-common column-B labels across sampled grids. Identity rows (Interviewer, Company Name, Title) lead, followed by the standardized rating sub-rows that expand under every factor (Importance / Satisfaction rating + paired Comments), and analytical buckets (KEY TAKEAWAYS, Strengths, Weaknesses).

Comments1489Importance rating456Satisfaction rating331Importance comments294Satisfaction comments287NHR ID (from Interview Log)158Interviewer158Company Name158Interviewee Name158Interviewee Title158State158# absolute157

How many rows per grid?

Counts every non-empty row in column B — a mix of three structurally different things: (1) section headers (BACKGROUND, RATINGS, WRAP-UP), (2) top-level interview questions, and (3) per-factor rating sub-rows. A 25-question interview with a 20-factor rating block expands to ~25 + 20×4 = ~105 rows from those structures alone, which is why most grids land in the 120+ bucket. The i2g extractor reads the engagement's blank template to learn which rows mean what.

  • 0-19 rows
    0% (0)
  • 20-49 rows
    1.9% (3)
  • 50-79 rows
    5.6% (9)
  • 80-119 rows
    8.1% (13)
  • 120+ rows
    84.4% (135)

How we use this corpus

  • System prompt grounding. Each extractor's system prompt encodes canonical section names, intro phrasings, and structural conventions so the model writes in NHR style by default.
  • Few-shot exemplars. Real (input → golden output) pairs from past engagements (Lauxera NEMT, Oaktree / Sembi) are embedded as multi-turn examples in every extraction call.
  • Conformance tests. A pytest harness checks every generated artifact against the corpus patterns — canonical section sequence, boilerplate phrases, three-tier objective shape, duration-phrase timing — catching regressions before they ship.
  • Template alignment. Word templates use the dominant section names from the corpus (BUDGET, APPROXIMATE TIMING AND RESOURCES, WORKING WITH ‘NHR’ with curly quotes) rather than generic placeholders.
  • Phase 2 i2g extractor. The paired transcript-to-grid corpus is the ground truth we'll evaluate the upcoming Interview-to-Grid extractor against — a real eval set, not synthetic data.

Last refreshed — proposals: May 21, 2026 · guides: May 22, 2026 · interview-to-grid: May 22, 2026. Re-runnable via scripts/analyze_*_corpus.py