Skip to main content

The playbook for running voice-agent interview campaigns and turning them into an automation backlog for the FDE team.

Quick start

  1. Templates: start from the pre-seeded “Administrative Staff — General” or build one per staff type, then Publish.
  2. Check Settings has an active consent document (v1.0 is pre-seeded).
  3. Create a campaign: pick the template version + consent doc, set recording and anonymization policy.
  4. Upload a CSV of invitees, Activate the campaign, and send invites (or use Copy link to hand links out directly).
  5. Staff complete interviews on their own schedule; analysis runs automatically within minutes of each completion.
  6. Work the Backlog: open top opportunities, generate dossiers, curate clusters, export Markdown/CSV for the FDE team.

1 · Interview templates

A template is the agent's interview plan: sections (with a goal and a time budget) containing questions (with a goal, a required flag, and a probe depth). The agent works through sections in order, adapts wording naturally, and probes inside each section.

  • Probe depth per question: 0 = ask and move on · 1 = one clarifying follow-up · 2 = dig for specifics (numbers, system names) · 3 = up to three follow-ups until the goal is genuinely met.
  • Probing notes steer follow-ups, e.g. “always pin down a number per week and ask how many other people do this same work.”
  • Key terms boost transcription of campus jargon (Concur, TritonLink, EASy, and so on). Add any acronyms your unit uses.
  • Preview prompt shows the exact composed instructions the agent receives. Read it before publishing.
  • Publishing freezes a version. Campaigns pin a version, so editing later creates v2 and never changes live interviews.

Good targets: 4 to 6 sections, about 25 to 30 minutes total budget. Interviews longer than about 35 minutes degrade. Trim sections that the Analytics drop-off chart shows bleeding people.

California requires all-party consent, and this system evidences it twice: the interviewee accepts a consent screen (click, timestamped) and the agent reads a spoken script verbatim and records the verbal yes before any substantive question. Both are pinned to the consent document version in force, so you can always prove what someone agreed to.

Both checks are enforced by the server, not just by the agent's instructions. The agent cannot mark a section complete until spoken consent is on record, and the participant's own words are stored as evidence. If the agent ends a call without spoken consent, the interview is paused for re-confirmation rather than completed, and it is never analyzed. A verbal decline after the conversation has started withdraws the interview.

Manage versions in Settings. Documents are immutable once used. To change wording, create a new version, activate it, and use it for new campaigns.

3 · Campaigns

A campaign points one template version at one group of staff. Key settings (fixed at creation):

  • Audio recording on/off, plus an optional retention period in days. When the period passes, recordings are deleted and the raw per-call transcript text is removed from analyzed interviews and stored webhook payloads. The redacted transcript used for analysis is always kept.
  • Anonymize reports: dossiers, exports, and the Q&A refer to people as “role, department” only, and the fde role loses access to raw transcripts for this campaign.
  • Target length (the agent paces to it) and a hard cap per call (a safety stop; hitting it pauses the interview and never loses work).
  • Link expiry and the reminder schedule.

Status: draft, then active, then paused or closed. Invites only send and links only work while the campaign is active.

4 · Invitations & links

Upload a CSV with columns email, first_name (required) and last_name, role_title, department (recommended; role and department feed the analysis). Extra columns are kept as metadata.

  • Each person gets a personal magic link. No account, no scheduling. The same link serves the whole lifecycle: start, pause, resume, done.
  • Send emails pending invites; Remind re-sends (also automatic per the campaign schedule); Revoke kills a link immediately; Copy link rotates the token and puts a fresh URL on your clipboard for hand-delivery.
  • Every send/copy rotates the token, so older links die. Links expire per campaign setting, but every pause extends the link by seven days so nobody gets locked out mid-conversation.

5 · What interviewees experience

  1. Landing page: what this is, time estimate, consent notice, explicit agree.
  2. Mic check with a live level meter.
  3. Voice conversation with the agent: a visible roadmap ticks off topics as they're covered; live captions; mute, Pause (resume anytime on any device with the same link), and End controls. Asking the agent to pause works the same way as the button.
  4. Thank-you page + a receipt email on completion.

Built-in guardrails: the agent never asks for or repeats student names or identifiers, reassures anyone worried about evaluation, and steers tangents back to process. Connection drops are non-events: progress is saved and the link resumes where they left off. Only one call can be live per interview; opening the link on a second device takes over and ends the first call.

6 · From transcript to backlog (automatic)

On completion, each interview flows through six stages:

  1. Redact: identifiers are stripped before anything else reads the transcript (regex + an AI entity pass that runs on campus infrastructure and refuses to fall back off campus).
  2. Extract: each described work process becomes a structured record. Every number must be backed by a verbatim quote. A validator checks each quote against the transcript, requires the numbers in it to match exactly, and nulls any numeric field it cannot verify, so vague answers produce “unknowns,” not invented data.
  3. Score: annual hours, feasibility (0 to 5), confidence, and a route-to-solution (script / RPA / LLM agent / AI assist / redesign first).
  4. Embed: vectors for clustering and the Q&A.
  5. Cluster: “travel reimbursements” in Biology and “expense reports” in Chemistry merge into one canonical process. Borderline matches get an AI adjudication. Your merge, split, and rename decisions are keyed on the interview, process name, and department, so they survive re-analysis with new prompts or models.
  6. Roll up: campaign-level totals, the people-multiplier, priority ranking, and theme synthesis. Only successful runs count.

Interview statuses: completed, then processing, then analyzed. Errors are retried automatically (up to three attempts); interruptions such as a deploy or a timeout resume from the last completed stage without using an attempt. Failures surface on the Dashboard; the interview detail page shows the reason, any model fallback that was used, and a Re-run analysis button, which is also how you re-process everything after methodology changes.

7 · Reading the results (per campaign)

  • Backlog: canonical processes ranked by priority = org hours/yr × (feasibility/5) × confidence. An asterisk on hours means it includes a capped extrapolation from “N other people do this too” claims. Filter by route or department; export CSV.
  • Dossier (click any backlog row): the FDE handoff document. Baseline metrics per source, systems & access modes, exceptions, verified quotes, route rationale, and the “what we don't know yet” checklist that should drive the engineer's first week. Generate or regenerate on demand; export as Markdown. Curation lives here too: rename, merge into another cluster, or split a wrongly-grouped source out. These decisions persist through every re-cluster, and a rename follows the cluster through a merge.
  • Heatmap: departments × hours/pain/feasibility, showing where to focus.
  • Systems: manual hours flowing through each system. Rows spanning many departments are platform-level opportunities (one integration, many beneficiaries).
  • Themes: non-process findings (training gaps, policy confusion, shadow IT) synthesized across the campaign, with source breadth, evidence, requirement implications, and follow-up questions. Often cheaper to fix with a doc or a class than an automation.
  • Analytics: completion funnel, section drop-off (trim what bleeds people), minutes, and cost per completed interview.
  • Ask: free-form questions over the corpus (“what do people say about Concur?”). Answers cite their sources and refuse to speculate beyond the interviews.
  • Service Owner report (campaign page, export): the executive readout with participation, themes, backlog, systems, data-quality caveats, and a per-interview appendix, as Markdown or HTML.

Roles & permissions

Roles and permissions
RoleCan doCannot do
ownerEverything, including managing console accessDemote or disable the last active owner
adminTemplates, campaigns, invites, curation, re-runs, transcriptsAdd/remove console users
fdeBacklog, dossiers, aggregates, Ask; transcripts and participant contact details on non-anonymized campaignsRaw transcripts and names on anonymized campaigns; any configuration
viewerRead-only dashboards, backlog, heatmap, systems, themes, and aggregate exports (review packet, backlog CSV, name-masked Service Owner report)Any changes; raw transcripts; participant names and emails (rows show role + department instead); per-person and dossier exports on non-anonymized campaigns

Add people in Settings (owner only). They sign in with an emailed magic link; there are no passwords. Disabling someone signs them out on their next request, and the console always keeps at least one active owner.

Privacy & framing: the load-bearing rules

  • Never performance evaluation. It's promised in the consent document, the invite emails, and the agent's own reassurances, and the product enforces it by having no per-person metrics anywhere. Don't undermine the promise in how you talk about the program.
  • Frame campaigns as “take tedious work off your plate,” not “find efficiencies.” Candor depends on it.
  • PII is redacted before analysis, and the redaction step runs only on campus infrastructure. Still treat transcripts as sensitive (about UC P3): limit who has admin/fde access, use anonymize reports for sensitive units, and set a retention period rather than keeping recordings and raw transcripts forever.
  • Interviewing staff about their work is fine; deploying automations that change duties is a labor-relations event (HEERA). Loop in HR/LR before the FDE team ships changes.

Costs & capacity

  • Voice: roughly $0.10 to 0.20 per minute, so $3 to 7 for a 30-minute interview.
  • Analysis: runs on UCSD's on-prem Triton AI models by default. Transcripts stay on campus infrastructure and the marginal LLM cost is $0 (embeddings of redacted text cost fractions of a cent). On cloud models it is about $0.26 per interview, and every run has a configurable hard cost ceiling. Any use of a fallback or cloud model is recorded on the run and shown on the interview page.
  • A 500-person campaign lands around $2 to 4k total. Actuals show up per-interview on the Analytics page.
  • Concurrency: each live interview is one Vapi call (10 concurrent included; more is a per-line add-on). Staff interview on their own schedule, so bursts are rare in practice.

Troubleshooting

  • Interview stuck in “processing”: background sweeps retry automatically every few minutes. If it lands in “failed,” the interview page shows the reason and a Re-run analysis button (commonly a missing or expired Triton gateway key, or a model returning invalid output after every fallback).
  • Interview marked “needs review”: the assistant configuration Vapi ran did not match what the server minted (prompt, tools, duration cap, recording settings, or metadata). It is parked in failed and is not analyzed until an admin looks at it.
  • Interview paused “to re-confirm consent”: the agent ended the call before a verbal yes was recorded. The participant can resume with the same link; the agent re-reads the consent script and continues once consent is recorded.
  • Someone lost their link / it expired: use Remind or Copy link; both issue a fresh token.
  • Half-finished interviews: paused interviews auto-finalize after 48 hours of inactivity. Whatever was captured gets analyzed; empty ones close as abandoned.
  • Wrongly merged or split processes: fix it on the dossier page (merge/split/rename); your decision sticks through all future re-clustering.
  • No emails in local development: without an email provider key configured, admin sign-in links appear directly on the “check your email” page, and invite emails are replaced by Copy link. Real voice calls locally additionally need Vapi keys and a webhook tunnel (see the README).
  • Transcription mangles campus jargon: add the terms to the template's key-terms list and re-publish.