Lab 02 · Chapter 3 — Run, test, diagnose, and deploy
← File-by-file build · Português · Overview
First prove everything that needs no network. Only then open one short microphone session while monitoring API usage.
1. Confirm the directory and secret
pwd
git check-ignore -v .env.local
The directory must end in labs/lab-02-realtime-voice-agent, and Git must print the rule that ignores .env.local.
2. Run offline gates
npm run lint
npm run typecheck
npm test
npm run build
npm run check
These commands request no microphone and issue no client secret. Fix the first failure before continuing.
3. Start without enabling the microphone
npm run dev
Open http://localhost:3000, but do not click Start live conversation yet. In another terminal:
cd openai-voice-playground/labs/lab-02-realtime-voice-agent
curl http://localhost:3000/api/health
PowerShell:
Invoke-RestMethod http://localhost:3000/api/health
The response must report ok: true, configured: true, model gpt-realtime-2.1, transport webrtc, and a 60-second issuance TTL. It must contain neither the API key nor a client secret.
4. Run one short Realtime smoke test
Use headphones and content with no personal data.
- Read the AI and privacy notice.
- Select the required consent.
- Choose language and voice.
- Click Start live conversation.
- Allow microphone access.
- Say one short sentence.
- Wait for a response.
- Interrupt once by speaking during the response.
- Toggle mute and verify the UI state.
- Send one text fallback message.
- Click End.
- Confirm the browser microphone indicator disappears.
Do not leave the tab connected. Explicit ending remains part of the test.
5. Troubleshoot by symptom
| Symptom | Likely cause | Diagnose | Fix | Confirm |
|---|---|---|---|---|
| packages missing or Node incompatible | no install, wrong directory, or Node < 22 | pwd, node --version, npm ls --depth=0 |
enter Lab 02 and run npm ci with Node.js 22+ |
npm run typecheck exits zero |
| port 3000 busy | another local server is running | inspect EADDRINUSE; use lsof -i :3000 or Get-NetTCPConnection -LocalPort 3000 |
stop the known process or use npm run dev -- --port 3001 |
the Next.js URL opens |
configured: false |
missing .env.local, invalid key, or no restart |
git check-ignore -v .env.local and curl localhost:3000/api/health |
use OPENAI_API_KEY=..., save, restart |
health reports configured: true without a credential |
| unavailable quota/credits or model | project billing, limit, or no gpt-realtime-2.1 access |
inspect request ID, Usage/Billing, and health model | enable billing/limit or consistently use an allowed model | one short planned smoke test connects within budget |
| browser never requests microphone | denied permission, insecure context, or no user gesture | inspect site permissions, navigator.mediaDevices, and console |
use HTTPS/localhost, allow permission, click again | microphone indicator appears only during the session |
| unsupported browser | WebRTC/media APIs unavailable or restricted | try current Chrome, Edge, Firefox, or Safari and inspect console | update/switch browser; avoid restricted webviews | session reaches connected state |
| client secret expires before connect | issued too early or delayed beyond TTL | inspect issuance time without logging the value | issue immediately before connect; never reuse it |
retry connects and response remains no-store |
| WebRTC fails | firewall, VPN, corporate network, HTTPS, or negotiation | inspect chrome://webrtc-internals, console, and Network without copying tokens |
try another network, remove a known VPN, confirm HTTPS, retry briefly | audio flows; End releases the connection |
403 cross_origin_request or CORS |
APP_ORIGIN differs from the domain |
compare browser origin with the protected variable | correct the full origin and redeploy | right origin gets a secret; another remains blocked |
| agent hears itself | speaker feedback reaches microphone | use headphones and observe unexpected turns | keep headphones and choose the right noise-reduction profile | agent responds only to the learner |
| transcript duplicates | history handled as append rather than snapshot | observe repeated items after history events | reconcile snapshots and keep state in memory | each turn appears once and disappears on refresh |
| microphone stays active | cleanup did not close session/tracks | click End and inspect the system indicator | call session.close() and stop every track on cleanup/unmount |
indicator disappears and a new session starts cleanly |
| build/import fails in CI | filename casing or stale generated guide | git ls-files | sort, npm run docs:check, npm run check |
match import and filename exactly; regenerate docs after displayed-code changes | local and remote checks pass |
For stale cache, Pages workflow, or generated documentation issues, use the shared troubleshooting guide.
Compare without replacing files:
git fetch origin
git diff --stat HEAD..origin/workshop/lab-02-v1-step-03-conversation
6. Commit
cd ../..
git status -sb
git add labs/lab-02-realtime-voice-agent
git commit -m "feat: complete realtime voice workshop"
.env.local, .next, node_modules, and *.tsbuildinfo must stay out.
7. Deploy the application to Vercel
GitHub Pages will host the tutorials. It cannot run /api/realtime/token or protect OPENAI_API_KEY; the application needs a server-side host such as Vercel.
- Import the repository into Vercel.
- Set Root Directory to
labs/lab-02-realtime-voice-agent. - Add:
OPENAI_API_KEY=protected_value
PLAYGROUND_ACCESS_TOKEN=a_long_random_phrase
APP_ORIGIN=https://your-domain.example
UPSTASH_REDIS_REST_URL=protected_value
UPSTASH_REDIS_REST_TOKEN=protected_value
Leave CLIENT_IP_HEADER absent on Vercel. Elsewhere, name the header overwritten by the trusted proxy.
- Deploy.
- Validate
/api/healthwithout secrets. - Confirm HTTPS and microphone permission.
- Run one short conversation.
- End explicitly and monitor usage/budget.
8. Final checklist
npm run checkpasses;- health contains no credential;
- the client secret is short-lived and
no-store; - connection, interruption, mute, text, and End work;
- the microphone is released;
- transcript does not persist after refresh;
- production has access, origin, and distributed quota controls;
- budget and alerts are active.
Done. Read the architecture article for deeper WebRTC, state-machine, retention, abuse, and server-side-control reasoning.